WebCompass A unified multimodal benchmark for evaluating LLMs' ability to generate, edit, and repair functional web pages. WebCompass spans three input modalities — text design documents, reference screenshots, and video demonstrations — and three task families — generation , editing , and repair . GitHub : NJU LINK/WebCompass Project Page : nju link.github.io/WebCompass Quick Start For editing/repair, the JSONL records carry the source code as text. The reference screenshots and any binary resources (logos, images, fonts) ship as parallel asset files inside editing/{sp,mp}/{instance id}/... and repair/{sp,mp}/{instance id}/... — you can fetch them with huggingface hub.snapshot download (see the GitHub repo for a ready to use download from hf.py that rebuilds the local file tree expected by the evaluator). Dataset Structure Task Types Task Description Configs Generation Generate a web page from scratch text generation , image generation , video generation Editing Add new features to an existing site editing (splits sp , mp ) Repair Fix a broken site to match a target screenshot repair (splits sp , mp ) Configs Config Split Samples Description text generation train 123 Generate from…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy