LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings
This paper introduces LongWebBench, a comprehensive benchmark designed to evaluate the structural fidelity and functional executability of long-horizon webpage generation by vision-language models, revealing that current state-of-the-art systems struggle with long-range coherence and interactive capabilities despite visual plausibility.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a talented architect to build a house based on a single, massive photograph of a mansion.
The Problem:
Until now, most tests for AI "architects" (called Vision-Language Models) only asked them to build tiny, single-room cottages. The tests checked if the front door looked right or if the paint color matched. But real websites are like massive mansions with dozens of rooms, multiple floors, and complex wiring. If you ask an AI to build a whole mansion from one photo, it might get the front door right but forget how to connect the stairs to the second floor, or build a kitchen that looks real but has no working sink.
The Solution: LongWebBench
The authors of this paper created a new, tougher test called LongWebBench. Instead of just checking if the house looks like the photo, they check two things:
- The Blueprint (Structure): Does the whole building make sense? Are the rooms in the right order? Is the roof connected to the walls?
- The Plumbing & Electricity (Function): If you turn on the faucet, does water come out? If you flip the light switch, does the light turn on?
How They Built the Test:
They gathered two huge collections of real-world "mansion photos" (long webpages):
- 490 "Look-Alike" Tests: They took photos of very long websites (some so long you'd have to scroll for 30 screens to see the bottom). They asked AI models to recreate the code for these pages. The test checked if the AI could keep the layout consistent from the top of the page to the very bottom.
- 507 "Do-It" Tests: They picked 129 websites and gave the AI specific jobs, like "Find a flight to Tokyo," "Add an item to the cart," or "Filter the search results." The AI had to write code that actually did these things in a real web browser, not just pretend to do them.
The Results: The "Reality Gap"
When they ran the best AI models through this new test, they found some surprising things:
- The "Long Scroll" Problem: As the websites got longer, the AI's ability to keep the structure correct got worse. It's like an architect who can design a perfect first floor but forgets how the second floor connects to it.
- The "Fake House" Problem: Many AIs built websites that looked beautiful and realistic (like a movie set), but when you tried to click a button or fill out a form, nothing happened. The "plumbing" was broken.
- The Input Issue: Even when the researchers broke the long photo into smaller pieces to make it easier for the AI to see, the AI still struggled to put the pieces together into a working whole.
The Takeaway
The paper concludes that we can't just judge AI by how pretty its drawings look. To be truly useful, an AI needs to be able to build long, complex websites that actually work when you interact with them. LongWebBench is the new ruler they are using to measure this, showing us that while AI is getting good at drawing, it still has a long way to go before it can build a fully functional digital house.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.