← Latest papers
🤖 AI

MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence

The paper introduces MMArch, a rigorous benchmark for architecture and civil engineering that evaluates multimodal large language models' ability to synthesize distributed visual evidence with engineering principles, revealing a significant performance gap between current AI systems and human experts.

Original authors: Chenxu Du, Kang An, Tengyue Wang, Zhongyu Yang, Xinqi Yang, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren

Published 2026-08-11
📖 6 min read🧠 Deep dive

Original authors: Chenxu Du, Kang An, Tengyue Wang, Zhongyu Yang, Xinqi Yang, Yuanchi Zhu, Hebao Zhu, Ziliang Wang, Faqiang Qian, Yunli Yang, Qibing Ren

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to be an architect. You don't just want it to recognize that a picture shows a house; you want it to understand why the house won't fall down. This is the world of "Multimodal Large Language Models" (MLLMs). Think of these models as super-smart digital brains that can read text and look at pictures at the same time. They are great at describing what they see, like saying, "That's a blue door," or "This chart shows a line going up." But in the real world of engineering, knowing what you see isn't enough. You have to connect the dots between a drawing, a physics rule, and a safety code to make a judgment. It's like being a detective who has to look at a crime scene photo, remember a specific law, and then decide if the suspect is guilty. The big question researchers have been asking is: Can these AI detectives actually solve the case, or are they just guessing based on the clues they see?

This is where a new study called MMArch steps in. The researchers built a giant, tricky test specifically for architecture and civil engineering to see if AI can really "think" like an engineer. They didn't just ask the AI to describe a picture; they gave it a puzzle where the answer was hidden across different parts of a drawing and required a specific rule of physics to solve. The results were a bit of a wake-up call. While the smartest AI models could get about half the questions right, a team of human experts got almost all of them right. The study found that the AI isn't failing because it can't "see" the picture; it's failing because it can't connect the picture to the right rule and then do the math to get the answer. It's like the AI can read the menu and see the ingredients, but it can't figure out how to cook the dish.

The Big Test: MMArch

The researchers created a benchmark called MMArch (Multimodal Architecture). Think of it as a final exam for AI, but instead of multiple-choice questions where you can just guess, it's a short-answer test. The exam is made up of 1,212 questions pulled directly from real, peer-reviewed scientific papers about buildings, bridges, and cities. These aren't made-up pictures; they are the actual diagrams engineers use to prove their designs work.

The test covers 10 different areas, from how buildings shake during earthquakes to how we design digital models of cities. To make sure the AI couldn't use shortcuts, the researchers used a clever "planner-writer" system. First, a "planner" looked at a paper and picked a specific answer (like "Rebar #6 and #7"). Then, a "writer" created a question that required looking at the picture and knowing the engineering rule to find that answer. The AI couldn't just read the text or look at one small part of the image; it had to look at the whole picture, find the right clues, and apply a rule to solve it.

The Results: AI vs. Humans

When they ran 18 different AI models (both free, open-source ones and expensive, top-tier commercial ones) through this test, the results were clear: AI is still a long way from being a real engineer.

  • The Human Experts: A panel of real architects and engineers got 94.6% of the answers right. They aced the test.
  • The Best AI: The smartest commercial AI model (GPT-5.5) got only 51.7% right.
  • The Open-Source AI: The best free model managed about 30%.

That's a gap of more than 40 percentage points. Even the best AI is barely passing, while humans are getting near-perfect scores. The researchers checked to make sure the AI wasn't just guessing or reading the text instead of the picture, and they confirmed that the AI really was struggling with the reasoning part.

Why Is AI Struggling?

The team didn't just stop at the scores; they looked at why the AI got things wrong. They found that the problem wasn't that the AI couldn't "see" the picture. The AI was actually pretty good at spotting the right lines and numbers in the diagrams. The real trouble happened when the AI tried to combine what it saw with what it knew.

They broke the mistakes down into five types:

  1. Perception Errors (20.3%): The AI just misread the picture (like confusing two similar lines).
  2. Principle Errors (21.7%): The AI saw the picture right but used the wrong rule (like thinking a wide loop in a graph meant the material was getting weaker, when it actually meant it was getting stronger).
  3. Composition Errors (35.8%): This was the biggest problem. The AI found the right numbers and knew the right rule, but when it tried to put them together to get the final answer, it dropped a step or messed up the math. It's like having all the ingredients and the recipe, but burning the cake because you forgot to mix them in the right order.
  4. Grounding Errors (14.5%): The AI knew what to look for but read the number wrong (like saying +138 instead of -142).
  5. Consistency Errors (7.7%): The AI figured out the right answer in its "brain" but wrote down the wrong one.

The biggest takeaway is that Composition Errors made up more than a third of all mistakes. This means the AI's biggest weakness is connecting the dots. It can see the evidence, and it can recall the rule, but it can't reliably put them together to solve a complex problem.

Does "Thinking Step-by-Step" Help?

You might have heard that telling AI to "think step-by-step" (a technique called Chain-of-Thought) helps it solve hard problems. The researchers tried this on MMArch. They asked the AI to explain its reasoning before giving the answer. The result? It didn't help much. For some models, it made them slightly better; for others, it actually made them worse. This suggests that the problem isn't that the AI needs to be told how to think, but that it simply doesn't know enough about the specific rules of architecture and engineering to build a correct chain of thought in the first place.

What's Next?

The study concludes that we can't just make AI bigger or give it more data to fix this. The issue isn't a lack of size; it's a lack of deep, specialized reasoning. The AI needs to learn how to bind visual evidence to domain knowledge in a way that current models just can't do yet.

MMArch is now available for other researchers to use as a training ground. It's a diagnostic tool to help us see exactly where AI is failing so we can build better models for the future. Until then, if you need to design a skyscraper that won't fall down in an earthquake, you're still better off hiring a human expert than asking a robot.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →