← Latest papers
💻 computer science

ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

This paper introduces ClawMark, a new benchmark designed to evaluate multi-turn, multi-day, multimodal coworker agents in a dynamic, stateful environment, revealing that while frontier models show partial progress, they struggle significantly with adapting to exogenous changes and achieving complete end-to-end task success.

Original authors: Fanqing Meng, Lingxiao Du, Zijian Wu, Guanzheng Chen, Xiangyan Liu, Jiaqi Liao, Chonghe Jiang, Zhenglin Wan, Jiawei Gu, Pengfei Zhou, Rui Huang, Ziqi Zhao, Shengyuan Ding, Ailing Yu, Bo Peng, Bowei Xi
Published 2026-04-28
📖 4 min read☕ Coffee break read

Original authors: Fanqing Meng, Lingxiao Du, Zijian Wu, Guanzheng Chen, Xiangyan Liu, Jiaqi Liao, Chonghe Jiang, Zhenglin Wan, Jiawei Gu, Pengfei Zhou, Rui Huang, Ziqi Zhao, Shengyuan Ding, Ailing Yu, Bo Peng, Bowei Xia, Hao Sun, Haotian Liang, Ji Xie, Jiajun Chen, Jiajun Song, Liu Yang, Ming Xu, Qionglin Qiu, Runhao Fu, Shengfang Zhai, Shijian Wang, Tengfei Ma, Tianyi Wu, Weiyang Jin, Yan Wang, Yang Dai, Yao Lai, Youwei Shu, Yue Liu, Yunzhuo Hao, Yuwei Niu, Jinkai Huang, Jiayuan Zhuo, Zhennan Shen, Linyu Wu, Cihang Xie, Yuyin Zhou, Jiaheng Zhang, Zeyu Zheng, Mengkang Hu, Michael Qizhe Shieh

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new assistant to help you run your office. Most tests for AI assistants today are like a pop quiz: they give the AI a single question, it answers, and the test is over. The world is frozen in time during that quiz.

But in real life, an assistant doesn't just answer one question and leave. They stay with you for days. While they are working, the world keeps changing around them. New emails arrive, your calendar shifts, files get updated, and new evidence (like photos or voice notes) pops up. If your assistant doesn't notice these changes, they might make decisions based on old information, leading to mistakes.

ClawMark is a new "test drive" designed specifically to see if AI assistants can handle this messy, changing reality.

The "Living Office" Test

Instead of a frozen quiz, ClawMark is like a simulated office that breathes.

  • The Setup: The AI is given a job (like handling an insurance claim or managing a project) that spans several "working days."
  • The Twist: Between each day, the environment changes on its own. Maybe a new email arrives, a spreadsheet gets updated, or a silent file is dropped in the folder without anyone telling the AI.
  • The Evidence: The AI isn't just reading text. It has to look at raw photos, listen to audio recordings, read scanned PDFs, and analyze video, just like a human would.

How They Grade the AI

In many AI tests, a human or another AI reads the answer and says, "That sounds good." ClawMark does something stricter: it uses a robot referee.

  • There are no "opinions" here. The test uses 1,537 tiny, automatic Python scripts (checkers) that look at the actual state of the computer after the AI is done.
  • Did the AI actually save the file? Did it update the calendar? Did it spot the hidden photo?
  • If the AI says "I did it" but the file isn't there, the robot referee gives a zero. It's like a math test where you get points only if the answer is written in the right box, not just if you said the right number out loud.

The Results: Good at Starting, Bad at Adapting

The researchers tested seven top-tier AI models (like the brains behind major chatbots) in this "living office." Here is what they found:

  1. The "Perfect Score" is Rare: Even the best AI only got a "perfect" end-to-end success on 20% of the tasks. This means that while they often made some progress, they rarely finished the whole job without a single mistake.
  2. The "Day 2 Drop": This is the most interesting finding. On the first day, the AIs did pretty well. But on the second day, when the environment changed (new emails, new files), six out of seven models got significantly worse.
    • Analogy: Imagine you are playing a video game. You beat Level 1 easily. But when the game suddenly changes the rules and adds new enemies in Level 2, the AI forgets how to play and starts stumbling. It struggles to "wake up" and realize the world has changed.
  3. Silent Changes are the Kryptonite: The AI failed most often when it missed a "silent mutation"—a change that happened without a loud announcement (like a file quietly updating in the background). They also struggled to actually save their work to the right place (backend writeback).

The Verdict

ClawMark shows that while AI assistants are getting smarter at solving single problems, they are still learning how to be persistent coworkers. They are good at the "first day" but often lose their footing when the office environment shifts around them.

The paper concludes that to build a truly reliable AI coworker, we need to focus less on how well it answers a single question and more on how well it adapts to a changing world and remembers to update its tools when the world changes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →