AgentIF-OneDay: A Task-level Instruction-Following Benchmark for General AI Agents in Daily Scenarios
This paper introduces AgentIF-OneDay, a novel benchmark comprising 104 diverse daily tasks across open workflow execution, latent instruction inference, and iterative refinement, designed to evaluate general AI agents' ability to follow natural language instructions and produce tangible file-based results, revealing that both API-based and RL-enhanced agents currently lead the field.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "One-Day" Test for AI Assistants
Imagine you hire a new personal assistant. You don't just want to see if they can answer a trivia question or write a poem. You want to see if they can handle your entire day—from organizing your work schedule to planning a family trip, all while following your specific, sometimes messy, instructions.
This paper introduces AgentIF-OneDay, a new "final exam" for AI agents designed to test exactly that. It moves away from testing AI on hard coding puzzles or deep scientific research and instead asks: "Can this AI actually help a regular person get their daily life and work done?"
The Three Types of "Daily Tasks"
The researchers created 104 different scenarios to test the AI. They grouped these tasks into three categories, which are like three different ways you might ask your assistant for help:
1. The "Follow the Recipe" Test (Open Workflow Execution)
- The Analogy: You give your assistant a very specific, step-by-step recipe for a complex dish. You say, "First, check the oven temperature. Then, look at the timer. Then, mix these three ingredients. Do not skip step 3."
- What it tests: Can the AI follow a long, detailed list of instructions without getting confused, forgetting a step, or making things up? It tests if the AI can stick to the plan you gave it, even if the plan is long and complicated.
2. The "Read Between the Lines" Test (Latent Instruction Inference)
- The Analogy: You hand your assistant a folder of old documents and say, "Make a new report that looks exactly like these." You don't write down the rules (like "use blue headers" or "put the date in the bottom right"). You expect the assistant to look at the old documents, figure out the hidden style rules, and apply them to the new report.
- What it tests: Can the AI look at a file (like a PDF or a spreadsheet) and understand the unspoken rules? It's about spotting patterns and implicit constraints that you didn't explicitly say out loud.
3. The "Edit and Improve" Test (Iterative Refinement)
- The Analogy: Your assistant drafts a presentation. You look at it and say, "The font is too small, and the second slide is missing a chart. Also, make the background lighter." You aren't asking for a new presentation from scratch; you want them to fix the existing one while keeping the good parts.
- What it tests: Can the AI remember what it just did, understand your feedback, and make precise changes without breaking the rest of the work? It tests how well the AI "remembers" the state of the project.
How They Graded the AI
Instead of just asking "Did it work?", the researchers used a very strict grading rubric (like a teacher's answer key).
- The "Judge": They used a super-smart AI (Gemini-3-Pro) to grade the results, but they checked it against human judges to make sure it was fair. They found the AI judge agreed with humans about 80% of the time.
- The Files: The test wasn't just about text. The AI had to handle real files—PDFs, Excel sheets, images, and code—just like a human would.
What They Found: The Results
They tested four of the top AI agents available today. Here is the takeaway:
The "API" vs. "Specialized" Debate:
- Some companies build AI agents by just plugging a powerful chatbot into a set of tools (like a generalist).
- Others build agents with special "brain training" (Reinforcement Learning) specifically for doing tasks.
- The Surprise: The paper found that both types are performing almost equally well. The "special training" isn't giving a huge advantage anymore. It seems that the basic "intelligence" to handle tasks is now built into the main AI models themselves.
The Winners:
- Manus came out on top overall. It was particularly good at following long, complex workflows.
- Genspark was excellent at following instructions and handling negative constraints (like "don't do X").
- ChatGPT-Agent was the best at work-related tasks.
- Minimax-Agent was the slowest, taking a long time to think, though it was good at logic.
The Weakness:
- The hardest part for all the AIs was Latent Instruction Inference (reading between the lines). Even the best agents struggled to perfectly figure out hidden rules from a file without being explicitly told what to look for.
The Bottom Line
This paper argues that we need to stop testing AI only on how "smart" it is at solving riddles. Instead, we need to test how useful it is in real life.
The good news is that AI agents are getting very good at handling the "boring" but essential parts of our day—managing files, following complex workflows, and editing work. The bad news is that they still sometimes struggle to "read the room" or figure out hidden rules in documents without a little more help from us.
The future of AI isn't just about making the model smarter; it's about building better products that help us navigate our specific, messy, daily lives.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.