← Latest papers
🤖 AI

ProEvent: An Event-centric Benchmark for Proactive Agents

The paper introduces ProEvent, the first event-centric benchmark for evaluating proactive agents' ability to autonomously track user events from instant messaging chats, revealing that current large language models struggle significantly with timing, event cancellation, and implicit reasoning.

Original authors: Guanzhen Li, Liangming Pan, Leye Wang

Published 2026-07-21
📖 5 min read🧠 Deep dive

Original authors: Guanzhen Li, Liangming Pan, Leye Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart digital assistant that doesn't just wait for you to ask for help, but actually watches your life like a hawk, guessing what you need before you even say a word. This is the dream of "proactive agents." Unlike a regular robot that sits idle until you press a button, a proactive agent is like a personal chef who starts chopping vegetables the moment they smell the garlic, or a friend who hands you an umbrella the second the sky turns gray. The big challenge in computer science right now is teaching these digital brains to do this without getting confused by the messy, chaotic reality of human life. We know computers are great at following strict rules, but real life is full of hidden hints, overlapping conversations, and people changing their minds without saying "cancel." If we want these agents to be truly helpful, they need to be able to track our future plans—like a hiking trip or a meeting—just by listening to our text messages, and they need to know exactly when to speak up and when to stay quiet.

Enter PROEVENT, a new "training ground" designed by researchers to test if these digital assistants are actually ready for the real world. Think of PROEVENT as a giant, high-stakes video game level where the player is an AI trying to manage a human's calendar. Instead of giving the AI a neat list of appointments, the researchers feed it a stream of messy, realistic text chats. The AI has to listen to conversations between friends, colleagues, and family members, figure out what events are being planned, and then update the calendar accordingly. But here's the twist: the chats are full of distractions. There are side conversations about nothing, people changing plans halfway through, and messages that sound like cancellations but aren't. The AI has to be a detective, a scribe, and a time-manager all at once.

The researchers built this test by creating thousands of fake but very realistic chat logs. They programmed "characters" with different personalities and goals, then simulated complex scenarios where people negotiate meeting times, cancel plans because they feel sick, or try to schedule things that never actually happen. They even added "noise"—random messages that have nothing to do with the main event—to see if the AI would get distracted. The goal was simple: can the AI correctly build and update a timetable just by reading these chats?

When they put eight different powerful AI models through this test, the results were a mix of "not bad" and "oh no." The biggest surprise was that the AIs tend to be over-eager. Imagine a dog that barks at every leaf that moves; these AIs were constantly trying to update the calendar even when nothing had changed. In fact, for some models like DeepSeek-V3.2, they tried to make changes more than 96% of the time when they should have just stayed silent. They were so afraid of missing a chance to help that they ended up creating chaos.

Another major hurdle was cancellation. If a user says, "I'm not feeling well, maybe I can't make it," a human knows that's a cancellation. The AIs were actually quite good at spotting that a deletion might be needed (catching about half of the actual cancellations), but they struggled to be sure about it. They frequently deleted events that were still happening or removed the wrong plans entirely. The paper suggests that even the most advanced models, like the one named GPT-5.1, only got the right answer in about 26.7% of the tricky scenarios. That means in roughly three out of four cases, the AI messed up the schedule.

The study also found that these digital brains have a weird blind spot: they don't really understand who they are helping. When a user says, "I can't go to the party," the AI often just removes the user's name from the guest list but leaves the party on the calendar. It's like a waiter who, when you say you're full, just takes your name off the reservation but leaves the table set for you. The AI is looking at the event from the outside, like a third-party observer, instead of stepping into the user's shoes and realizing, "Oh, I (the user) am not going, so this event is over for me."

In short, while these AI agents are getting smarter, they aren't quite ready to be your proactive life manager just yet. They are too jumpy, they get confused by complex conversations, and they struggle to understand the difference between a "maybe" and a "no." The researchers hope that by using PROEVENT to show exactly where these AIs fail, future versions will learn to be more patient, more accurate, and finally, more like a helpful human friend who knows exactly when to hand you that umbrella.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →