← Latest papers
🤖 AI

Situation Graph Prediction for User Perspective Modeling

This paper introduces Situation Graph Prediction (SGP), a novel task and synthetic data generation strategy for modeling user perspectives from multimodal artifacts, demonstrating through a diagnostic study that inferring latent internal states is significantly more challenging than surface-level extraction for current foundation models.

Original authors: Jisung Shin, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama

Published 2026-08-12
📖 7 min read🧠 Deep dive

Original authors: Jisung Shin, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

=== SUMMARY ===
Imagine you are trying to understand a friend who is going through a tough week. You see them typing short, angry texts, hearing a shaky voice on a call, and noticing they skipped their usual gym routine. A simple computer program might just see "fewer messages" and "lower activity" and think, "Oh, they are just busy." But a truly smart friend knows that those surface clues actually hide a deeper story: maybe they are feeling overwhelmed, anxious, or spiraling toward a crisis. This is the difference between just watching what someone does and understanding who they are in that moment. Scientists call this "Perspective-Aware AI." It's the dream of building personal assistants that don't just follow orders but actually grasp your evolving goals, feelings, and the unique way you see the world. The big problem? We can't easily teach computers to do this because real people's private data is off-limits, and we rarely have a "reference guide" that labels exactly what someone is feeling inside just by looking at their digital footprints.

This paper introduces a clever new way to teach computers this skill, called Situation Graph Prediction (SGP). Think of it as a detective game where the computer has to look at a messy pile of clues (like emails, chat logs, or voice notes) and reconstruct a hidden, structured map of what the person is actually thinking and feeling. Since we can't use real people's private secrets to train the AI, the authors built a "fake" world. They started by drawing the hidden map first (deciding a person is "stressed" and "at work"), and then used a super-smart AI to write the fake clues (like a curt email) that would naturally result from that map. This "structure-first" method ensures the clues always match the hidden feelings perfectly, creating a safe, privacy-friendly training ground.

When the researchers tested this on three of the world's most advanced AI brains (GPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4), they found something fascinating. Even with these super-intelligent models, it is much harder for the AI to guess the hidden feelings (like "anxiety" or "determination") than it is to just list the visible facts (like "office" or "email"). The study suggests that while AI is getting better at reading the surface of our lives, figuring out the deep, internal story behind our actions is still a really tough challenge. The researchers created a small dataset of 75 fake scenarios to prove this point, showing that the gap between "what we see" and "what we feel" is real and consistent, no matter which AI model you use.

The Detective's Puzzle: From Clues to Inner Worlds

Let's dive deeper into how this works. Imagine you are a detective trying to solve a mystery, but you can't talk to the suspect. You only have a pile of evidence: a text message that says "I'm fine," a photo of a messy desk, and a voice note that sounds a bit shaky. A basic computer might just read the text and say, "The person is fine." But a Perspective-Aware AI wants to build a Situation Graph.

Think of a Situation Graph as a 3D mind-map of a specific moment in someone's life. It's not just a list of facts; it's a web of connections.

  • The Surface Nodes: These are the easy-to-see things, like "Office," "Tuesday," or "Email."
  • The Latent Nodes: These are the invisible, internal states, like "Feeling Stressed," "Valued," or "Confused."

The goal of Situation Graph Prediction (SGP) is to take the messy pile of clues (the "artifacts") and work backward to build that complete mind-map. The paper frames this as an "inverse inference problem." Usually, we know the cause (the person is stressed) and we see the effect (they send a short email). SGP asks the AI to do the reverse: "I see a short email; what must the person be feeling inside?"

The "Fake World" Solution

Here is the tricky part: In the real world, we can't ask people, "Hey, please draw a map of your feelings right now." That's too private and too annoying. So, how do you train an AI to do this without real data?

The authors came up with a brilliant "structure-first" strategy. Instead of asking an AI to imagine a person and then guess their feelings (which is messy and unreliable), they flipped the script:

  1. Draw the Map First: They used a computer to generate a perfect, valid "Situation Graph" (e.g., "Person is at a Job Interview," "Feeling Nervous," "Goal is to Get Hired").
  2. Generate the Clues: Then, they asked an AI to write the "evidence" that would naturally come from that map. So, if the map says "Nervous," the AI writes a shaky voice note or a very polite, overly formal email.

This creates a perfect training set where the "answer key" (the map) is guaranteed to match the "clues" (the text/audio). It's like building a fake crime scene where the detective knows exactly who the culprit is, so they can practice solving the case without ever hurting a real person.

The Pilot Study: Testing the AI Detectives

To see if this actually works, the researchers built a small "pilot" dataset. They created 75 different scenarios involving a fictional character named Elise Navarro, a 28-year-old marketing analyst living in Toronto. They tracked her life over 75 different moments between 2021 and 2025, covering everything from professional struggles to health milestones.

They then handed these 75 cases to three of the most powerful AI models available today: GPT-4o, Gemini 2.5 Flash, and Claude Sonnet 4. They asked the AIs to look at Elise's clues and reconstruct her Situation Graph.

The results were a mix of "not bad" and "still a long way to go."

  • The Good News: When the AIs were given a few examples to learn from (a technique called Retrieval-Augmented In-Context Learning), they got much better at the job. Their ability to get the details right jumped significantly.
  • The Big Discovery: Even the smartest AIs found it much easier to spot the surface facts (like "Elise is at work") than to guess the hidden feelings (like "Elise is feeling overwhelmed").

The researchers measured this using a "gap" score. They found that for all three models, the ability to guess the hidden feelings was consistently worse than spotting the visible facts. For example, with Claude Sonnet 4, the "gap" was about 0.70, meaning the AI was significantly better at reading the surface than the soul. This suggests that while AI is getting great at reading our digital footprints, it still struggles to understand the why and the how we feel behind them.

Why This Matters (And What It Isn't)

This paper doesn't claim to have solved the problem of perfect AI empathy. In fact, it suggests the opposite: that this is a really hard problem that current technology is only just beginning to tackle. The study shows that even with the best models, there is a consistent struggle to move from "what happened" to "what it means."

The authors are careful to point out that their data is synthetic (fake) and small. They aren't saying, "We have a perfect system." They are saying, "We have a new way to test this, and the tests show that guessing human feelings is still much harder than just reading a text message."

By creating this "Situation Graph" framework and the synthetic data to test it, the researchers have given the AI community a new tool. It's like giving detectives a new set of training dummies that perfectly mimic human behavior, allowing them to practice their skills without violating anyone's privacy. The hope is that by understanding exactly where the AI fails (the gap between surface and latent), we can build better, more trustworthy personal agents that truly understand us, not just what we type.

So, the next time you talk to a personal AI, remember: it might know you're at the office and that you sent an email, but figuring out that you're actually having a bad day? That's the next big mountain for AI to climb.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →