← Latest papers
💬 NLP

Can AI Guess What You Know? Performance Comparison of Large Language Models for Human Domain Knowledge Estimation From Communication Logs

This study evaluates the ability of seven Large Language Models to infer individual domain knowledge from Slack communication logs by comparing their zero-shot estimates against self-reported skills, finding that while Gemini 2.5 Flash achieved the highest accuracy, inference performance is only weakly dependent on message volume and highlights the need for privacy-preserving and structure-aware approaches.

Original authors: Ko Watanabe, Shoya Ishimaru

Published 2026-05-25
📖 4 min read☕ Coffee break read

Original authors: Ko Watanabe, Shoya Ishimaru

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you walk into a massive, bustling office. You have a tricky problem, but you don't know who to ask. Is it the person in the corner who talks about coding all day? Or the one who seems to know everything about project management? Usually, you'd have to wander around, ask a few people, and hope you find the right expert. This "guessing game" wastes time and money.

This paper asks a simple question: Can an AI act as a super-observant intern that reads everyone's chat history and instantly tells you who knows what?

Here is the breakdown of their experiment, explained simply:

The Setup: The "Chat Detective"

The researchers took a giant pile of real chat messages from a company's Slack account (about 27,000 messages from 43 people over several years). They fed these messages into seven different "super-brains" (Large Language Models like GPT, Claude, and Gemini).

Think of these AI models as different types of detectives:

  • Some are fast but maybe a bit shallow (like a quick scanner).
  • Some are slow but deep thinkers (like a professor analyzing a book).
  • Some are generalists, while others are specialists.

The AI's job was to read the chats and create a "skill report" for each person. For example, if someone kept talking about "Python" and "Data Science," the AI would guess, "This person knows Python."

The Reality Check: The "Self-Report"

To see if the AI was actually good at its job, the researchers asked the real humans to grade themselves. They created a simple website where each person looked at the list of skills the AI found and said, "Yes, I know this," or "No, I don't."

Then, the researchers compared the AI's guess against the human's actual answer. It's like a teacher grading a student's guess of the test answers against the real answer key.

The Results: Who Won the Contest?

The researchers tested seven different AI models. Here is how they ranked:

  1. The Winner: Gemini 2.5 Flash was the best guesser. It was off by about 21 points on a 0–100 scale. (If the human said they were 80% expert, the AI guessed around 59% or 101%).
  2. The Runner-Up: Gemini 2.5 Pro was close behind.
  3. The Middle Pack: The Claude models (Haiku and Sonnet) were okay, but not as sharp as the Gemini models.
  4. The Strugglers: The GPT models (including the very famous GPT-4o and GPT-5) performed the worst. They were off by about 33 points on average.

The Big Takeaway: Even the best AI wasn't perfect. It could get the general idea of who knows what, but it couldn't pinpoint the exact level of expertise with high precision.

The Surprising Twist: More Text ≠ Better Guesses

You might think, "If I give the AI a million messages instead of a thousand, it will get smarter, right?"

Not necessarily. The researchers found that simply having more chat history didn't automatically make the AI's guesses more accurate.

  • Why? Because a lot of chat messages are just "Okay," "Got it," or "Lunch at 1?" These don't tell you what someone knows.
  • The Limit: Once the AI had a minimum amount of chat to work with, adding more didn't help much. It's like trying to guess someone's favorite movie by reading their grocery list; no matter how long the list is, it won't tell you much about their taste in cinema.

The Bottom Line

The study proves that AI can peek into our chat logs and make a decent guess about our skills. It's a useful tool for getting a rough map of "who knows what" in a company.

However, it's not a magic crystal ball yet. The AI still makes mistakes, and the famous GPT models didn't even come in first place in this specific test. The researchers also noted that for this to work in real life, companies would need to be very careful about privacy, as sending private company chats to outside AI servers is a sensitive issue.

In short: AI is getting better at reading the room, but it still needs to learn how to distinguish between a person's actual expertise and just their daily chatter.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →