← Latest papers
💬 NLP

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings

This paper introduces LEX-EC, a black-box audit framework that combines prevalence diagnostics, lexical ablation, and prompt sensitivity analysis to distinguish between marginal distribution effects and genuine trait-associated signals in zero-shot large language model personality classification across diverse text genres.

Original authors: Brittany Harbison, Ashok K. Goel

Published 2026-07-28
📖 3 min read☕ Coffee break read

Original authors: Brittany Harbison, Ashok K. Goel

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to figure out someone's personality just by reading a single note they wrote. You might guess they are "outgoing" because they used exclamation points, or "serious" because they listed their job title. This is the world of Large Language Models (LLMs): super-smart computer programs that read text and try to guess human traits like whether someone is friendly, anxious, or creative. Scientists call these guesses "Big Five" personality labels. But here is the tricky part: just because a computer says, "This person is definitely outgoing," doesn't mean it actually knows the person. It might just be guessing based on the length of the note, the topic (like talking about parties), or even the writer's age and gender. The big question for researchers is: Are these AI detectives actually reading the person, or are they just spotting easy clues and making lucky guesses?

This paper introduces a new detective tool called LEX-EC to solve this mystery. Think of LEX-EC as a "lexical evidence-channel audit." In plain English, it's a way to test if an AI is really understanding personality or just cheating. The researchers took three different types of writing—long, free-flowing essays, short graduate student introductions, and tiny Facebook status updates—and asked various AI models to guess the writers' personalities. Then, they played a game of "hide and seek" with the words. They used a digital eraser to wipe out all the specific topics (like "I love hiking" or "I live in Atlanta") and demographic details (like "I am 22 years old"), leaving behind only the "skeleton" of the language: the small connecting words, emotional words, and style markers. They then asked the AI to guess the personality again using only this stripped-down text.

The results were a mix of surprises and confirmations. When the AI looked at long essays, it could still find a faint signal of personality even after the specific topics were erased, suggesting it was picking up on some deeper style habits. However, when the AI looked at short Facebook statuses (which averaged only about 17 words), the personality clues vanished almost completely once the topics were removed. This suggests that for very short texts, the AI might have been relying entirely on the specific things people talked about, rather than who they actually are.

The study also tested how the AI "explains" its guesses. When the researchers told the AI to focus on how the person wrote (their style and feelings) rather than what they wrote, the AI's explanations changed to match the new instructions. However, the AI didn't necessarily become more accurate; it just changed its story. The authors found that the AI's ability to guess personality wasn't a single super-power; it depended heavily on how much text was available and what kind of writing it was. In short, the paper suggests that while AI can spot personality patterns in long, rich texts, it struggles mightily with short snippets, often mistaking the topic for the person. This means we should be very careful about using these tools to judge people based on just a few sentences, as the AI might be seeing ghosts in the machine rather than the real human behind the text.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →