Quantifying the Sources of Instability in LLM-Based Stance Analysis of Public Discourse
This study demonstrates that while aggregate sentiment proportions in LLM-based stance analysis of public discourse remain stable across different preprocessing pipelines, the specific conclusions regarding the coupling between affective valence and epistemic modality are highly sensitive to both pipeline variations for low-sample speakers and, more significantly, to systematic disagreements between LLM and keyword-lexicon measurement methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to figure out what a famous person is really feeling and how sure they are about their own words. You have a giant pile of video interviews. But before you can start your investigation, you have to run the raw video through a machine that does three things: it listens to the audio, figures out who is speaking (the guest or the interviewer), and chops the speech into sentences.
This paper asks a simple but tricky question: Does the way you set up your machine change the final verdict?
The researchers, led by Bo Chen, set up a massive experiment with 256 YouTube interviews featuring 41 famous people from finance, politics, and media. They ran these videos through two different "machines" (preprocessing pipelines) and then used two different "detectives" (measurement methods) to analyze the text.
The Two Machines: The "Cut-and-Paste" vs. The "Audio Ear"
Think of the first machine as a Text-Only Detective. It grabs the auto-generated subtitles from YouTube, cleans them up, and tries to guess who is talking based on the words. It's fast, but it sometimes gets confused by the messy way YouTube captions are written.
The second machine is the Audio-Only Detective. It listens to the actual sound waves, identifies the voices, and writes down what was said. This is a more expensive, high-tech approach.
The researchers wanted to see if switching from the Text-Only Detective to the Audio-Only Detective would change the story they told about the speakers.
The Two Detectives: The "AI Brain" vs. The "Word List"
Once the text was ready, they needed to measure two things:
- Valence: Is the speaker happy (positive) or sad/angry (negative)?
- Modality: Is the speaker being super confident (emphatic) or unsure (hedging)?
To measure this, they used two different tools:
- The AI Brain (LLM): A powerful artificial intelligence that reads the text and writes a report on the speaker's feelings and confidence.
- The Word List (Keyword Lexicon): A rigid rulebook that just counts how many times specific "angry" or "confident" words appear.
The Big Surprise: The Machine Matters Less Than You Think (For Some)
Here is the first twist. The researchers found that how you cut the video matters a lot, but only if you don't have enough video to begin with.
If a famous person only has a tiny handful of videos (5 or fewer), the results are a total mess. Switching from the Text-Only Detective to the Audio-Only Detective can completely flip the results, turning a "confident" speaker into a "hesitant" one. It's like trying to guess the weather by looking at a single cloud; it's just too much noise.
However, if a speaker has a lot of videos (16 or more), the machine choice barely matters. For the four best-sampled stars in the study, changing the machine only changed the results by a tiny amount (an average shift of 0.13). The story remained the same regardless of which machine you used.
The Real Villain: The Detective Choice
This is where the plot thickens. While the machine (preprocessing) didn't change the story much for the well-sampled stars, the detective (measurement method) absolutely did.
The researchers found that the AI Brain and the Word List often told completely opposite stories about the same person, even when they were reading the exact same text.
- In nearly 49% of the cases, the two detectives disagreed on whether the speaker was being confident or hesitant.
- For some very famous, well-sampled speakers (like the economist Ken Rogoff with 17 videos), the AI said one thing, and the Word List said the exact opposite.
This suggests that the biggest source of instability isn't how you clean the data, but which tool you choose to analyze it. One tool might see a speaker as "angry and certain," while the other sees them as "calm and unsure."
The Illusion of Safety
You might think, "Well, maybe if we just look at the average of everyone, it all balances out?" The researchers checked this, and the answer is a hard no.
They found that if you just look at the total percentage of negative words across all videos, the numbers look super stable. No matter which machine or which detective you use, the total count of "bad words" barely moves (less than 6 percentage points of change).
But this stability is a trap. It hides the fact that for individual speakers, the conclusions are totally different. It's like saying a classroom is "average" because the math genius and the student who failed the test balance each other out. You miss the fact that one student is struggling and the other is thriving.
The Takeaway
The paper concludes with a warning for anyone studying public discourse: Don't just trust one setup.
- Check your sample size: If you only have a few videos, your results are likely unreliable no matter what you do.
- Don't rely on just one detective: If your AI tool and your word-counting tool disagree, you haven't found the truth yet. You need to see if they agree before you publish your findings.
- Separate the noise: You have to figure out if your results are changing because you changed the machine (preprocessing) or because you changed the tool (measurement).
In short, the way you prepare your data matters, but the tool you use to measure it matters even more. And if you only look at the big averages, you might miss the whole story.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.