← Latest papers
📊 statistics

Human- vs. AI-generated tests: dimensionality and information accuracy in latent trait evaluation

This study compares AI-generated and human-developed versions of the Body Awareness Questionnaire, finding that while surface wording is similar, AI adaptations exhibit significant differences in dimensionality and information accuracy, underscoring the need for rigorous statistical validation of AI-driven measurement tools.

Original authors: Mario Angelelli, Morena Oliva, Serena Arima, Enrico Ciavolino

Published 2026-02-16
📖 6 min read🧠 Deep dive

Original authors: Mario Angelelli, Morena Oliva, Serena Arima, Enrico Ciavolino

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Can an AI Write a Better (or Worse) Test?

Imagine you are a teacher who needs to create a quiz to measure how "aware" students are of their own bodies (like noticing if their heart is racing or if they are tense). You have a perfect, time-tested quiz created by human experts.

Now, you ask a super-smart AI (like ChatGPT) to write a new quiz that looks just like the human one. The AI does its job, and the new quiz looks almost identical on the surface. The words are similar, the questions are about the same topics, and the students seem to answer them just fine.

The paper asks: Is the AI's quiz actually the same as the human one underneath the surface?

The researchers (Mario, Morena, Serena, and Enrico) decided to put both quizzes under a microscope. They didn't just look at the words; they looked at the math behind the answers to see if the AI was secretly measuring something different or doing a sloppy job.


The Analogy: The "Shadow Puppet" vs. The "Real Hand"

Think of the Human-Generated Test as a real hand casting a shadow on a wall. The shape of the shadow perfectly matches the hand because the hand is solid and consistent.

Think of the AI-Generated Test as a puppet made of sticks and cloth. From a distance (or if you just look at the words), the puppet looks exactly like the hand. But if you poke it, the joints move differently. The "bones" (the internal logic) aren't the same.

The researchers found that while the AI's puppet looked like the human hand, its internal structure was wobbly and inconsistent.

What They Found (The Three Big Differences)

1. The "One-Track Mind" vs. The "Confused Mind" (Dimensionality)

  • The Concept: A good psychological test should measure one main thing (like "Body Awareness").
  • The Finding: The human test was like a laser beam, focusing on one clear target. The AI test, however, was like a flashlight with a cracked lens. Sometimes it measured body awareness, but sometimes it accidentally measured something else (like general anxiety or confusion).
  • The Metaphor: If the human test is a straight arrow hitting the bullseye, the AI test is an arrow that wobbles in the air and hits a slightly different spot, even though it was aimed at the same target.

2. The "Difficulty Ladder" (Item Ordering)

  • The Concept: In a test, some questions are easy, and some are hard. In a good test, the order of difficulty should be consistent. If Question A is harder than Question B for humans, it should be harder for the AI version too.
  • The Finding: The researchers found the AI scrambled the ladder.
    • Human Test: "Question 1 is easy, Question 18 is hard."
    • AI Test: "Question 1 is hard, Question 18 is easy."
  • The Metaphor: Imagine a video game level. The human designers made Level 1 easy and Level 10 hard. The AI designer looked at the list and accidentally made Level 1 a boss fight and Level 10 a tutorial. The names of the levels were the same, but the experience was completely flipped.

3. The "Flashlight Beam" (Information Accuracy)

  • The Concept: A test is most useful when it gives you the most accurate information about a person. Some tests are great at spotting people who are very aware, but bad at spotting people who are not aware.
  • The Finding: The human test was like a spotlight. It was very sharp and precise in the middle range (where most people fall). It told the researchers exactly where a person stood.
    The AI test was like a floodlight. It spread its light everywhere. It wasn't wrong, but it was "blurry." It gave information over a wide range, but it wasn't as sharp or precise as the human version.
  • The Metaphor: If you are trying to measure the temperature of a cup of coffee, the human test is a precise thermometer that tells you it's exactly 60°C. The AI test is a guess that says "It's somewhere between 40°C and 80°C." Both are technically "right," but the human one is much more useful.

The "Prompt" Problem

The researchers tried two different ways of asking the AI to write the test:

  1. The Strict Prompt: "Here is the exact list of questions. Rewrite them but keep the meaning." (This made a better AI test, but it still had the "wobbly" internal structure).
  2. The Vague Prompt: "Write a test about body awareness with 18 questions." (This made a much worse test that was confused and inconsistent).

The Lesson: Even if you tell the AI exactly what to do, it doesn't "understand" the concept the way a human expert does. It's just predicting the next word based on patterns it saw in its training data.

Why Should You Care?

This paper is a warning label for the future of research and psychology.

  • Don't trust the surface: Just because an AI-generated survey looks like a real one doesn't mean it measures the same thing.
  • The "Black Box" risk: We can't always see why the AI made a question hard or easy. It might be picking up on hidden biases or stereotypes in its training data.
  • Human Supervision is Key: We can use AI to help write tests, but we must have human experts check the "math" (the internal logic) to make sure the AI isn't hallucinating a fake structure.

The Bottom Line

AI is a powerful tool, like a very fast, very creative intern. But if you ask that intern to build a bridge (a psychological test), you can't just look at the paint job. You have to check the steel beams underneath. In this study, the AI's beams were a bit shaky, even if the paint job looked perfect.

In short: AI can mimic the words of a human test, but it struggles to mimic the soul (the statistical structure) of one. We need to keep humans in the loop to ensure the tests we use are actually telling the truth.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →