← Latest papers
🤖 machine learning

ASD-Bench: A Four-Axis Comprehensive Benchmark of AI Models for Autism Spectrum Disorder

This paper introduces ASD-Bench, a comprehensive four-axis benchmark evaluating diverse machine learning and foundation models on a curated dataset of 4,068 AQ-10 records across three age cohorts to reveal distinct age-specific diagnostic patterns, the limitations of single-metric evaluations, and the need for cohort-tailored deployment strategies in automated Autism Spectrum Disorder screening.

Original authors: Shubhankit Singh, Hassan Shaikh, Kuldeep Raghuwanshi, Keshav Bulia

Published 2026-05-13
📖 5 min read🧠 Deep dive

Original authors: Shubhankit Singh, Hassan Shaikh, Kuldeep Raghuwanshi, Keshav Bulia

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build a "metal detector" to find a specific type of hidden treasure (Autism Spectrum Disorder, or ASD) in a crowd of people. For a long time, researchers have built these detectors, but they've mostly tested them in three ways that don't tell the whole story:

  1. They only tested one type of detector at a time.
  2. They only checked if it found the treasure, not if it was sure it found it, or if it would break if the ground shook.
  3. They mostly tested it on adults, ignoring children and teenagers who might hide the treasure differently.

This paper, ASD-Bench, is like a massive, organized "Tough Guy Tournament" for 17 different AI detectors. The researchers wanted to see which one is the best at spotting ASD using a simple 10-question checklist (called AQ-10), but they did it in a much smarter way.

Here is the breakdown of their tournament, explained simply:

1. The Three Teams (Age Groups)

The researchers didn't just test everyone together. They split the crowd into three distinct teams, realizing that a child, a teenager, and an adult might "hide" their traits differently:

  • The Kids (1–11 years): The easiest group to spot.
  • The Teens (12–16 years): The hardest group. They are like "chameleons" who have learned to blend in, making the treasure harder to find.
  • The Adults (17–64 years): Surprisingly easy to spot, but only because the dataset was small and the patterns were very clear.

2. The Four Rules of the Game (The "Four-Axis" Benchmark)

Instead of just counting how many treasures the detectors found (Accuracy), the judges used four different scorecards:

  • Score 1: The Hit Rate (Performance): Did it find the treasure? (Standard accuracy).
  • Score 2: The Confidence Check (Calibration): If the detector says "90% sure," is it actually 90% sure? Some detectors were great at finding treasure but terrible at knowing how sure they were (like a person who guesses right but thinks they are a genius).
  • Score 3: The "Why" Factor (Interpretability): Can the detector explain which question on the checklist made it decide? (e.g., "I said 'Yes' because the kid hates social chit-chat").
  • Score 4: The Shake Test (Robustness): If you accidentally scribble over a few answers or the data gets a little noisy, does the detector still work, or does it panic and give a wrong answer?

3. The New Scorecard: "The HAP"

The researchers invented a new scoring system called HAP (Heuristic Aggregate Penalty).

  • The Analogy: Imagine a doctor's waiting room.
    • False Positive: The detector says a healthy kid has ASD. The kid gets a second check-up. It's annoying, but not a disaster.
    • False Negative: The detector says a kid with ASD is healthy. The kid misses out on help. This is a tragedy.
  • The HAP Rule: The HAP score treats missing a real case (False Negative) as 5 times worse than a false alarm. It also punishes detectors that are "moody" (inconsistent). If a detector is great on Monday but terrible on Tuesday, HAP gives it a bad score.

4. The Tournament Results

Here is what happened when they ran the 17 different AI models against the three age groups:

  • The Adults: It was a blowout. 10 out of 17 models got a perfect score. However, the researchers warn this might be because the adult group was small and the data was too easy, like a test with all the answers on the back of the page.
  • The Kids: The models did very well. The most important clue for kids was A9: "Do I enjoy social chit-chat?" Kids with ASD usually hate it, and the AI noticed this immediately.
  • The Teens: This was the boss battle. The models struggled the most here. The "social chit-chat" clue stopped working. Instead, the models started looking at A5: "Do I notice patterns in things all the time?" This suggests that as kids grow up, they learn to hide their social struggles (masking), but their obsession with patterns remains visible.
  • The "Confidence" Trap: One model, AdaBoost, got a perfect score for finding adults, but it was wildly overconfident and wrong about how sure it was. The paper says: "Don't trust a model that is confident but wrong."

5. The Winners (Who should you use?)

The paper doesn't say "Use this one forever." Instead, it says "Use the right tool for the right job":

  • For Adults (in a perfect, quiet room): A complex model called TabTransformer works best.
  • For Adults (in a noisy, messy room): A simpler model called MLP is more stable.
  • For Kids: A model called XGBoost is the champion.
  • For Teens: A "Foundation Model" called TabPFN is the best at finding the treasure, even though it's a bit fragile. If you need it to be tough against noise, use TabNet instead.

6. The Big Warning (The "Fine Print")

The authors are very careful to say: "This is a proof-of-concept, not a medical diagnosis tool."

  • The Data: They used a dataset where people filled out a questionnaire on an app. It wasn't a doctor confirming the diagnosis.
  • The Limit: Because the data came from one specific app and one specific time, we don't know if these models would work in a real hospital in a different country.
  • The Takeaway: This paper is a map showing us how to build better detectors and what to measure, but the detectors themselves aren't ready to replace a doctor yet.

In a nutshell: The researchers built a rigorous gym for AI models to test their strength, balance, and honesty. They found that children, teens, and adults are different puzzles, and the best AI for one group might be the worst for another. They also proved that being "accurate" isn't enough; a medical AI must also be honest about its confidence and tough enough to handle messy real-world data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →