← Latest papers
🤖 AI

A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video

This paper proposes a deterministic evidence layer that stabilizes vision-language model outputs for autism screening by converting naturalistic home videos into timestamped, labeled event tables scored via a transparent weight-of-evidence mechanism, achieving robust and interpretable diagnostic performance on preschool children.

Original authors: Wenqi Li, Mindi Ruan, Chuanbo Hu, Shuo Wang, Xin Li

Published 2026-10-08
📖 5 min read🧠 Deep dive

Original authors: Wenqi Li, Mindi Ruan, Chuanbo Hu, Shuo Wang, Xin Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Diagnosing autism spectrum disorder often begins with a specialist watching a child play, looking for subtle signs in how they interact with others or repeat certain movements. This process is highly effective but relies on a scarce resource: trained experts. Because there are not enough specialists to see every child who might need help, families often face long waits for an evaluation, and early intervention is delayed. In recent years, scientists have looked to home videos—recorded by parents on their phones during ordinary play—as a way to bridge this gap. The hope is that artificial intelligence could watch these videos and spot the same patterns a human expert would, offering a quick, accessible first step to decide which children need urgent attention. However, a major hurdle has emerged: even the most advanced AI models, when asked to make the same judgment twice, can give different answers. This inconsistency makes them unreliable for medical screening, where a decision must be the same every time to be trusted.

A team of researchers has tackled this problem by building a new kind of screening system that separates what the computer sees from how it decides. Instead of asking a single artificial intelligence to watch a video and immediately say "autism" or "no autism," they created a two-step process that acts more like a careful human observer followed by a strict accountant. First, a powerful vision model watches the video and writes down a detailed list of events, noting exactly what the child did, when it happened, and what the parent or caregiver did just before. Crucially, this model is forced to stick only to what is visible in the video, avoiding guesses about the child's inner thoughts. It also has to list examples of normal behavior, not just the unusual ones, to provide a complete picture. This list is then passed to a second, text-only model that adjusts the importance of each event based on the child's age, understanding that some behaviors are normal for a toddler but concerning for a preschooler. Finally, a deterministic computer program, which follows fixed mathematical rules without any guesswork, adds up the evidence from this list to produce a final risk score.

The researchers tested this system on 43 home videos of preschool children, recorded by their families without any special instructions or scripts. The videos included children who had been diagnosed with autism by specialists and children who were developing typically. When the system watched these videos, it achieved a high level of accuracy, correctly identifying the risk level in about 86 percent of the cases. More importantly, the system was remarkably stable. When the researchers ran the same videos through the system three times, the final decision remained the same for nearly three-quarters of the clips every single time. In contrast, a standard artificial intelligence model without this special structure changed its mind on nearly a third of the clips between runs. The new system was particularly good at avoiding false alarms; it never flagged a typically developing child as high-risk in every single run, whereas a standard model flagged nine out of thirty-one such children in every run.

The key to this success was not making the artificial intelligence smarter, but rather changing how it was used. The researchers found that the instability in standard models comes from the way they generate answers, which can vary slightly even when the settings are identical. By moving the final decision out of the AI and into a fixed scoring system that simply adds up the evidence the AI collected, they removed that source of error. The system also introduced strict rules for the AI, requiring it to name the specific moment a child failed to respond to a parent's call or a gesture, rather than just guessing that the child was unresponsive. This "grounded" approach meant the AI could only report what it actually saw, and the final score was built on a transparent record of those observations.

While the results are promising, the researchers are careful to note the limits of their work. The system was tested on a relatively small group of children from one location, and it missed three children with autism in every single run, suggesting that the current technology still struggles to detect certain subtle signs. The system works best as a triage tool to help prioritize children for specialist evaluation, not as a final diagnosis. The study demonstrates that by combining the pattern-recognition power of modern AI with rigid, transparent rules for decision-making, it is possible to create a screening tool that is both accurate and reliable. This approach offers a path forward for using home videos to reduce the waiting time for autism evaluations, provided that future work can expand the system to include more diverse children and improve its ability to catch every case.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →