← Latest papers
🧬 biology

Ontology-constrained multi-LLM scoring of hypothesis support in the predictive processing literature

This paper introduces an ontology-constrained, multi-LLM pipeline that synthesizes fragmented predictive coding literature by scoring studies against a defined glossary of hypotheses, thereby generating auditable disagreement measurements and quantitative evidence spaces that conventional meta-analysis cannot resolve.

Original authors: Hamed Nejat, Alexander Maier, Jesse Spencer-Smith, André M. Bastos

Published 2026-06-05
📖 5 min read🧠 Deep dive

Original authors: Hamed Nejat, Alexander Maier, Jesse Spencer-Smith, André M. Bastos

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are trying to understand a massive, chaotic library where thousands of scientists have written books about how the human brain predicts the future. The problem is, these books are written in different languages, use different maps, and sometimes even contradict each other. Some authors talk about "prediction errors" as if they are electrical sparks, while others describe them as statistical probabilities. Trying to read all of them and figure out what the scientific community actually agrees on is like trying to assemble a giant puzzle where the pieces are from different boxes.

This paper describes a new way to solve that puzzle using a team of AI "librarians."

The Problem: A Fragmented Library

The authors focus on a specific field called Predictive Coding. This is the idea that our brains are constantly guessing what will happen next and only paying attention when reality surprises us (like hearing a loud noise in a quiet room).

Over the years, scientists have studied this using many different tools: brain scans, electrical recordings, computer models, and behavioral tests. Because the methods are so different, it's hard to combine their findings into one clear picture. Traditional ways of summarizing science (called meta-analysis) usually require all studies to be done the exact same way, which isn't possible here.

The Solution: A Council of AI Librarians

Instead of one human expert trying to read everything, the authors built a local multi-LLM pipeline. Think of this as a "council" of ten different AI models (like a team of ten distinct experts) working together in a secure, private room (local execution) so no sensitive data leaves their computers.

Here is how they made the AI team work:

  1. The Rulebook (The Ontology): Before the AI started reading, the human authors created a strict "glossary" or rulebook. It defined 36 specific concepts the AI should look for, grouped into three main ideas (hypotheses):

    • Predictive Suppression: Does the brain quiet down when it expects something?
    • Feedforward Error Propagation: When the brain is surprised, does it send a "shout" of error signals up the hierarchy?
    • Ubiquity: Is this "prediction machine" used everywhere in the brain, or just in specific spots?
  2. The Reading Process: The AI team read 31 scientific papers. They didn't just read the text; they also looked at the figures and charts (using a special vision tool) to understand the data.

  3. The Scoring: For every paper, each of the 10 AI models gave a score from -1 (strong disagreement) to +1 (strong agreement) for each of the three main ideas. They had to cite exactly where in the paper they found their evidence, acting like a strict peer reviewer.

The Results: Mapping the Landscape

Once the AI team finished scoring, the authors treated the results like a 3D map.

  • The "Local" vs. "Global" Oddball: They tested the papers in two different scenarios:
    • Local Oddball: A simple surprise (like a beep in a series of beeps).
    • Global Oddball: A complex, long-term pattern surprise (like a rule change in a game).
  • The Findings:
    • Agreement: The AI council found that most papers agreed that the brain does "predictive suppression" (it quiets down for expected things) and sends error signals forward.
    • Disagreement: There was much less agreement on whether these mechanisms happen everywhere in the brain (Ubiquity).
    • The Twist: The agreement was much stronger for simple "Local" surprises than for complex "Global" ones. The "Global" context was messier and more scattered.

New Tools for Measuring Science

The authors invented some clever ways to measure how "together" the scientific literature is:

  • Hypothesis-Space Temperature: Imagine the scientific papers are gas particles in a jar. If they are all huddled together in one corner, the "temperature" is low (high agreement). If they are bouncing wildly all over the jar, the "temperature" is high (high disagreement).
    • They found the "temperature" was low for simple local surprises (scientists agree).
    • The "temperature" was high for complex global surprises (scientists disagree).
  • The Outlier: They identified one specific paper (Westerberg et al., 2025) that was very different from the rest. It disagreed with the standard view, especially regarding complex global patterns. The AI council successfully flagged this paper as an "outlier" without human bias.

Why This Matters

The paper demonstrates that a team of AI models, guided by a strict human-made rulebook, can read a messy, fragmented scientific field and turn it into a clean, quantitative map.

  • Auditability: Because the AI models are running locally and keeping logs of their reasoning, humans can check their work.
  • Measuring Disagreement: Instead of just saying "scientists disagree," this method measures how much they disagree and where the disagreement happens.
  • Scalability: This approach can handle more papers than a human team ever could, allowing science to keep up with the rapid growth of new research.

In short, the authors built a machine that can read a chaotic library, sort the books into a neat 3D map, and tell us exactly where the scientists are in sync and where they are still arguing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →