← Latest papers
🧠 neurology

Using MRI whole brain atrophy and clinically reported outcomes in combination to assess interim treatment response in multi-arm multi-stage trials in progressive multiple sclerosis

This paper proposes and evaluates a multivariate mixed model that combines MRI whole brain atrophy with three clinical outcomes to enhance the resilience and statistical power of interim treatment response assessments in multi-arm multi-stage trials for progressive multiple sclerosis.

Original authors: Burnell, M., Nicholas, J., Burton, R., Chataway, J., Apap Mangion, S., Carpenter, J.

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Burnell, M., Nicholas, J., Burton, R., Chataway, J., Apap Mangion, S., Carpenter, J.

Original paper dedicated to the public domain under CC0 1.0 (https://creativecommons.org/publicdomain/zero/1.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The Big Picture: The "Taste Test" Problem

Imagine you are a chef developing a new soup. You have three different experimental recipes, and you want to know which one is the best before you spend a fortune cooking it for a huge banquet.

In a standard cooking contest, you might wait until the very end to taste the final soup (the "primary outcome"). But in medical trials for serious diseases like Progressive Multiple Sclerosis (PMS), waiting until the end takes years and costs millions.

So, scientists use a "Multi-Arm Multi-Stage" (MAMS) trial. Think of this as a cooking competition with two rounds:

  1. Round 1 (Interim): A quick taste test to see if the soup has any promise. If a recipe tastes terrible here, you stop making it immediately to save money and time.
  2. Round 2 (Final): If a recipe passes Round 1, you cook it longer and do a full, detailed tasting to see if it's actually the best.

The Problem: Relying on Just One Ingredient

In the OCTOPUS trial (the study this paper discusses), the scientists needed a way to judge the "Round 1" taste test.

Originally, they decided to judge the soup based on only one ingredient: the rate at which the brain shrinks (called "whole brain atrophy"), measured by an MRI scan.

  • The Logic: They thought, "If the drug works, the brain shouldn't shrink as fast."
  • The Risk: The authors realized this was like judging a complex soup only by how much salt is in it. Sometimes, a great soup might have the perfect amount of salt, but other ingredients (like texture or flavor) might be off. Conversely, a bad soup might accidentally have the right amount of salt.

In real-world trials, they found cases where a drug successfully stopped patients from getting disabled (the "flavor"), but the brain shrinkage (the "salt") didn't change much. If they had relied only on the MRI scan, they might have thrown away a good drug too early.

The Solution: The "All-Hands" Approach

To fix this, the authors proposed a new strategy: Don't just look at the MRI; look at the MRI plus three other clinical tests.

They combined:

  1. The MRI: Brain shrinkage.
  2. The EDSS: A doctor's score of disability.
  3. The Walk Test: How fast a patient can walk 25 feet.
  4. The Peg Test: How fast a patient can put pegs in a board.

The Analogy: Instead of judging the soup by just the salt, you now taste for salt, sweetness, spice, and texture all at once.

The Technical Magic: The "Universal Translator"

Here is the tricky part: You can't just add these things together easily.

  • You can't add "millimeters of brain shrinkage" to "seconds it takes to walk." It's like trying to add "miles" to "apples."

The authors invented a statistical "Universal Translator" (a multivariate mixed model).

  1. Standardization: They converted all the different measurements into a common language (like converting everything to "points of improvement").
  2. Weighting: They looked at how these tests usually move together. If the walk test gets worse, does the peg test usually get worse too? They used this relationship to blend the data.
  3. The Result: They created a single "Super-Score" that represents the overall health of the patient, combining the MRI and the three physical tests into one number.

What Did They Find?

They ran the numbers using data from a previous trial (MS-STAT2) to see if this new "Super-Score" would have worked better than just the MRI.

  • The Power Boost: By adding the three physical tests to the MRI, they increased the chance of spotting a working drug from 83% to 90%.
  • The Safety Net: More importantly, this method makes the trial more resilient. If the drug works great on walking but the brain shrinkage is weirdly slow, the "Super-Score" still catches the success. If the drug fails on walking but helps the brain, it still catches that. It spreads the risk so you don't accidentally reject a good drug just because one specific test didn't show a result.

The Verdict

The authors conclude that while this method doesn't magically turn a bad drug into a good one, it makes the "Round 1" check much smarter.

  • Old Way: "Did the brain shrink less?" (If no, stop the trial).
  • New Way: "Did the brain shrink less, OR did the patient walk better, OR did the peg test improve?" (If any of these show promise, keep the trial going).

This approach ensures that in the high-stakes game of finding a cure for Progressive MS, scientists are less likely to throw away a winning recipe just because they were only looking at the salt.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →