← Latest papers
💬 NLP

LADDER: Language-Driven Slice Discovery and Error Rectification in Vision Classifiers

LADDER is a novel framework that leverages large language models to discover coherent, language-driven error slices and generate pseudo-attributes for comprehensive bias mitigation in vision classifiers, overcoming the limitations of traditional clustering and attribute-based methods by integrating domain knowledge and advanced reasoning without requiring explicit annotations.

Original authors: Shantanu Ghosh, Rayan Syed, Chenyu Wang, Vaibhav Choudhary, Binxu Li, Clare B. Poynton, Shyam Visweswaran, Kayhan Batmanghelich

Published 2026-08-07
📖 4 min read☕ Coffee break read

Original authors: Shantanu Ghosh, Rayan Syed, Chenyu Wang, Vaibhav Choudhary, Binxu Li, Clare B. Poynton, Shyam Visweswaran, Kayhan Batmanghelich

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to recognize animals. You show it thousands of pictures of cats and dogs, and it gets really good at guessing. But then, you notice something weird: the robot isn't actually learning what a cat is. Instead, it is relying on shortcuts. It's looking at the background. If the cat is on a green lawn, the robot thinks, "Green lawn? Must be a cat!" But if you show it a cat on a red carpet, the robot panics and guesses "dog." This is called a "bias," and it's a sneaky problem in artificial intelligence. The robot is relying on shortcuts instead of doing the hard work of understanding.

For a long time, scientists tried to fix this by making a giant checklist of things to look for, like "is there grass?" or "is there a tree?" But this is like trying to catch a thief by only checking for red hats; what if the thief is wearing a blue one? The old methods were too rigid. They couldn't think outside the box, and they often missed the subtle clues that made the robot fail. This is where a new idea comes in: what if we could just ask the robot, "Hey, why did you get that wrong?" and have it explain itself using words? That's the big question this paper tackles.

Enter LADDER, a clever new tool that acts like a detective for AI mistakes. Instead of forcing the robot to fit into a pre-made checklist, LADDER uses a super-smart language brain (called a Large Language Model, or LLM) to figure out why the robot is failing. Think of it this way: if the robot is a student taking a test, the old methods were like a teacher who only checked if the student wrote the right answer. LADDER is like a teacher who reads the student's scratch paper, sees they were distracted by a bird outside the window, and says, "Ah! You didn't fail because you don't know math; you failed because you were looking at the bird!"

Here is how LADDER works its magic. First, it looks at the pictures the robot got right and the ones it got wrong. It then asks a language expert (the LLM) to read the descriptions or "captions" of those pictures. The language expert is like a detective with a magnifying glass, scanning the text for clues. It might say, "Wait a minute! Every time the robot got the 'pneumothorax' (a lung condition) diagnosis right, the report mentioned a 'chest tube.' But when it got it wrong, the chest tube was missing!" The robot wasn't looking at the lung; it was just looking for the tube.

Once LADDER spots these sneaky patterns, it doesn't just point them out; it fixes them. It creates a new set of rules (called "pseudo-labels") to teach the robot to ignore the chest tubes and actually look at the lungs. It's like telling the student, "Next time, ignore the bird and focus on the math problem." The paper tested this on all sorts of tricky tasks, from spotting birds in forests to finding cancer in medical scans. They tried it on over 200 different robot brains and found that LADDER was much better at finding these hidden mistakes than the old checklist methods.

The best part? LADDER doesn't need a human to write down every single thing it should look for. It figures it out on its own by reading the text that comes with the images. In fact, when they tested it on medical images, LADDER found biases that even the old methods missed, like the robot getting confused by the age of the patient or the specific type of machine used to take the X-ray. The paper suggests that by using this language-driven detective work, we can make AI much fairer and more reliable, especially in important fields like medicine where getting the answer right matters the most. It's not a magic wand that solves everything instantly, but it's a powerful new way to listen to our AI and help it learn the right lessons.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →