← Latest papers
💻 computer science

ProDG: Prototypes for Data-Free Generative Post-Hoc Explainability

This paper introduces ProDG, a novel data-free framework that leverages generative models to synthesize high-fidelity visual prototypes directly from frozen model weights, thereby enabling robust post-hoc interpretability for privacy-sensitive domains without requiring access to any external data.

Original authors: Piotr Borycki, Magdalena Trędowicz, Jacek Tabor, Łukasz Struski, Przemysław Spurek

Published 2026-05-21
📖 4 min read☕ Coffee break read

Original authors: Piotr Borycki, Magdalena Trędowicz, Jacek Tabor, Łukasz Struski, Przemysław Spurek

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot that can look at a picture and tell you exactly what it is—like identifying a specific breed of dog or a rare bird. But here's the catch: the robot is a "black box." It gives you the answer, but it won't tell you why it thinks that. It's like a chef who serves you a perfect meal but refuses to tell you what ingredients they used.

For a long time, scientists tried to peek inside the black box. Some methods just highlighted "hot spots" on the image (like a heat map), but those spots were often blurry and didn't make much sense to humans. Others tried to find real examples from the robot's training data that looked like the answer, saying, "See? This picture looks like that one, so that's why I chose it."

The Problem with the Old Way
The problem with finding real examples is that you need access to the robot's original training data. But in many real-world situations—like in hospitals with private patient records or banks with sensitive financial data—that data is locked away. You can't look at it. So, the old methods hit a wall: they can't explain the robot if they can't see the data.

The New Solution: ProDG
This paper introduces ProDG (Prototypes for Data-Free Generative Post-Hoc Explainability). Think of ProDG as a "magical sketch artist" that doesn't need to see the original photo album to understand the robot's mind.

Here is how it works, using simple analogies:

1. The "Brain Surgery" (Feature Disentanglement)

The robot's brain (its neural network) is messy. When it sees a "wolf," it might activate 50 different parts of its brain at once, all jumbled together.

  • The Fix: ProDG performs a gentle "brain surgery." It uses a special mathematical tool (called an Orthogonal Feature Disentanglement Module) to untangle the mess. It rearranges the robot's internal wiring so that each specific wire (or "channel") is dedicated to just one clear idea, like "wolf eyes" or "wolf snout," without the other ideas getting in the way.
  • The Magic: It does this without changing the robot's actual answers. It's like reorganizing a library so books are easier to find, but the books themselves and the librarian's knowledge remain exactly the same.

2. The "Dream Generator" (Generative Prototypes)

Usually, to explain the robot, you'd need to dig through a pile of real photos to find a perfect example of a "wolf snout." But since we can't see the real photos, ProDG uses a Generative Model (like a high-tech AI artist).

  • The Process: ProDG talks to this AI artist. It says, "Hey, I need you to draw a picture that makes the 'wolf snout' wire in the robot's brain light up as brightly as possible."
  • The Result: The AI artist doesn't copy a real photo; it dreams up a brand new, perfect "wolf snout" image from scratch. It keeps trying different sketches until it finds the one that makes the robot's brain say, "Yes! That is exactly what I was thinking!"

3. The "Concept Bank" (Prompt Optimization)

To make sure the AI artist draws the right thing, ProDG uses a Concept Prompt Bank. Imagine this as a dictionary of instructions.

  • Instead of just saying "draw a wolf," ProDG fine-tunes the instructions (the "prompts") to be incredibly specific. It tweaks the instructions slightly to ensure the artist draws a variety of different "wolf snouts" (some with fur, some with scars, some in different lighting) so the explanation isn't just one boring, static image. This prevents the artist from getting stuck drawing the exact same picture every time.

Why This Matters

  • Privacy First: Because ProDG creates these explanations from scratch using the robot's internal "weights" (its brain structure) and a generative artist, it never needs to see the original private data. You can explain a medical diagnosis without ever looking at the patient's private records.
  • Clearer Explanations: Instead of showing a blurry red blob on a picture (like older methods), ProDG shows you a clear, high-quality image of the specific feature the robot is looking at (e.g., "I see the antennae, so I know this is an insect").

The Bottom Line

ProDG is a new way to open the black box. It takes a frozen, unchangeable AI, untangles its messy brain, and then uses a generative AI artist to "dream up" perfect examples of what the robot is thinking. This allows us to understand complex AI decisions even when the original data is strictly off-limits, keeping privacy safe while making the AI's reasoning crystal clear.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →