Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks
This paper introduces TSEF, a dual-target attack demonstrating that temporal consistency in time series explanations is a misleading proxy for robustness, as it is possible to adversarially decouple predictions from plausible explanations to achieve targeted misclassifications.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "Trustworthy" Lie
Imagine you have a high-tech security guard (a Time Series Classifier) that watches a video feed of a factory machine to tell if it's working normally or about to break. Because the guard is a "black box" (we don't know exactly how its brain works), the factory owners also use a Spotlight (an Explainer).
The Spotlight shines on the parts of the video the guard is looking at. If the guard says "Danger!" and the Spotlight shines on a smoking pipe, the workers trust the warning. They assume: If the explanation looks right, the decision must be right.
The Paper's Discovery:
The researchers found a way to trick the system so that the guard changes its mind (e.g., from "Safe" to "Danger"), but the Spotlight still shines on the exact same spot it did before. The workers see the "Danger" warning and the "smoking pipe" explanation and think, "Okay, the system is working perfectly," while the system has actually been completely fooled.
The Problem: The "High-Dimensional Paradox"
The paper explains why this is hard to do with time series data (like heartbeats or stock prices).
- The Old Way (Single-Target Attack): Imagine trying to change the guard's mind by poking it randomly all over the body. You might succeed in making it scream "Danger," but because you poked everywhere, the Spotlight gets confused. It starts shining on random, scattered spots. The workers see the "Danger" warning but also see a messy, confusing Spotlight. They get suspicious: "Wait, why is the light everywhere? Something is wrong."
- The New Problem: The researchers call this the "High-Dimensional Paradox." Time series data has so many points (time steps and sensors) that if you try to change the prediction without being careful, the "noise" spreads out too much. The explanation becomes messy and fails to look like a natural, focused reason.
The Solution: TSEF (The "Master of Disguise")
The authors created a new attack tool called TSEF (Time Series Explanation Fooler). Think of TSEF as a master stage magician who can change the outcome of a trick without the audience noticing the sleight of hand.
TSEF does two things at once, like a two-step dance:
Step 1: Find the Weak Spot (The Temporal Mask).
Instead of poking the whole machine, TSEF looks for a tiny, specific window of time where the machine is most sensitive. It's like finding the one loose screw that, if wiggled, makes the whole machine shake. It ignores the rest of the machine.- Analogy: It's like a thief who knows exactly which one window is unlocked, rather than trying to break every window in the house.
Step 2: The Smooth Edit (The Frequency Filter).
Once it finds that weak spot, TSEF doesn't just add random noise. It edits the pattern of the signal (like changing the rhythm or the wave shape) in a way that looks natural.- Analogy: Instead of scribbling on a painting with a marker (which looks messy), it carefully repaints a small section to change the story, but the brushstrokes look perfectly smooth and consistent with the rest of the art.
The Result: The "Cover-Up"
When TSEF attacks the system:
- The Prediction Changes: The guard flips from "Normal" to "Abnormal" (or vice versa).
- The Explanation Stays the Same: The Spotlight continues to shine on the exact same "smoking pipe" it was shining on before.
The result is a perfect cover-up. The system is lying about the danger, but the "proof" (the explanation) looks 100% honest and consistent.
Why This Matters (According to the Paper)
The paper argues that we have been making a dangerous mistake. We assume that if an AI's explanation is stable (doesn't change much) and consistent (matches what we expect), then the AI is robust (hard to hack).
The paper proves this assumption is wrong.
- The Trap: An attacker can force the AI to make a specific wrong decision while keeping the explanation looking perfectly normal.
- The Consequence: If a doctor or engineer relies on the explanation to verify the AI's decision, they will be fooled. They will see a "plausible" reason for a wrong decision and trust it.
Summary of the "Magic Trick"
- Old Attacks: Break the prediction, but leave a messy, obvious trail of evidence (bad explanation).
- TSEF (New Attack): Breaks the prediction and fakes the evidence so perfectly that the explanation looks exactly like it did before the attack.
The paper concludes that we cannot trust "explanation stability" as a safety check. Just because the AI gives a good reason for its decision doesn't mean the decision is safe or correct; the AI could be a master of disguise.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.