← Latest papers
💬 NLP

L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention

The paper proposes L2V-CoT, a training-free method that enhances Vision-Language Models' reasoning capabilities by transferring Chain-of-Thought representations from Large Language Models via latent intervention in the frequency domain, leveraging their shared low-frequency latent structures without requiring architectural alignment or extensive training.

Original authors: Yuliang Zhan, Xinyu Tang, Han Wan, Jian Li, Ji-Rong Wen, Hao Sun

Published 2026-03-23
📖 4 min read☕ Coffee break read

Original authors: Yuliang Zhan, Xinyu Tang, Han Wan, Jian Li, Ji-Rong Wen, Hao Sun

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Smart Text" vs. The "Dumb Picture"

Imagine you have two friends:

  1. The Text Genius (LLM): This friend is a master at solving complex math problems, logic puzzles, and step-by-step reasoning. They can break a huge problem down into tiny, manageable steps.
  2. The Picture Lover (VLM): This friend is great at describing what they see in a photo or a chart. However, when you ask them to solve a tricky math problem inside a picture, they often get stuck. They tend to guess the answer immediately without thinking it through, because they haven't been taught how to "think step-by-step" with images.

The problem is that teaching the Picture Lover to think like the Text Genius usually requires massive amounts of expensive data and years of training. It's like trying to teach a fish to fly by building it wings; it's hard and inefficient.

The Discovery: They Speak the Same "Secret Language"

The researchers asked a bold question: Do these two friends actually think in the same way deep down, even if they look different on the surface?

To find out, they used a tool called Linear Artificial Tomography (LAT). Think of this like an X-ray machine for the brain. Instead of looking at the words the models say, they looked at the electrical signals (the "thoughts") happening inside the models while they were solving problems.

The Surprise:
They discovered that when the Text Genius thinks logically, it produces a specific type of "low-frequency" brain signal. When they looked at the Picture Lover, they found the exact same type of low-frequency signal, but it was buried under a lot of "static noise" (confusion caused by mixing pictures and text).

The Analogy:
Imagine the Text Genius is playing a clear, pure musical note (the logic). The Picture Lover is trying to play the same note, but they are also playing it through a broken speaker that adds static and distortion. The core melody (the logic) is actually there; it's just hard to hear.

The Solution: L2V-CoT (The "Brain Boost")

Instead of retraining the Picture Lover from scratch, the researchers invented a method called L2V-CoT. Think of this as a hearing aid or a neural transplant that happens in real-time.

Here is how it works, step-by-step:

  1. Extract the Pure Melody: They take the Text Genius and ask it to solve a problem. They record the "pure" low-frequency logic signal (the clear musical note).
  2. Filter the Noise: They use a digital filter (Low-Pass Filter) to clean up the signal, removing any high-frequency "static" that doesn't belong to the logic.
  3. Resize the Signal: Since the Text Genius and Picture Lover have different brain sizes (different dimensions), they use a "resampling" trick to shrink or stretch the signal so it fits perfectly into the Picture Lover's brain.
  4. Inject the Logic: When the Picture Lover is asked a question, the researchers secretly inject this clean, pure logic signal into the Picture Lover's brain while it is thinking.

The Result:
Suddenly, the Picture Lover starts thinking like the Text Genius! It stops guessing immediately. Instead, it starts "slow thinking," breaking the problem down into steps, just like the expert.

Why This is a Big Deal

  • No Training Required: Usually, to make a model smarter, you have to feed it millions of books and pictures and wait months for it to learn. This method works instantly. It's like giving someone a cheat sheet instead of making them go to school for four years.
  • Works on Any Model: It doesn't matter if the Picture Lover is a small model or a big model. The "logic signal" fits them all.
  • Better than Fine-Tuning: In their tests, this "instant injection" method actually performed better than models that had been trained for weeks on expensive data.

The Bottom Line

The paper proves that reasoning is a universal language that exists in the "low-frequency" part of AI brains, regardless of whether the AI sees text or pictures. By simply "tuning" the Picture Lover's brain to pick up the Text Genius's frequency, we can make visual AI much smarter, faster, and cheaper to build.

In short: They didn't teach the fish to fly; they just gave it a jetpack that uses the same fuel as the bird.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →