← Latest papers
💻 computer science

Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

This paper demonstrates that linear representations of temporal horizon in the Qwen3-32B model can be identified via contrastive linear probes and utilized through activation addition to effectively steer the model's intertemporal preferences and planning capabilities in both short-term and long-term directions.

Original authors: Michal Mráz, Justin Shenk

Published 2026-08-05
📖 6 min read🧠 Deep dive

Original authors: Michal Mráz, Justin Shenk

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are talking to a very smart, very fast robot that has read almost every book on the internet. You might think this robot just spits out facts, but it's actually more like a person with a personality. It has preferences, habits, and even a sense of "when" to do things. Sometimes it wants to grab a cookie right now (short-term), and other times it's willing to wait for a bigger cake later (long-term). This paper is about a new way to peek inside the robot's brain to find the specific "switch" that controls this sense of time. Scientists call this "intertemporal preference," which is just a fancy way of asking: "Do you want it now, or later?" Understanding this is crucial because if we can't control how a robot weighs immediate rewards against future benefits, it might give us bad advice about saving money, planning a trip, or even making big life decisions.

The researchers in this study decided to test this on a specific, powerful robot brain called Qwen3-32B. They didn't just ask the robot questions; they tried to physically "steer" its thoughts while it was thinking. Think of the robot's brain as a giant river of electricity flowing through layers of pipes. The team found a specific direction in that river that represents "long-term thinking." By gently pushing the electricity in that direction (or pulling it back), they could make the robot suddenly change its mind. They proved that they could make the robot much more patient, willing to wait years for a huge reward, or much more impatient, desperate for a small reward right this second. It's like having a remote control for the robot's patience.

The Secret "Time Dial" in the Robot's Brain

The team started by teaching the robot to answer questions about time. They gave it pairs of questions where one answer was about "what to do in the next 30 days" and the other was about "where to be in 10 years." As the robot processed these answers, the researchers watched the electricity flowing through its brain. They discovered that the difference between a "short-term" answer and a "long-term" answer wasn't messy or random; it was a straight, clean line. They could draw a vector (a mathematical arrow) that pointed from "now" to "later."

To test if this arrow was real, they used a technique called Contrastive Activation Addition (CAA). Imagine the robot's brain is a car driving down a road. The researchers found a specific "steering wheel" hidden inside the engine. When they turned this wheel slightly to the left (adding a negative push), the car swerved toward "immediate action." When they turned it to the right (adding a positive push), the car swerved toward "future planning." They didn't just guess this; they measured it. They found that by adding a tiny bit of this "time vector" to the robot's brain while it was answering, they could force it to change its mind on binary choices (A vs. B) with very high reliability, shifting the model's preferences significantly in both directions.

The Money Test: Can We Change Its Wallet?

The real magic happened when they tested this on a completely different kind of question: money. They didn't teach the robot about money; they only taught it about "short vs. long" timeframes. Then, they asked it a classic human dilemma: "Would you rather have $100 today or $1,000 in a year?"

Without any steering, the robot had a certain "indifference point"—a specific amount of money in the future that made it say, "Eh, I'll wait." But when the researchers applied their "long-term steering," the robot's patience skyrocketed. Suddenly, it was willing to wait for a future reward even if the extra amount it gained by waiting was relatively small. In other words, the "premium" it demanded to wait dropped significantly. Conversely, when they applied "short-term steering," the robot became greedy, demanding a massive reward just to wait a single day.

The numbers were wild. At the strongest steering level, the robot's preference shifted so drastically that at a 10-year horizon, it required more than 56 times the future reward to be willing to wait when steered toward the short-term, compared to when it was steered toward the long-term. This suggests that the robot's sense of time isn't just a fixed setting; it's a dial that can be turned up or down, even on tasks it wasn't explicitly trained for.

Planning Trips: Does Patience Make a Better Travel Agent?

Finally, the team wondered if making the robot more patient would actually make it better at complex tasks. They used a benchmark called TravelPlanner, where the robot has to create a multi-day trip itinerary with hotels, flights, and restaurants. This is hard because you have to think about the whole week, not just the next hour.

They found a sweet spot. When they gave the robot a moderate "long-term" push, its travel plans became more sensible and realistic. It stopped making silly mistakes that a short-sighted planner would make. However, if they pushed the "long-term" dial too hard (the strongest setting), the robot got confused and started making up nonsense. It seems that while a little bit of future-thinking helps, too much can break the robot's ability to focus on the details.

What This Means for Us

The big takeaway is that AI models have a measurable "time horizon," and we can change it. This isn't just a cool trick; it's a safety issue. If an AI is giving you advice on investing, health, or climate change, its advice depends heavily on whether it cares about next week or next century. The researchers showed that we can measure this preference and steer it.

However, they are careful to say this isn't a magic wand that fixes everything. The "time dial" might also be tangled up with other things, like how abstract or urgent the robot feels. And if we turn the dial too far, the robot might stop making sense entirely. But the fact that we can find this switch and flip it suggests that the way AI thinks about the future isn't a mystery we can't touch. It's a feature we can understand, measure, and potentially control, which is a huge step forward for making sure AI stays helpful and aligned with our long-term goals.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →