← Latest papers
💬 NLP

VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models

This paper introduces VALUEFLOW, a unified framework that addresses key gaps in value-based alignment by integrating hierarchical value extraction, a large-scale intensity-labeled database, and a calibrated evaluation system to enable the pluralistic and steerable control of value intensities in Large Language Models.

Original authors: Woojin Kim, Sieun Hyeon, Jusang Oh, Jaeyoung Do

Published 2026-06-08
📖 4 min read☕ Coffee break read

Original authors: Woojin Kim, Sieun Hyeon, Jusang Oh, Jaeyoung Do

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a very smart, but somewhat rigid, robot how to act like a human. The problem isn't just that the robot needs to be "nice"; it's that humans have complex, sometimes conflicting, internal compasses called values.

Some people care deeply about kindness, others about justice, and some about power. The tricky part is that these values aren't just "on" or "off" switches. You can be slightly kind, or extremely kind. You can be mostly just, but a little bit selfish.

Current AI alignment methods are like trying to teach the robot with a simple "Good/Bad" button. It's too blunt. The paper VALUEFLOW introduces a new, much more sophisticated toolkit to help AI understand and express these human values with the right intensity.

Here is how VALUEFLOW works, broken down into three simple parts:

1. The Map: HIVES (The Hierarchical Value Embedding Space)

Imagine you have a giant library of human thoughts, but the books are thrown everywhere. Some books are about "Kindness," others about "Helping Neighbors," and others about "Donating Money." A simple search might confuse "Kindness" with "Justice" because they are related but different.

HIVES is like a super-organized library map. It doesn't just list values; it understands the hierarchy. It knows that "Kindness" is a big umbrella, and under it, you have "Caring" and "Helping." It organizes these values into a 3D space where similar values are close together, and very different ones are far apart. This helps the AI understand the structure of human values, not just the words.

2. The Reference Library: VIDB (The Value Intensity Database)

Now, imagine you want to tell the AI, "Be moderately kind, but very strict about justice." How does the AI know what "moderately" looks like?

Previously, AI judges would just guess a score (like "7 out of 10"). But different AIs guess differently, and their scores are unstable.

VIDB is a massive, pre-calibrated library of text examples. It contains thousands of sentences, each labeled with a precise "intensity score" (from -10 to +10) for specific values.

  • The Analogy: Think of VIDB as a set of standard weights on a scale. If you want to know how "heavy" a new sentence is in terms of "Justice," you don't just guess. You compare it against these standard weights.
  • How it works: Instead of asking the AI "Is this text just?" (which leads to bad guesses), the system asks the AI to rank the new text against a few standard examples from the library. "Is this text more just than Example A, but less just than Example B?" This ranking method is much more stable and reliable than guessing a number.

3. The Steering Wheel: Intensity-Aware Steering

This is the part where you actually control the AI.

In the past, if you told an AI "Be kind," it might be too kind or not kind enough. With VALUEFLOW, you can give the AI a dial.

  • The Analogy: Imagine driving a car. Old methods were like having a gas pedal that only had "Off" and "Full Throttle." VALUEFLOW gives you a steering wheel with a speedometer. You can say, "Drive with a 'Kindness' intensity of +8" or "Drive with a 'Justice' intensity of -2" (meaning, reject injustice).

The paper tested this on many different AI models and found:

  • Some models are easier to steer than others. Some respond well to your instructions, while others are stubborn.
  • Values fight each other. If you tell the AI to be "Very Kind" but also "Very Strict," it has to find a balance. The paper found that sometimes one value wins out, and sometimes they mix in predictable ways.
  • It works for groups. They used this to figure out what values different groups of people (like different political parties or cultures) hold, and then steered the AI to sound more like those groups. This made the AI's predictions about what those groups would say much more accurate.

The Big Picture

The authors built a complete system that:

  1. Maps values into a structured space (HIVES).
  2. Measures how strong a value is in a text by comparing it to a standard library (VIDB).
  3. Steers the AI to express those values at specific strengths.

They aren't saying this will solve all AI safety problems or that we should use this to manipulate people. Instead, they are providing a scientific toolkit to study how AI behaves when we ask it to be "a little bit" of something versus "a lot" of something. It's a step toward making AI that can understand the nuance of human values, rather than just following a simple "good vs. bad" rulebook.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →