Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
The paper introduces Skywork-Reward-V2, a suite of state-of-the-art reward models trained on the 40-million-pair SynPref-40M dataset, which was created using a human-AI synergistic pipeline to overcome the limitations of existing preference data and achieve superior performance across multiple benchmarks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a very smart but sometimes stubborn robot how to write a story, solve a math problem, or just chat nicely with people. You can't just tell the robot, "Be good." You have to show it examples of what "good" looks like and what "bad" looks like.
In the world of AI, this process is called Reinforcement Learning from Human Feedback (RLHF). The robot needs a Reward Model (RM)—think of this as the robot's "taste teacher." This teacher looks at two different answers the robot gave and says, "I like this one better because it's more helpful," or "I hate this one because it's rude."
For a long time, these "taste teachers" were struggling. They were brittle, easily confused, and couldn't handle the messy, complicated way humans actually think. They were like a food critic who only knows how to judge pizza but gets confused when you serve them sushi.
Skywork-Reward-V2 is a new, super-powered "taste teacher" that fixes these problems. Here is how they did it, explained simply:
1. The Problem: Too Much Junk, Not Enough Gold
The researchers realized the problem wasn't the robot's brain; it was the training data.
- The Old Way: They tried to feed the teacher millions of examples, but many were labeled by other robots (AI) or were just random internet comments. It was like trying to teach a chef by showing them a mix of Michelin-star recipes and burnt toast from a gas station. The teacher got confused.
- The New Insight: They hypothesized that if they could find the "gold" in the trash, they could make a much better teacher.
2. The Solution: The "Human-AI Synergy" Pipeline
Instead of just hiring thousands of humans (which is slow and expensive) or just letting robots label everything (which is inaccurate), they built a two-stage assembly line that combines the best of both worlds.
Stage 1: The "Gold Standard" Workshop (Human-Led)
Imagine a small, elite team of expert human editors.
- They take a small pile of data and label it with extreme care.
- The Secret Sauce: These humans aren't just guessing. They are allowed to use search engines, math tools, and even other super-smart AIs to fact-check their answers before they label them.
- Analogy: It's like a judge in a cooking competition who doesn't just taste the food but also checks the recipe, the temperature of the oven, and the freshness of the ingredients before giving a score.
- The Result: A small pile of "Gold Data" that is 100% trustworthy.
Stage 2: The "Mass Production" Factory (AI-Led)
Now, they have the "Gold Standard" from Stage 1. They use this to train a "Gold Teacher" (a Reward Model).
- They take a massive mountain of "wild" data from the internet (40 million pairs!).
- The "Gold Teacher" looks at this mountain and says, "This looks like the Gold Standard, keep it." OR "This looks wrong, flip it." OR "This is confusing, let's ask a robot to double-check it using the Gold Standard as a guide."
- Analogy: Think of the Gold Teacher as a foreman on a construction site. He doesn't lay every brick himself. Instead, he trains a fleet of robots to lay bricks, but he constantly checks their work against his own blueprint. If a robot makes a mistake, the foreman corrects it.
- The Result: They turned that messy mountain of 40 million pairs into a clean, curated pile of 26 million high-quality examples.
3. The Outcome: The Skywork-Reward-V2 Series
Using this super-clean data, they trained a family of 8 new "Taste Teachers" (Reward Models).
- Size: They range from tiny (0.6 billion parameters) to medium-sized (8 billion parameters).
- Performance: Even the tiny ones are smarter than the massive 70-billion-parameter teachers from other companies.
- Why? Because Quality > Quantity. A teacher trained on 26 million perfect examples is better than one trained on 100 million confusing examples.
4. Why This Matters (The "So What?")
- Better Robots: These new teachers can now guide AI to be more helpful, safer, and more honest. They don't get tricked by fancy-sounding but wrong answers (resistance to "style bias").
- Cost-Effective: You don't need to spend millions of dollars hiring humans to label everything. You only need a small team of experts to set the rules, and then the AI can do the heavy lifting.
- The "Best-of-N" Superpower: If you ask an AI to generate 10 different answers and pick the best one, these new teachers are incredibly good at spotting the winner, even if the difference is tiny.
The Big Takeaway
This paper proves that we don't need to wait for bigger, more expensive computers to make better AI. Instead, we need better data curation.
By combining human expertise (to set the standard) with AI scale (to process the data), Skywork created a system that unlocks the hidden potential in existing data. It's like realizing that the whole ocean is full of fish, but you just needed a better net to catch the good ones.
In short: They built a better "taste teacher" by teaching it how to learn from the best human experts, and then letting it teach millions of other robots how to be good, too.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.