← Latest papers
💻 computer science

A dataset of rated conceptual arguments

This paper introduces a dataset of 951 expert-rated critiques of 442 position texts on conceptual questions (such as AI safety and ethics) to evaluate large language models' ability to assess individual arguments, finding that performance on these tasks correlates with general model capability rankings.

Original authors: Emery Cooper, Caspar Oesterheld, Linh Chi Nguyen, Alexander Kastner, Ethan Perez

Published 2026-07-31
📖 3 min read☕ Coffee break read

Original authors: Emery Cooper, Caspar Oesterheld, Linh Chi Nguyen, Alexander Kastner, Ethan Perez

Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you're trying to teach a robot how to be a philosopher. You can easily teach it math or coding because those subjects have a clear "answer key." If the robot says 2+2=52+2=5, you know it's wrong. But what happens when the question is, "Is free will real?" or "What is the best way to treat people fairly?" There is no answer key for these "conceptual questions." No one knows the ground truth, and experts often disagree forever. This is the tricky corner of science where artificial intelligence is trying to take its next big step.

Usually, to teach a computer, we need to show it right and wrong answers. But when the answers are fuzzy, how do you grade the computer? You can't just ask, "Did you get the right conclusion?" because nobody knows what the right conclusion is. Instead, this paper suggests a clever workaround: stop grading the final answer and start grading the arguments used to get there. Think of it like a debate tournament. Even if we don't know who is "right" about the meaning of life, we can still agree that one debater made a clearer, more logical, and more central point than the other. The paper builds on the idea that while we can't solve the big philosophical mysteries yet, we can teach AI to spot good reasoning versus bad reasoning, even in the fog of uncertainty.

This is exactly what the researchers set out to do. They created a massive new dataset—a digital library of 951 "rated" debates—where human experts acted as judges. They didn't just ask, "Who won?" Instead, they broke down every argument into specific ingredients: How central was the point to the main topic? How strong was the attack? Was the critique clear, or was it confusing? Did it contain false facts? They then used this dataset to test how well modern AI models could play the role of the judge.

The results were a mix of good news and a bit of a reality check. The study found that smarter, more powerful AI models generally make better judges of philosophical arguments. If you have a "heavyweight" AI, it's better at spotting a weak argument than a "lightweight" one. This suggests that as AI gets better at general thinking, it naturally gets better at spotting the quality of reasoning in fuzzy, conceptual debates.

However, there was a surprising twist. The researchers tested a popular feature in modern AI called "thinking" or "reasoning," where the model pauses to think through a problem step-by-step before answering. You might expect that giving the AI more time to think would make it a better judge. But the paper found that, surprisingly, this "thinking" mode didn't really help much. Sometimes it helped a tiny bit, sometimes it made things slightly worse, but on average, the difference was negligible. The authors suggest this might be because these "thinking" models are currently trained mostly on math and coding, where answers are black-and-white. They haven't quite learned how to use that extra thinking time for the gray areas of philosophy and ethics.

In short, the paper proves that we can build a dataset to teach AI how to argue better, even when there's no single right answer. It shows that while AI is getting good at judging these debates, simply telling it to "think harder" isn't the magic button we hoped for. The path to better AI reasoning in philosophy might require more than just more processing power; it might need a different kind of training entirely.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →