SEF-CLGC at SemEval-2026 Task 11: Logical Notation Impact on Language Model Performance
This paper presents the SEF-CLGC pipeline, which integrates formal logical notations with Small Language Models to achieve a 27.80% content score on SemEval-2026 Task 11 while significantly reducing content bias in reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a very smart, but tiny, robot how to solve logic puzzles. Usually, when people build these robots, they make them enormous and feed them the entire internet to learn. But this paper is about a different approach: using Small Language Models (SLMs). Think of these as "pocket-sized" robots that are frugal with energy but still capable of complex thinking if trained correctly.
The researchers entered a competition called SemEval-2026 Task 11. The challenge was simple on the surface: look at a logical argument (called a syllogism) and decide if it is True (valid) or False (invalid).
Here is the tricky part: Humans (and big AI models) often get tricked by how the argument sounds. If the conclusion sounds like a real-world fact (e.g., "All cats are mammals"), we might say "True" even if the logic is broken. This is called content bias. The goal was to teach the robot to ignore the "flavor" of the words and focus only on the "recipe" (the logic structure).
The Secret Sauce: Translating the Puzzle
The team used a pipeline they built called SEF-CLGC. Here is how it works, using a cooking analogy:
- The Raw Ingredients (Natural Language): The puzzle starts as a sentence in English, like "All cars are vehicles. No animal is a car..."
- The Translation (FOL): First, they use a powerful translator (an AI called ChatGPT) to turn that English sentence into First-Order Logic (FOL). This is like translating a recipe from "add a pinch of salt" to a precise chemical formula. It removes all the ambiguity.
- The Flavor Variations (Logical Notations): Once they have the chemical formula, they translate it into different "dialects" of logic. Some look like standard math, some look like computer code (CLINGO), and others are custom-made shortcuts.
- Analogy: Imagine you have a song. You can play it on a piano (Natural Language), a synthesizer (FOL), or a drum machine (CLINGO). The song is the same, but the instrument changes how it sounds.
- The Training: They fed these different versions of the logic puzzles to their tiny robot models. They wanted to see if teaching the robot to read the "chemical formulas" (logic) instead of just the "English sentences" would help it stop getting tricked by the content.
The Results: Small but Mighty
The researchers tested their models and found some surprising things:
- The "Hybrid" Approach Won: The best model wasn't one that only spoke English or one that only spoke pure logic. The winner was a model that saw both the English sentence and the Logic formula together. It was like giving the robot a map and a compass.
- Beating the Bias: By training on these logical "dialects," the models became much better at ignoring the real-world facts and focusing on the rules. They reduced the "content bias" significantly.
- The Score: Their best model achieved a score of 27.80% on a specific metric designed to measure how well they balanced accuracy with fairness (ignoring bias). While that number might look low, in this specific, bias-heavy test, it was a strong performance for such a small model.
- What Didn't Work: Some of the custom, highly abstract logical languages were too confusing for the tiny robots. It's like trying to teach a beginner to read using a language with no vowels; the robot just got lost.
The Takeaway
The paper claims that you don't need a massive, energy-hungry supercomputer to solve complex logic problems. By using Small Language Models and teaching them to speak "logic languages" alongside natural language, you can create a system that is:
- Frugal: It uses very little computing power.
- Fairer: It is less likely to be tricked by how an argument sounds.
- Competitive: It can hold its own against much larger models in logic tasks.
The authors also noted a limitation: they relied on a commercial AI to do the initial translation from English to Logic. If that commercial AI changes its mind or gets updated, it could mess up the whole pipeline, like if the translator suddenly started speaking a different dialect.
In short, they proved that with the right "translation tools," a small, efficient robot can learn to think logically without getting distracted by the story it's reading.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.