← Latest papers
💬 NLP

LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation

This paper introduces the IHLC system for the LT-EDI 2026 Shared Task, which combines LoRA fine-tuning for gender-inclusive rewriting and a compute-efficient activation steering technique using PCA-derived principal directions for counter-narrative generation, achieving high scores while analyzing key limitations like semantic drift and bias leakage.

Original authors: Akhil Rajeev P, Manoj Balaji J

Published 2026-07-28
📖 3 min read☕ Coffee break read

Original authors: Akhil Rajeev P, Manoj Balaji J

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are talking to a very smart, very well-read robot that has read almost every book on the internet. This robot is great at writing stories and answering questions, but it has a problem: it learned from human books, and sometimes those books contain old-fashioned, unfair ideas about boys and girls. If you ask the robot to write a story about a "fireman," it might stubbornly insist on that word, even though "firefighter" is better. If you ask it to argue against a mean statement like "girls are bad at math," it might accidentally agree with the mean idea or sound too robotic. This is the world of Artificial Intelligence (AI) and Natural Language Processing (NLP), where scientists try to teach computers to speak more fairly and inclusively. The big challenge is figuring out how to tweak the robot's brain so it naturally chooses kind, neutral words without having to rebuild its entire brain from scratch.

In this paper, a team of researchers named IHLC tried to solve this problem for a special contest called LT-EDI 2026. They focused on two main tasks: first, simply rewriting biased sentences to be fair (like changing "fireman" to "firefighter"), and second, creating a "counter-narrative"—a polite, persuasive response to a mean or biased statement. For the first task, they used a standard method called LoRA, which is like giving the robot a small, specialized cheat sheet to help it learn new rules quickly. But for the second task, they tried something much more experimental and clever called Activation Steering.

Think of the robot's brain as a giant, complex machine with thousands of gears turning inside. Usually, to change how the machine works, you have to take it apart and replace the gears (this is called "fine-tuning"). But the researchers asked: "What if we could just give the machine a gentle nudge while it's running?" They found a specific "nudge" direction by comparing how the robot's brain reacts to a mean sentence versus a kind one. They calculated the difference between these two reactions and created a "steering vector"—imagine it as a magnetic force that pushes the robot's thoughts toward kindness and inclusivity. They injected this force into the robot's brain while it was writing, without changing any of its original gears.

The results were a mix of success and interesting glitches. For the simple rewriting task, their system did very well, scoring 80.00% and ranking 3rd out of 9 teams. For the harder task of writing counter-narratives, they scored 78.12% and ranked 6th out of 7 teams. The system was great at being polite and understanding the context, but sometimes it got a little too excited by the "nudge." When they pushed the steering too hard, the robot started repeating itself, fusing words together (like "exhibitimpulsiveness"), or drifting away from the original meaning to give a long lecture instead of a direct answer.

The researchers found that while this "nudge" method is a powerful and efficient way to make AI more inclusive, it's not perfect. It sometimes leaves behind tiny bits of the original bias, or it changes the sentence so much that it loses its original flavor. They suggest that in the future, they might need to combine this "nudge" with other tools to keep the robot's answers both fair and faithful to the original question. It's a promising step toward making AI a kinder conversationalist, but it shows that even with a gentle push, the robot still needs careful guidance to get it just right.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →