Histoires Morales: A French Dataset for Assessing Moral Alignment
This paper introduces "Histoires Morales," a culturally adapted French dataset derived from Moral Stories to address the lack of resources for assessing moral alignment in French language models, demonstrating that while these models generally adhere to moral norms by default, they remain vulnerable to manipulation through user-preference optimization.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Large language models are powerful tools that can write, translate, and converse by predicting the next word in a sentence based on patterns they have learned from vast amounts of text. As these systems become more integrated into daily life, a critical question has emerged: do they understand human values? Specifically, can they distinguish between right and wrong in the same way people do? While researchers have made significant progress in testing these models in English and Chinese, a major gap remained for French, the fifth most spoken language in the world. Without a way to measure how these models handle moral reasoning in French, it is difficult to know if they will behave ethically for the hundreds of millions of people who speak it. This uncertainty matters because a model that fails to grasp cultural nuances or moral norms could offer harmful advice or reinforce biases in a way that feels natural to the machine but wrong to the user.
To address this gap, a team of researchers from France created a new resource called HISTOIRESMORALES. This dataset consists of 12,000 short stories written in French, each describing a social situation where a person faces a choice between a moral action and an immoral one. The stories cover a wide range of everyday scenarios, from how one treats friends and family to how one handles money or animals. For instance, one story might describe a person who ignores a phone call from their parents to go out with friends, while another might depict someone who pays a veterinarian for their work versus someone who refuses to pay. Each narrative includes the context, the character's intention, the action they take, and the consequence of that action. The researchers did not simply translate these stories from English; they carefully adapted them to fit French culture. This meant changing names like "John" to "Jean," converting currency from dollars to euros, and swapping American activities like playing baseball for more common French pastimes like playing tennis. They also ensured that idioms and social cues felt natural to a native French speaker, rather than sounding like a direct, awkward translation.
The team built this dataset by starting with an existing collection of English stories and using a combination of automated translation and human review. They first used a language model to translate the text, then asked native French speakers to correct errors and ensure the cultural context was accurate. This process involved multiple rounds of checking, where human annotators evaluated whether the translations preserved the original meaning, used correct grammar, and adapted cultural references appropriately. The researchers found that simply translating words was not enough; the model often missed the tone or the cultural weight of a situation. By adding specific examples of errors and corrections to the translation instructions, they significantly improved the quality of the final dataset. The result is a collection of stories that feels authentic to French speakers, allowing researchers to test how well artificial intelligence understands the moral fabric of French society.
Using this new dataset, the researchers tested several large language models to see how they performed on moral reasoning tasks. They asked the models to predict whether a given action was likely to happen or to choose between a moral and an immoral option. The results showed that while the models generally aligned with human moral norms, they were not perfect. In fact, the models performed slightly better when processing the stories in English than in French, suggesting that their training data might be richer or more balanced in English. More surprisingly, the researchers discovered that these moral alignments were fragile. By using a technique called direct preference optimization, which involves showing the model a small number of examples of preferred behavior, they could easily shift the model's moral compass. With as few as 84 examples, the models could be trained to prefer immoral actions over moral ones, or vice versa. This indicates that the models do not have a deep, unchangeable understanding of right and wrong; rather, their behavior is highly sensitive to the data they are exposed to.
The study concludes that while current language models are generally helpful and aligned with human values, their moral reasoning is not robust, especially when operating in languages other than English. The ability to easily influence a model's moral stance with a small amount of training data suggests that these systems are not yet reliable enough to be trusted with complex ethical decisions without careful oversight. The creation of HISTOIRESMORALES provides a vital tool for researchers to continue monitoring and improving these systems, ensuring that as they become more capable, they also become more trustworthy across different cultures and languages. The work highlights that building ethical artificial intelligence requires more than just translating code; it demands a deep understanding of the cultural and moral contexts in which these tools will be used.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.