← Latest papers
🤖 AI

Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH

This paper outlines an ongoing initiative within the LLMs4EU project and ALT-EDIC infrastructure to adapt foundation models for Social Sciences and Humanities research by integrating knowledge graphs and multilingual corpora, employing a rigorous quantitative and qualitative evaluation framework to ensure reliable, ethically compliant, and domain-sensitive AI support for scholarly tasks.

Original authors: Adam Faci, Alessio Miaschi, Anne Combe, Pascal Cuxac, Francesca Frontini, Nicolas Larrousse, Stéphane Pouyllau

Published 2026-07-08
📖 5 min read🧠 Deep dive

Original authors: Adam Faci, Alessio Miaschi, Anne Combe, Pascal Cuxac, Francesca Frontini, Nicolas Larrousse, Stéphane Pouyllau

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Teaching AI to Read Like a Humanist

Imagine you have a super-smart robot librarian (a Large Language Model, or LLM) that has read almost everything on the internet. It's great at answering questions about science, math, and English news. But, if you ask it about history, philosophy, or literature in French or Italian, it often gets things wrong. It tends to ignore books, focus only on English journal articles, and miss the deep, messy, interpretive way human scholars actually think.

This paper describes a project called ReSearch_SSH (part of a bigger European initiative called LLMs4EU) that is trying to fix this. The goal is to take that "general" robot librarian and train it specifically to understand the Social Sciences and Humanities (SSH). They want an AI that respects different languages, understands that a book is just as important as a journal article, and knows that the same text can be interpreted in many different ways depending on who is reading it.

The Problem: The "One-Size-Fits-All" Trap

Currently, most AI tools for research are like a generic map of the world. They are drawn mostly in English and highlight only the biggest cities (major journals). If you try to navigate a small village in Italy or a specific historical debate in France using this map, you get lost.

The authors argue that this isn't neutral. It pushes out scholars who work in other languages or who study things like old manuscripts, digital blogs, or local history. It forces everyone to think like a "hard scientist" (looking for simple facts and citations) rather than a humanist (looking for context, nuance, and different perspectives).

The Solution: A Specialized Toolkit

Instead of building a brand new search engine from scratch, the team is upgrading an existing French research platform called ISIDORE. Think of ISIDORE as a massive, well-organized library basement. The team is installing a new "smart assistant" inside it.

Here is how they are building it, using three main tools:

1. The "Specialized Reading List" (Data)

You can't teach a chef to cook French cuisine if you only give them American recipes. To train the AI, they are feeding it a massive, curated diet of Social Science and Humanities texts.

  • The Menu: They are using about 3 million documents from a database called ISTEX. This includes books, articles, and data in over 60 languages, though French and English are the main courses.
  • The Side Dishes: To make sure the AI understands the specific jargon of Digital Humanities, they are adding Italian conference papers and specialized journals.
  • The Goal: This ensures the AI learns the specific "dialect" of scholars, not just the "dialect" of the internet.

2. The "Fact-Checking Map" (Knowledge Graphs)

This is the most important part. Standard AI often "hallucinates" (makes things up) because it guesses based on patterns. This project uses a Knowledge Graph, which is like a giant, interconnected web of facts.

  • The Analogy: Imagine the AI is a tour guide. A normal AI might guess where a landmark is. This AI, however, is holding a detailed map (Wikidata and other graphs) that links authors to their books, books to their themes, and themes to historical events.
  • How it works: When the AI answers a question, it doesn't just guess; it looks at the map, finds the exact page in the book, and says, "I found this answer in this specific document." This makes the answer traceable and reliable.

3. The "Human Feedback Loop" (Evaluation)

The team isn't just testing the AI with computer scores; they are testing it with real human experts.

  • The Panel: They have assembled a group of Digital Humanities scholars from France and Italy.
  • The Test: These experts will ask the AI complex questions and check if the answers make sense in a real research context. They are checking: "Is this answer actually useful for a historian? Did it miss a crucial nuance?"
  • The Result: If the AI fails, the experts tell it why, and the AI learns to think more like a human scholar.

The Rules of the Game (Legal & Ethics)

The paper emphasizes that this isn't a "wild west" experiment. They are building this inside a strict legal framework (like the EU AI Act and GDPR).

  • No Wild Guessing: The AI is designed to be a "research assistant," not a decision-maker. It suggests, but humans decide.
  • Privacy: They are careful not to use personal data to profile people.
  • Copyright: They are respecting the rules of the books they use. They aren't just stealing content; they are using licensed data and ensuring the AI points back to the original source so the authors get credit.

Where Are They Now?

The project is currently in the preparation phase.

  • They are gathering and cleaning the data (the "ingredients").
  • They are setting up the legal and ethical rules (the "kitchen safety protocols").
  • They are about to start the first round of training the model on this specific data.

The Bottom Line

This paper is a blueprint for building an AI that doesn't just "talk" like a human, but thinks like a scholar. By combining a specialized library of books, a fact-checking map, and a team of human experts, they hope to create a tool that helps researchers discover new ideas without losing the nuance, language, and ethics that make the Humanities special. It's about making sure the AI serves the scholar, rather than forcing the scholar to change how they work to fit the AI.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →