← Latest papers
💬 NLP

HebID: Detecting Social Identities in Hebrew-language Political Text

This paper introduces HebID, the first multilabel Hebrew corpus for detecting nuanced social identities in political discourse, which is used to benchmark language models and analyze the alignment between elite identity expression and public priorities in Israel.

Original authors: Guy Mor-Lan, Naama Rivlin-Angert, Yael R. Kaplan, Tamir Sheafer, Shaul R. Shenhav

Published 2026-02-24
📖 4 min read☕ Coffee break read

Original authors: Guy Mor-Lan, Naama Rivlin-Angert, Yael R. Kaplan, Tamir Sheafer, Shaul R. Shenhav

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking through a bustling marketplace where politicians are shouting their messages to the crowd. They talk about everything from the economy to religion, but they often do it by wearing "invisible badges" that signal who they are and who they are talking to. Some wear a badge that says "I'm a traditionalist," others wear "I'm a modern liberal," and some wear "I'm a security hawk."

For a long time, computers (AI) have been great at reading English texts and spotting these badges. But when it comes to Hebrew, the language of Israel, the computers were mostly blind. They didn't have a dictionary or a rulebook to understand these subtle social identities in Hebrew political speech.

Enter HEBID. Think of HEBID as the first-ever giant, high-tech translator's guidebook specifically designed to spot these social identity badges in Hebrew political text.

Here is a breakdown of how they built it and what they found, using simple analogies:

1. Building the Rulebook (The Dataset)

The researchers didn't just guess which badges mattered. They went to the source: the people.

  • The Survey: They asked thousands of regular Israeli citizens, "What groups do you feel you belong to?" (Like: "I'm a right-winger," "I'm religious," "I care about social justice").
  • The Selection: Based on what people actually said mattered to them, the researchers picked the top 12 most popular "badges."
  • The Training: They took 5,500 sentences from Israeli politicians' Facebook posts and had human experts manually tag every sentence with the correct badges. For example, a sentence might wear two badges at once: "Rightist" and "Security-oriented."
  • The Result: A massive, labeled library of Hebrew sentences that teaches computers what these specific identities look like in real life.

2. The Race to Find the Best Detective (The Models)

Once they had the training data, they needed to teach a computer to spot these badges automatically. They ran a "detective race" with three types of AI:

  • The Old School Detectives (Encoders): These are standard AI models that read text and classify it.
  • The New Super Detectives (Decoder LLMs): These are the latest, massive AI models (like the ones powering chatbots) that can generate text.
  • The Winner: The Hebrew-tuned Super Detective (specifically a model called DICTALM2.0) won the race. It was like giving the detective a native Hebrew speaker's intuition. It understood the cultural nuances and political slang better than the older models, achieving a success rate of about 74% (which is very high for this complex task).

3. Putting the Detective to Work (The Findings)

Once the AI was trained, the researchers used it to analyze three different "rooms" in the political house:

  1. Facebook: Where politicians talk to voters directly.
  2. The Knesset (Parliament): Where they give formal speeches to each other.
  3. The Public: What regular people say in surveys.

Here are the cool things they discovered:

  • The "Election Bump": Just like a store puts up "SALE" signs before a holiday, politicians suddenly start wearing more identity badges right before elections. The AI spotted that identity talk spikes dramatically during election seasons.
  • The "Bundle" Effect: Some badges always come together. If a politician wears the "Rightist" badge, they are very likely to also be wearing the "Security" and "Zionist" badges. It's like a uniform; you don't see the "Leftist" badge paired with the "Ultra-Orthodox" badge very often.
  • The Gender Gap: The AI noticed that men and women politicians wear different badges. Men tended to talk more about "Security" and "Capitalism," while women were much more likely to wear the "Socially-Oriented" (caring about welfare and community) badge.
  • The "Two Worlds" Problem: Sometimes what politicians say and what the public feels don't match. For example, the public cares a lot about "Honesty" and "Democracy," but politicians on Facebook talk about those less than they talk about "Rightist" or "Security" issues. However, in the Parliament, politicians talk more about social issues than they do on Facebook.

Why Does This Matter?

Think of HEBID as a new pair of glasses for social scientists.

  • Before: They could only see the "big picture" (e.g., "This party is Left or Right").
  • Now: They can see the "fine details" (e.g., "This politician is trying to appeal to religious voters and security voters simultaneously").

This tool isn't just for Hebrew. It shows researchers how to build similar tools for other languages that aren't English, helping us understand how people in different cultures use language to signal who they are and who they want to vote for.

In short: HEBID is a smart, culturally-aware AI that finally learned how to read the "invisible badges" of Israeli politics, revealing how politicians try to connect with different groups of people in ways we couldn't measure before.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →