← Latest papers
💬 NLP

Ebisu: Benchmarking Large Language Models in Japanese Finance

This paper introduces Ebisu, a novel benchmark comprising two expert-annotated tasks designed to evaluate the challenges large language models face in understanding the complex linguistic and cultural nuances of Japanese finance, revealing that even state-of-the-art models struggle with implicit commitments and hierarchical terminology extraction despite scale or adaptation.

Original authors: Xueqing Peng, Ruoyu Xiang, Fan Zhang, Mingzi Song, Mingyang Jiang, Yan Wang, Lingfei Qian, Taiki Hara, Yuqing Guo, Jimin Huang, Junichi Tsujii, Sophia Ananiadou

Published 2026-02-03
📖 5 min read🧠 Deep dive

Original authors: Xueqing Peng, Ruoyu Xiang, Fan Zhang, Mingzi Song, Mingyang Jiang, Yan Wang, Lingfei Qian, Taiki Hara, Yuqing Guo, Jimin Huang, Junichi Tsujii, Sophia Ananiadou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to understand a very polite, highly complex conversation between a Japanese company and its investors. In this conversation, the company rarely says "No" directly. Instead, they might say, "We will consider the possibilities," or "It is difficult to give a definitive answer right now." To a native speaker, these phrases are clear signals of a refusal. But to an AI trained mostly on English, where people tend to be more direct, these phrases look like a "Maybe" or even a "Yes."

This paper, titled EBISU, introduces a new "exam" designed to test how well Artificial Intelligence (AI) can understand these subtle, culturally specific financial conversations in Japan.

Here is a breakdown of the paper using simple analogies:

1. The Problem: The "Lost in Translation" Trap

The authors argue that current AI models are like tourists who have studied a Japanese textbook but have never actually visited the country.

  • The Language Hurdle: Japanese is a "head-final" language (the verb comes at the very end of the sentence) and uses three different writing systems mixed together. It's like trying to read a sentence where the most important word is hidden at the very end, and the words are written in three different fonts simultaneously.
  • The Cultural Hurdle: In Japanese business culture, being too direct is often considered rude. Companies use "high-context" communication, meaning the real meaning is hidden in the tone, the silence, or the specific way a sentence ends, rather than in the explicit words.
  • The Result: Even the smartest AI models get "lost" in these financial documents. They miss the subtle "No" and the complex financial terms because they are looking for direct English-style answers.

2. The Solution: The EBISU Benchmark

To fix this, the researchers created EBISU, a specialized test with two distinct challenges (tasks). Think of it as a two-part driving test for AI.

  • Task 1: The "Reading Between the Lines" Test (JF-ICR)

    • The Scenario: An investor asks a tough question, and the company gives a polite, vague answer.
    • The Challenge: The AI must decide: Is the company saying "Yes, we will do it," "Maybe, but it's risky," or "No, absolutely not"?
    • The Trap: The AI often mistakes a polite "No" for a "Maybe" because it doesn't understand the cultural rules of indirect refusal.
    • The Data: They used real transcripts from 4 major Japanese companies over 3 years, carefully labeled by human experts who know the difference between a polite hedge and a hard refusal.
  • Task 2: The "Finding the Needle in a Haystack" Test (JF-TE)

    • The Scenario: Japanese financial reports are like dense forests of text. Important financial terms are often buried in long, nested phrases and mixed with loanwords from English and Chinese that have changed meaning over time.
    • The Challenge: The AI must find specific financial terms and rank them by importance.
    • The Trap: Because the words are mixed (Kanji, Hiragana, Katakana) and nested inside each other, the AI often gets confused about where one term ends and another begins. It's like trying to find specific ingredients in a soup where the labels are written in three different languages and the ingredients are glued together.

3. The Results: The AI Struggles

The researchers tested 22 different AI models, including the most famous ones from the US (like GPT-4 and Claude) and models specifically trained on Japanese or financial data.

  • The Scorecard: The results were surprisingly poor. Even the "smartest" AI models only got about 40% of the answers right.
  • Bigger isn't Better: Making the AI "bigger" (adding more data and parameters) helped a little bit, but not enough to solve the problem.
  • Specialization Didn't Help: Surprisingly, models that were specifically trained on Japanese financial data performed worse than general models in some cases. It's as if studying a specific textbook made the student overconfident and less flexible.
  • The Bias: AI models trained on English data tend to be too optimistic. When a Japanese company is being vague, the English-trained AI assumes they are agreeing, when they are actually refusing.

4. The Conclusion

The paper concludes that we cannot just "scale up" AI or translate English financial rules to Japanese. To truly understand Japanese finance, AI needs to learn the cultural nuances and linguistic quirks of the language, not just the vocabulary.

In short: The EBISU benchmark is a reality check. It shows that while AI is great at many things, it is currently terrible at understanding the polite, indirect, and complex way Japanese companies talk about money. The authors have released their test data and tools so other researchers can try to build better "cultural translators" for finance.

Note: The authors explicitly state that these models are for research and evaluation only. They warn that even if an AI scores well on this test, it should not be used to make real investment decisions or for legal compliance without human experts double-checking everything.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →