← Latest papers
💬 NLP

Darshana Graph: A Parallel Commentary Corpus for Comparative Indian Philosophy, with Stylometric and Exploratory Graph Analyses

This paper introduces Darshana Graph, a unique parallel commentary corpus of over 125,000 records from Hindu, Buddhist, and Jain traditions that enables direct cross-commentator comparison, and demonstrates its utility through stylometric analyses of argumentative styles and a constrained LLM pipeline for extracting and validating philosophical relationships.

Original authors: Joy Bose

Published 2026-06-17
📖 5 min read🧠 Deep dive

Original authors: Joy Bose

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a massive, ancient library where the same few foundational books have been read and argued over for a thousand years. In this library, different groups of scholars (Hindus, Buddhists, and Jains) have written commentaries on the exact same verses. Usually, if you go to a digital library today, you pick up one book and read one scholar's take on it. You never see the other scholars' arguments side-by-side with the original text.

Darshana Graph is a new digital project that changes this. It's like building a giant, organized spreadsheet where every single verse from these ancient texts is a row, and the columns are filled with the different interpretations from eighteen different historical scholars.

Here is a simple breakdown of what the researchers did and what they found:

1. The Big Collection (The Corpus)

The researchers gathered over 125,000 text records.

  • The "Elephant in the Room": Most of these records (about 92%) come from the Buddhist Pali Canon. Think of this as the massive foundation of the building.
  • The "Special Sauce": The real magic lies in the remaining 8,500 records. These are Hindu and Jain texts where the same original verse is lined up next to commentaries from different schools of thought (like Advaita, Dvaita, and Jainism).
  • Why it matters: It's like having a debate transcript where everyone is quoting the exact same line from the rulebook, but explaining what it means in completely different ways. This alignment has never been done at this scale before.

2. Analysis One: The "Writing Style" Detective Work

The researchers didn't use complex AI to guess meanings. Instead, they used simple math to measure how the scholars argued, not just what they argued. They looked at things like:

  • How many times did they quote the original scripture?
  • How often did they say, "My opponent is wrong"?
  • How long were their sentences?

What they found:

  • The Trade-off: They noticed a pattern: Scholars who quoted the original text a lot tended to argue against opponents less. Conversely, those who argued against opponents a lot quoted the text less. It's like two different strategies: "I win because the book says so" vs. "I win because your idea is flawed."
  • The "Refutation" Trend: In one specific family of thought (the Dvaita school), they saw a funny trend over time. The founder wrote very little arguing against others. But as time went on, later scholars in that same line wrote much more arguing against others. It's as if the debate got more heated and defensive as the generations passed.
  • Buddhist Texts: Even within just the Buddhist texts, they found that some collections were written like short, punchy slogans (like the Dhammapada), while others were long, flowing essays.

3. Analysis Two: The AI "Relationship Mapper"

For the second part, they used a small, smart computer program (a Large Language Model) to act like a librarian trying to map out the relationships between ideas.

  • The Rules: To keep the AI from making things up (hallucinating), they gave it a strict, pre-approved list of relationship types it could use, like "IS THE SAME AS" or "IS DIFFERENT FROM." If the AI tried to use a word not on the list, the system automatically threw it out.
  • The Result: They built a "knowledge graph" (a web of connected ideas).
  • The Big Discovery: The AI correctly identified the biggest historical fight in Indian philosophy: Is the human soul (Atman) the same as the ultimate reality (Brahman)?
    • One school said: "Yes, they are identical."
    • Another school said: "No, they are totally different."
    • The AI mapped these disagreements perfectly, showing exactly where the schools disagreed based on the text.

4. The Honest "Fine Print" (Limitations)

The authors are very open about where their project isn't perfect yet:

  • The AI isn't perfect: Sometimes the AI got confused and used a generic label ("IS A QUALIFIED ASPECT OF") too often when it wasn't sure.
  • Missing Pieces: They couldn't distinguish between two major branches of Jainism in their data, and they didn't have enough data to compare different scholars within the same school using their new AI methods.
  • No "Truth" Guarantee: They didn't hire a team of humans to check every single AI result. They did a quick spot-check, but they admit the graph needs more human verification before we trust it 100%.

The Bottom Line

This paper introduces a massive new tool (Darshana Graph) that lets us see ancient philosophical debates side-by-side for the first time. They used simple math to show how different scholars argued, and they used a carefully controlled AI to map out where different schools of thought disagreed.

They aren't claiming to have solved all the mysteries of Indian philosophy. Instead, they are saying: "Here is the data, here is what we found, and here are the mistakes we made. We hope this helps others dig deeper." They have released all their data and code for anyone to use.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →