← Latest papers
💬 NLP

A Geometric Profile of Semantic Information in Text: Frame-Conditional Uniqueness and a Trade-Off Triangle for Scalar Summaries

This paper introduces a geometric framework for measuring semantic information in text via a three-coordinate profile of novelty, breadth, and integration, proving a fundamental trade-off that prevents any single scalar summary from satisfying all desirable properties while demonstrating that a rank-normalized configuration effectively outperforms existing baselines.

Original authors: Dmitriy Kompaneets

Published 2026-06-11
📖 6 min read🧠 Deep dive

Original authors: Dmitriy Kompaneets

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant library of books, and you want to answer a simple question: "How much meaning is actually in this specific text?"

For decades, scientists have tried to answer this using math that counts how unpredictable the words are (like guessing the next letter in a word). But as this paper points out, that's like measuring a painting by counting how many different shades of blue are on the canvas. It tells you about the colors, but not the picture. Two paintings can have the exact same mix of colors but tell completely different stories.

This paper proposes a new way to measure meaning by looking at the "shape" of the ideas inside a text, using a tool called sentence embeddings (which turn sentences into points in a giant, multi-dimensional map).

Here is the core of their discovery, broken down into simple concepts:

1. The "Shape" of Meaning (The Profile)

Instead of trying to squeeze a whole book's meaning into a single number (like a test score), the authors say we need a three-part profile. Think of it like describing a person not just by their height, but by three distinct traits:

  • Novelty (The "Newness"): How far does this text wander from the "average" boring conversation? If you write a text that sounds exactly like a generic news template, its novelty is low. If you write something surprising and unique, it moves far away from the center of the map.
  • Breadth (The "Spread"): How many different distinct ideas are in there? Imagine a text that talks about cats, then rockets, then baking, then quantum physics. That has high breadth. A text that talks about cats for 10 paragraphs has low breadth.
  • Integration (The "Glue"): How well do those ideas stick together? A text that jumps randomly from cats to rockets to baking has low integration (it's a messy bag of facts). A text that weaves those topics together into a coherent story has high integration.

The Analogy:

  • Poetry is like a tight, beautiful knot. It has high novelty (unique words) and high integration (everything fits), but low breadth (it stays on one theme).
  • A chaotic dialogue is like a scattered pile of toys. It has high breadth (lots of different topics) but low integration (nothing connects).
  • Legal text is like a wide, flat net. It covers a lot of ground (high breadth) and is very structured (high integration), but it's not very "new" (low novelty).

2. The "Semantic Quantum" (The Pixel of Meaning)

The paper introduces a concept called the Semantic Quantum. Think of this as the "pixel" of meaning.

  • If you have a text with 100 sentences, but 90 of them are just rephrasing the same idea, the text isn't actually 100 ideas deep.
  • The authors use a "clustering" tool to group similar sentences together. If 10 sentences are almost identical, they count as one "quantum" (one idea).
  • This prevents the math from being tricked by repetition. It's like counting the number of unique ingredients in a soup, rather than counting every single grain of salt.

3. The "No-Go" Triangle (The Impossible Dream)

Here is the paper's most important warning. The authors tried to combine those three traits (Novelty, Breadth, Integration) back into a single number (a "Scalar") to make it easy to use.

They proved mathematically that you cannot have a perfect single number. It's a "Trade-Off Triangle." You can only pick two of the following three superpowers for your number:

  1. Stability: If I rewrite the text slightly (paraphrase), the number stays the same.
  2. Ranking Power: If I compare two different texts, the number correctly tells me which one is "richer."
  3. Cross-Model Fairness: If I use a different AI tool to measure the text, the number stays comparable.

The Catch:

  • If you make a number that is stable (doesn't change with small edits), it becomes bad at ranking different texts.
  • If you make a number that is great at ranking, it becomes unstable (small changes make the score jump wildly).
  • If you try to make it fair across different AI tools, you lose the other two.

Because of this, the paper suggests we stop looking for one "magic number" and instead use two different tools for two different jobs:

  • Tool A (Sminmax): Good for theoretical math and checking if a summary is stable.
  • Tool B (Srank): Good for ranking documents (like a search engine) where you just need to know "which one is better," even if the exact score jumps around a bit.

4. The "Volume" Discovery

The authors found something cool about the "Breadth" part of the profile. They discovered that the "spread" of ideas in a text is mathematically identical to a concept from physics called a Determinantal Point Process (DPP).

  • Simple Analogy: Imagine you are picking a team for a project. You don't want three people who all think exactly the same way. You want a team where everyone brings a different perspective, but they still work well together.
  • The math shows that the "Breadth" of a text is essentially the volume of the best possible diverse team you could pick from that text's sentences. This gives the "Breadth" score a solid mathematical foundation, proving it's not just a guess.

Summary

The paper argues that meaning is too complex to be a single number.

  • It gives us a 3D map (Novelty, Breadth, Integration) to see the shape of a text.
  • It warns us that trying to flatten this map into one score always involves a compromise (the Trade-Off Triangle).
  • It provides two practical tools (one for stability, one for ranking) so we can measure text meaning in a way that actually makes sense for the job we are doing.

The authors tested this on 5 classic novels (like Moby Dick and Pride and Prejudice), 23 types of synthetic text, and various AI models, and found that this 3D approach works much better than previous methods at telling the difference between a coherent story and a random bag of facts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →