← Latest papers
🤖 AI

A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video

This paper presents a transparent, deterministic hybrid framework for extracting personal names from broadcast news videos that prioritizes auditability and hallucination-free traceability over the marginal accuracy gains of generative multimodal models.

Original authors: Andrea Filiberto Lucas, Dylan Seychell

Published 2026-08-04
📖 4 min read☕ Coffee break read

Original authors: Andrea Filiberto Lucas, Dylan Seychell

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a fast-paced news broadcast. The screen is a chaotic dance of moving pictures, flashing headlines, and colorful graphics. Suddenly, a name pops up on the screen to tell you who is speaking, but it's there for only a second, written in a tricky font, maybe with a weird symbol, and the camera is shaking. Your brain has to work overtime to catch it. This is the world of "visual information extraction," a branch of computer science where machines try to read and understand text that isn't just sitting in a book, but is painted onto moving video. The core idea is simple: teach computers to act like super-fast, super-attentive readers who can spot a name in a crowd of pixels, even when the crowd is moving and the lighting is bad. Why does this matter? Because we are drowning in video content. From TV news to TikTok clips, there is so much information flying by that humans can't possibly keep track of every name, place, or event mentioned. If we can't read the screen quickly, we miss the story.

Now, meet the team from the University of Malta who decided to build a machine that could solve this puzzle. They created a new tool called the "Accurate Name Extraction Pipeline" (ANEP). Think of ANEP as a very strict, very organized detective squad. Instead of guessing, this squad follows a rigid, step-by-step checklist: first, they find the graphic on the screen; second, they clean up the image to make the text clear; third, they read the text; and fourth, they use a rulebook to confirm if the text is actually a person's name. They built this system because they wanted something "deterministic," which is a fancy way of saying "predictable and auditable." If you ask the detective squad to do the same job twice, they will get the exact same result and can show you exactly how they found it, step by step.

To test their idea, the researchers created a massive library of 1,500 news frames called the "News Graphics Dataset" (NGD), capturing all the messy, different ways news channels show names. They trained their detective squad on this library and then pitted them against the new, flashy "Generative AI" models (like the ones you might have heard of that write stories or draw pictures). These new AI models are like creative artists; they look at the whole picture and guess what the name might be based on patterns they've learned. They are fast and often very good, but they are also a bit of a "black box"—you can't always see how they came up with an answer, and sometimes they might "hallucinate" (make up a name that isn't there).

The results were a fascinating showdown between the organized detective and the creative artist. The creative AI (specifically a model called Gemini 1.5 Pro) was the fastest and got the highest overall score for accuracy, with an F1 score of 84.18%. However, the paper argues that this speed comes with a cost: you can't trace its steps, and it might make mistakes you can't easily spot. The detective squad (ANEP) was slower, taking about 542 seconds to process the same video that the AI did in 94 seconds. Its accuracy score was slightly lower at 77.08%, but it had a crucial superpower: transparency. It never made up names, and if it made a mistake, you could look at its notes and see exactly where the process went wrong.

The researchers found that while the creative AI is impressive, the detective squad is better for situations where you need to be 100% sure of the facts, like in journalism or legal investigations. They also surveyed 413 people and found that 59% of them struggle to read names on fast news screens, proving that there is a real need for these tools. The paper concludes that while the flashy AI models are getting better, we still need the reliable, step-by-step systems to ensure that the information we extract from our screens is trustworthy and can be verified. It's not about which tool is "smarter," but about which tool you trust to tell the truth when the news is breaking.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →