← Latest papers
💬 NLP

Benchmarking Vision-Language Models for French PDF-to-Markdown Conversion

This paper introduces a French-focused benchmark for evaluating Vision-Language Models on challenging PDF-to-Markdown conversion tasks, utilizing model-disagreement sampling and specialized normalization metrics to demonstrate that while proprietary models excel on complex handwritten and form-based documents, several open-weights systems remain competitive for standard printed layouts.

Original authors: Bruno Rigal, Victor Dupriez, Alexis Mignon, Ronan Le Hy, Nicolas Mery

Published 2026-02-13
📖 5 min read🧠 Deep dive

Original authors: Bruno Rigal, Victor Dupriez, Alexis Mignon, Ronan Le Hy, Nicolas Mery

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a massive library of old French documents. Some are crisp, modern reports; others are handwritten notes from the 19th century, forms filled out in messy cursive, or pages with tiny text and complex charts. Your goal? To turn all these messy images into clean, organized digital text (Markdown) that a computer can easily read and use to answer questions.

This paper is essentially a report card for the smartest AI computers (called Vision-Language Models) trying to do this job. The authors, working for French postal and data companies, wanted to see which AI is the best "digital librarian."

Here is the breakdown of their study, explained with some everyday analogies:

1. The Problem: The "Perfect Score" Trap

Imagine you are grading a student's essay. If you use a strict computer program to compare the student's essay word-for-word against the teacher's answer key, you might fail the student just because they used a comma instead of a semicolon, or because they broke a line in a slightly different spot.

The authors realized that existing tests for AI were doing exactly this. They were too obsessed with exact spelling and formatting rather than meaning.

  • The Old Way: "You missed a comma here, and you put the table in a different order. You get a zero."
  • The New Way (This Paper): "Did you get the main ideas right? Is the table readable? Did you miss any important facts? If yes, you pass, even if the formatting is slightly different."

They built a new test that cares about semantic truth (did you get the meaning?) rather than string perfection (did you type every character exactly right?).

2. The Test: The "Adversarial" Challenge

To make sure they were testing the AI's limits, they didn't just pick random pages. They used a clever trick:

  • They asked two different AIs to transcribe the same page.
  • If the two AIs disagreed wildly, that page was marked as "Hard."
  • They collected 60,000 documents and cherry-picked the messiest ones: tiny text, handwritten forms, dense tables, and pages with lots of graphics.

Think of it like a driving test where they don't just test you on an empty parking lot; they send you into a blizzard with a broken windshield to see if you can actually drive.

3. The Contestants: The "Big Tech" vs. The "Open Source"

They put 15 different AI models through the gauntlet.

  • The Big Tech Giants (Proprietary): These are the expensive, closed-source models from companies like Google (Gemini) and OpenAI.
  • The Community Heroes (Open-Weights): These are free models that researchers and developers built and shared.

The Results:

  • The Handwriting & Forms Challenge: This was the hardest part. The Google Gemini 3 Pro was the clear winner, like a seasoned detective who can read a messy, 100-year-old handwritten letter. Many of the free models completely failed here, getting confused by the cursive.
  • The Standard Text Challenge: For clean, printed documents, several of the free models performed almost as well as the expensive ones. They are great for everyday tasks.
  • The Graphics Challenge: Most AIs struggled to describe charts and scientific figures. They often ignored them or got the data wrong. One model, Chandra, stood out for handling these better, but it was also the slowest.

4. The Speed vs. Accuracy Trade-off

There is a classic rule in the tech world: Fast and Cheap, or Good and Slow.

  • The Speedsters: Some models (like Docling and MinerU) were incredibly fast, processing a page in under a second. But they often missed details or got confused by complex layouts.
  • The Slow-and-Steady: The best models (like Gemini 3 Pro and Chandra) took longer (a few seconds per page) but produced much higher quality results.
  • The "Glitch" Factor: The authors noticed that when the image quality was low (tiny text), some smaller models would get "stuck" in a loop, repeating the same sentence over and over, or skipping entire sections of the page.

5. The Takeaway

The paper concludes that:

  1. If you need to read messy handwriting or complex forms: You currently need the expensive, top-tier models (like Gemini). The free ones just aren't there yet.
  2. If you are dealing with clean, printed text: You can use the free, open-source models to save money.
  3. The Test is a Living Thing: The authors aren't just publishing a score; they are inviting the whole community to help improve the test, add new types of documents, and keep the leaderboard up to date.

In a nutshell: This paper is a guide for anyone trying to turn French PDFs into data. It tells you that while AI is getting amazing at reading documents, it still struggles with the "human messiness" of handwriting and complex layouts, and you have to choose your tool based on whether you need speed or perfection.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →