← Latest papers
💻 computer science

MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition

This paper introduces MarkushGlyph and OCSRGlyph, two image-to-text translation models that significantly improve chemical structure recognition by leveraging a vision-language approach for Markush structures and enhanced stereochemistry handling for single molecules, alongside a new metric for evaluating Markush translation accuracy.

Original authors: Alex Andonian, Samuel G Rodriques, Andrew D White, Siddharth M Narayanan

Published 2026-07-31
📖 3 min read☕ Coffee break read

Original authors: Alex Andonian, Samuel G Rodriques, Andrew D White, Siddharth M Narayanan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of chemistry as a massive, bustling library where every book, patent, and research paper is filled with intricate drawings of molecules. For human chemists, these drawings are like a secret language they can read instantly, seeing how atoms connect to form medicines, plastics, or fuels. But for computers, these images are just a jumble of pixels; a robot can't "see" a benzene ring the way a person does. To make these drawings useful for computers—so they can search databases, predict new drugs, or train artificial intelligence—we have to translate the pictures into a text code that machines can understand. This process is called Optical Chemical Structure Recognition (OCSR). It's like taking a sketch of a house and turning it into a precise set of blueprints written in a language a construction robot can read.

There are two main types of drawings in this library. The first is a single, specific molecule, like a photo of one exact car. The second is a "Markush" structure, which is more like a blueprint for a whole family of cars. Instead of drawing one specific vehicle, a Markush drawing shows a shared frame with empty spots labeled "R1" or "R2," meaning "any part that fits here." This is how patent lawyers describe families of potential medicines without listing every single variation. The challenge has always been that while computers are getting good at reading the single-car photos, they still struggle to understand the complex, flexible family blueprints.

This paper introduces two new AI tools, named OCSRGlyph and MarkushGlyph, designed to solve these translation problems. Think of them as a new generation of super-smart translators. OCSRGlyph is a specialist focused on the single-molecule photos. The researchers found that the AI kept making tiny mistakes with the "handedness" of the molecules (a property called stereochemistry, where a molecule can be a left-handed or right-handed version of itself, like a pair of gloves). To fix this, they fed the AI extra training examples of these tricky shapes. The result is a translator that gets the single-molecule code right 93.8% of the time, which is a new record for accuracy.

MarkushGlyph, on the other hand, is the more ambitious sibling. It's a "vision-language" model, meaning it looks at the entire patent page—including the drawing and the surrounding text—as one big picture, rather than trying to read the text and the image separately. Previous methods tried to solve this puzzle in stages, like a team where one person reads the text, another draws the graph, and a third tries to put them together, often losing details in the hand-off. MarkushGlyph skips the middleman and reads everything at once. The authors show that this single-step approach is much better at understanding the "family blueprints" than the old multi-stage methods. They also introduced a stricter way to grade the answers, catching errors that the old grading system missed, such as when the AI accidentally swapped two labels or added a feature that wasn't there.

In short, the paper demonstrates that by treating these chemical drawings as a direct image-to-text translation problem and using modern AI techniques, we can get computers to read chemical patents with unprecedented accuracy. The team didn't just build a better translator; they also built a better test to make sure the translation is truly correct, ensuring that when a computer reads a chemical family, it understands exactly what the chemist intended, down to the smallest detail.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →