Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance
The paper introduces TIGER, a novel two-stage framework that resolves the trade-off between image quality and text readability in scene text super-resolution by prioritizing glyph structure restoration to guide subsequent image enhancement, supported by the newly proposed UZ-ST dataset.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to read a street sign from the top of a skyscraper. The sign is there, but from so far away, it looks like a blurry, fuzzy smudge. In the world of computer vision, this is a classic puzzle called "Super-Resolution." It's the art of taking a tiny, grainy, low-quality picture and using math to guess what the big, sharp, high-definition version should look like. Usually, computers are pretty good at this when the picture is of nature—like turning a fuzzy photo of a grassy field into a crisp one. They can invent plausible details, like individual blades of grass or the texture of a leaf, because nature is full of patterns.
But text is a different beast entirely. A blurry letter isn't just a fuzzy blob; it's a specific code. If a computer guesses the wrong shape for a single stroke in a Chinese character, it might accidentally turn a word for "peace" into a word for "war," or make the whole thing look like gibberish. For years, scientists have been stuck in a frustrating tug-of-war: if they try to make the whole picture look pretty and sharp, the text often turns into nonsense. If they try to fix the text, the rest of the image looks weird and blocky. It's like trying to fix a broken clock by painting the whole wall; you might get a nice wall, but the time is still wrong.
This is where a new team of researchers steps in with a clever solution called TiGeSR. Instead of trying to fix the blurry picture all at once, they decided to split the job into two distinct steps, following a "text-first, image-later" philosophy. Think of it like a master calligrapher and a painter working together. First, the calligrapher ignores the messy background and focuses entirely on reconstructing the perfect, sharp outlines of the letters based on what they think the words say. Once those perfect letter shapes are drawn, the painter uses them as a strict guide to fill in the rest of the picture, ensuring the background looks natural but the letters stay exactly where they belong.
The team didn't just invent a new tool; they also realized that to teach a computer this skill, they needed a much harder test than anyone had ever used before. Existing tests were like looking at a sign from a few steps away. The researchers built a new dataset called UZ-ST, which simulates looking at signs from extreme distances—up to 14.29 times further away than the standard tests. It's like asking a student to read a menu not just from across the room, but from the other side of the street.
When they put their two-stage method to the test, the results were impressive. On these super-challenging, ultra-zoomed images, TiGeSR managed to restore text that other methods completely failed to read. While other AI models either turned the text into alien symbols or left the background looking like a patchwork quilt, TiGeSR kept the letters sharp and the image coherent. The researchers found that by explicitly separating the task of "fixing the letters" from "fixing the picture," they could achieve both goals at once, proving that you don't have to choose between readability and beauty—you can have both if you tackle them in the right order.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.