← Latest papers
💻 computer science

ReART: Reference-Guided Retrieval and Refinement for Emotion-Aware Art Generation

ReART is a reference-guided retrieval and refinement framework that enhances emotion-aware art generation by decomposing captions into structured visual fields for targeted reference retrieval and employing an AAS-driven loop to diagnose and repair alignment failures, achieving top-tier performance in the AffectiveArt 2026 Grand Challenge.

Original authors: Qianqian Tang, Jiayi Gao, Ting Lei, Yang Liu

Published 2026-08-25
📖 4 min read☕ Coffee break read

Original authors: Qianqian Tang, Jiayi Gao, Ting Lei, Yang Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Art has always been more than a picture of a thing; it is a way of feeling. For centuries, painters have used brushstrokes, color choices, and the way light falls on a scene to convey emotions like solemnity, joy, or melancholy, often without painting a single smiling face or a tear. Today, computers are learning to create these images from written descriptions, a process that has become startlingly good at producing realistic scenes. However, a new challenge has emerged: teaching a machine to understand that a "solemn" atmosphere requires specific, subtle visual decisions, such as muted colors or a flattened sense of space, rather than just adding a sad object to the picture. The difficulty lies in the fact that human language is often too vague to describe these fine details. When a person writes a prompt asking for a "solemn Renaissance painting," the computer does not know exactly which shades of gray to use or how thick the brushstrokes should be to match that feeling. It is like asking someone to paint a mood without giving them a specific palette or a clear vision of the texture.

To solve this, researchers at Peking University and Wuhan University developed a system called ReART, which acts as a bridge between vague emotional words and concrete visual evidence. Instead of trying to guess the right look from text alone, the system breaks down a description into specific visual parts, such as the main subject, the layout of the scene, the texture of the paint, and the overall mood. It then searches a large library of existing artworks to find real examples that match each of these specific parts. If the prompt asks for a certain type of brushwork, the system finds a painting with that exact texture and uses it as a guide. This allows the computer to see what "solemn" or "dramatic" actually looks like in the hands of a master artist, rather than just guessing based on words.

The process happens in two main steps. First, the system generates an initial image using the text description and the visual clues it gathered from the library. It then acts as its own critic, checking the result against the original request. If the image gets the subject right but the colors feel too bright, or if the composition is correct but the texture looks too smooth and digital, the system identifies exactly where it went wrong. It does not simply throw the image away and try again. Instead, it creates a specific plan to fix only the broken parts while protecting the parts that are already correct. It selects new visual examples that match the specific error and carefully edits the image to correct the flaw, ensuring the rest of the painting remains untouched. This cycle of checking and fixing continues until the image aligns perfectly with the emotional and visual goals of the prompt.

The results of this approach were tested in a major competition for emotional art generation, where the system was judged on how well it captured the content, the artistic style, and the fine-grained emotional attributes. The researchers found that their method produced images that were significantly closer to real human-made art than previous attempts. In the competition, the system achieved a perfect score on the emotional alignment test, meaning it successfully matched the intended feelings in every single case it was given. It also ranked second overall in the entire challenge, proving that breaking down a complex artistic request into smaller, manageable visual pieces and using real examples to guide the creation is a powerful way to teach machines the subtle language of art. The work demonstrates that when computers are given the right visual references and a way to check their own work, they can learn to create art that resonates with human emotion, not just by following rules, but by understanding the visual texture of feeling.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →