Through Van Gogh's Eyes: Global Style Transfer with Diffusion Mod
This paper introduces Global Style Transfer (GST), a novel artistic image synthesis paradigm that leverages Global Style Guidance and Content Alignment Guidance to aggregate multiple artworks from a target artist into a single content image, thereby overcoming the limitations of existing methods in capturing an artist's global style distribution while preserving semantic structure and reducing text-induced bias.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a time-traveling art critic who has just stepped into a room full of digital canvases. You want to know: if a famous painter like Van Gogh were alive today, how would they paint a modern city street or a cat sitting on a fence? This is the heart of "artistic image synthesis," a field where computers try to mimic the unique visual fingerprints of human artists. For a long time, computers had two main ways to do this. The first was like a strict photocopy machine: it would take one specific painting and force a new photo to look exactly like it, but it often missed the bigger picture of the artist's whole career. The second way was like asking a robot to "draw in the style of Van Gogh" using only words. While flexible, this often made the robot rely heavily on the most famous, iconic paintings (like The Starry Night) over and over again, ignoring the rest of the artist's diverse work. The big question scientists are trying to answer is: how can we teach a computer to understand an artist's entire soul—their whole range of styles, colors, and brushstrokes—without just copying one famous picture or relying on a text prompt that might trick the computer?
This paper introduces a clever new solution called Global Style Transfer (GST), which acts like a master art student who studies a whole library of an artist's work before picking up a brush. Instead of looking at just one painting or listening to a text command, the researchers' system gathers hundreds of artworks from a single artist (like Van Gogh, Monet, or Renoir) and learns the "global style" that ties them all together. They call this a "Many-to-One" approach because it takes many source paintings and blends them into one unified style guide for a single new image.
To make this work, the team built two special tools. The first is Global Style Guidance (GSG). Think of this as a secret decoder ring that lives inside the computer's brain. Instead of asking the computer to "be Van Gogh" with words, this tool looks at the visual patterns in hundreds of paintings and calculates a "style offset"—a mathematical nudge that tells the computer how to shift its colors and textures to match the artist's true, broad personality. Crucially, the researchers found that by training this tool with a boring, fixed phrase like "A painting" for every single artwork, they could strip away the influence of text. This means the computer learns purely from what it sees, not from what it reads, avoiding the bias where it only copies the most famous, text-associated works.
The second tool is Content Alignment Guidance (CAG). Imagine you are painting a portrait of your friend in the style of Van Gogh. You want the swirling, thick brushstrokes, but you still want it to look like your friend, not a random stranger. CAG is the safety net that holds the friend's face in place while letting the artist's style warp and twist the rest of the image. It uses a "training-free" method (meaning it doesn't need extra learning time) to constantly check that the new image keeps the same basic structure as the original photo, even while the style changes.
The results are quite promising. When the researchers tested their system on over 80,000 artworks from the WikiArt database, they found that GST produced images that felt much more like the artist's true, diverse self compared to older methods. While other techniques often got stuck on a single famous look or produced blurry, inconsistent results, GST managed to capture the full "spectrum" of the artist. For instance, when generating Van Gogh-style images, the system didn't just copy The Starry Night; it created a variety of swirling, textured scenes that felt authentic to his entire body of work. The data suggests that this method successfully avoids "memorizing" just a few famous patterns and instead learns a richer, more faithful representation of the artist's visual identity, all while keeping the original subject of the photo recognizable.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.