← Latest papers
💻 computer science

ScribbleEdit: Synthetic Data for Image Editing with Scribbles and Text

This paper introduces ScribbleEdit, a large-scale synthetic dataset that combines natural language instructions with freehand scribbles to train and improve multimodal image editing models, enabling them to achieve precise spatial and semantic control that existing off-the-shelf models struggle to attain.

Original authors: Anya Ji, George Ma, Téa Wright, Yiming Zhang, David M. Chan, Alane Suhr, Somayeh Sojoudi

Published 2026-05-06
📖 4 min read☕ Coffee break read

Original authors: Anya Ji, George Ma, Téa Wright, Yiming Zhang, David M. Chan, Alane Suhr, Somayeh Sojoudi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to give instructions to a very talented but slightly confused artist. You want to change a photo, but you have two ways to talk to them:

  1. The Texty Way: You say, "Put a red apple in the middle." The artist understands "red apple" perfectly, but they might put it on the left, the right, or make it huge. They lack your specific vision of where it goes.
  2. The Scribble Way: You grab a pen and draw a messy, shaky circle on the photo where you want the apple. The artist knows exactly where to put it, but they have no idea if it should be red, green, shiny, or fuzzy.

The Problem: Current AI image editors are great at one or the other, but they struggle when you try to use both at the same time. They get confused by the messy scribble combined with the text, mostly because they haven't been trained on enough examples of this specific combination.

The Solution: ScribbleEdit
The authors of this paper built a massive "training gym" for AI called ScribbleEdit. Think of it as a giant library of practice exercises where the AI learns how to listen to both the text and the scribble simultaneously.

Here is how they built this library:

  • The Setup: They took thousands of photos of objects (like apples, cats, or cars).
  • The Eraser: They used a smart "eraser" tool to remove the object, leaving a blank spot in the background.
  • The Guide: They took the original photo and paired it with a messy, hand-drawn scribble (like a child's drawing) that roughly outlined where the object should go.
  • The Script: They used another AI to write a sentence describing the object (e.g., "Place a shiny red apple here").
  • The Result: They created over 11,000 examples where the AI had to look at the blank spot, the messy scribble, and the sentence, and then "paint" the object back in exactly where the scribble pointed, looking exactly like the text described.

The Experiment
The researchers took two different types of AI artists and put them through a test:

  1. The "Off-the-Shelf" Artist: These are pre-trained models that have never seen ScribbleEdit.
  2. The "Trained" Artist: These are the same models, but after they studied the ScribbleEdit library.

The Results

  • Before Training: The untrained models were terrible at this. When given a scribble, they often ignored it, drew weird shapes that looked like scribbles, or just failed to edit the image at all. They didn't understand the "language" of a messy line.
  • After Training: Once the models studied the ScribbleEdit dataset, they got much better. They learned to:
    • Place the object exactly where the scribble indicated (spatial accuracy).
    • Make the object look like the text described (semantic accuracy).
    • Keep the rest of the photo looking natural.

The Takeaway
The paper shows that while AI is getting good at editing photos, it still struggles to understand our messy, human way of drawing. By creating a synthetic dataset that teaches AI how to combine "rough drawings" with "clear instructions," the researchers proved that AI can learn to be a much more precise and obedient editor.

What They Didn't Do (Important Limitations)
The paper is very specific about what they did not do:

  • They didn't invent a new way to draw; they used existing human drawings from a database called "Sketchy."
  • They only tested "adding an object back in" (object addition). They didn't test complex things like changing a person's face or editing only a tiny part of a shirt.
  • They didn't claim this works for medical imaging or other specialized fields; it's strictly about general photo editing.

In short, ScribbleEdit is a new training manual that teaches AI how to listen to our messy doodles and clear words at the same time, making photo editing feel more like talking to a helpful human assistant and less like guessing a riddle.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →