fog: Expressing Motion and Emotion through Function Composition of AI-Generated Code
This paper introduces "fog," a function composition framework that leverages AI-generated code and an interactive animation editor to enable users to create expressive, Heider-Simmel-style motion and emotion, which was validated through perceptual studies showing 68% semantic recognition accuracy and improved user control.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're trying to tell a story using nothing but floating shapes on a screen—a triangle, a circle, a square. You want the circle to look angry as it chases the triangle, or the triangle to look scared as it runs away. Sounds simple, right? But ask an AI to "make the circle look angry," and you might get a circle that just spins in place or moves too fast, missing the feeling entirely.
That's the problem fog (a new tool by researchers Vivian Liu and Lydia Chilton) is trying to solve. Instead of just shouting vague instructions at a computer, fog treats motion and emotion like a giant, mix-and-match LEGO set built from code.
The Big Idea: Motion is a Recipe, Not a Magic Spell
The researchers suggest that motion and emotion aren't just random chaos; they are made of specific ingredients. Think of it like cooking. If you want a spicy dish, you don't just tell the chef, "Make it spicy." You add specific amounts of chili, pepper, and heat.
In fog, the "ingredients" are Verbs (like chase or flee), Adverbs (like hesitantly or aggressively), Gestures (like waving or shaking), and Emotions (like anger or fear). The magic happens when you snap these code-blocks together.
For example, you can tell the AI to make a shape hesitantly(chase) another shape. The system doesn't just guess; it builds a specific recipe where the shape moves forward but stops and starts nervously, maybe shaking a little. Or, you could combine anger with a lunge, creating a motion that builds up energy, strikes hard, and then recoils.
The "Magic" Test: Can People Actually See the Feeling?
The team didn't just build this and hope for the best. They ran a massive taste test with 452 different animations. They showed these moving shapes to hundreds of people and asked, "What is this shape doing? Is it chasing, avoiding, or fighting?"
The results were pretty cool. People were able to guess the correct meaning 68% of the time. That might sound like a C+ in school, but remember, if they were just guessing randomly among four options, they would only get it right 25% of the time. So, fog's animations were 2.68 times better than a lucky guess.
However, the paper is careful to point out that this isn't a perfect solution. When they tried to mix two things together—like making a shape hesitantly and chase—the accuracy dropped a bit (to about 58%). Sometimes the "ingredients" clashed, like trying to mix oil and water. The researchers found that if the code for "hesitant" and the code for "chase" tried to control the same part of the movement, they sometimes canceled each other out. It's a reminder that while the system is powerful, it's not yet a mind-reader.
The "Playground" for Creators
The researchers also built a playground (an animation editor) to see how real people would use this. They invited 10 people—some who know animation like the back of their hand, and some who have never coded a thing.
They compared fog to a standard "just type what you want" AI tool. The difference was huge.
- The Old Way: Users had to type a prompt, wait for the AI to generate a whole new scene, and then wait again if they didn't like it. It took about 1.75 minutes just to see a new version of the animation.
- The fog Way: Users could tweak the motion instantly. They could drag a slider to make a shape move faster, draw a path for it to follow, or mix and match emotions on the fly. The time to see a new version dropped to just 0.36 minutes.
The pros loved the control, saying it felt like having "atomic" building blocks to fine-tune every detail. The beginners loved the safety net; instead of staring at a blank screen wondering what to type, they could pick from a grid of pre-made "verbs" and "adverbs" to get started. One beginner even managed to animate a circle that looked anxious and then calmed down, saying, "I can see the emotion in the circles."
What It's NOT (And What It Can't Do Yet)
It's important to know what this paper doesn't claim.
- It's not a movie maker: The animations are still just simple shapes (circles and triangles) moving around. The paper explicitly says that while 68% accuracy is great, you can't expect 100% perfection. A simple circle just can't express every possible human emotion or complex story.
- It's not a "fix-all" for AI: The researchers tried to have the AI "self-refine" its own code to fix the clashing ingredients, but that didn't really help much. The accuracy barely moved.
- It's not for complex robots (yet): The system works best for these abstract shape stories. The paper notes that if you want to animate a real robot with joints and lighting, or a character with a face, fog isn't ready for that level of detail just yet.
The Takeaway
fog suggests that we can teach computers to understand emotion by breaking it down into code functions, like a recipe book for feelings. It's not a perfect, magic wand that solves every animation problem, but it's a powerful new way to play. It turns the scary, vague process of "telling an AI what to do" into a fun, hands-on game of mixing and matching motion ingredients, letting both experts and beginners tell stories with shapes that actually feel alive.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.