Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces
This paper audits six vision-language models to evaluate their consistency in ranking 3D objects by affective Kansei descriptors, revealing that while models show partial agreement above chance levels, their convergence varies significantly by object category and semantic alignment, prompting a UI prototype to guide the selection of reliable semantic controls for generative design interfaces.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the early stages of design, whether sketching a new chair or modeling a car, creators often rely on words to guide their tools. They might ask for something "more elegant," "less bulky," or "more minimalist." Today, artificial intelligence systems are beginning to understand these requests, using powerful computer programs that can see images and read text to generate new shapes based on such descriptions. These systems, known as vision-language models, act as a bridge between a designer's vague idea and a concrete 3D object. However, a critical question remains: do these different computer programs actually agree on what those words mean? If one program thinks a tall, thin bottle is "elegant" while another thinks a short, wide one is, the tool becomes unreliable. Designers need to know if the digital assistant is consistent before they trust it to shape their work.
A team of researchers from Honda Research Institute Europe and Julius-Maximilians-Universität Würzburg set out to answer this question by auditing six of the most advanced vision-language models available. They wanted to see if these independent programs, built by different companies and trained on different data, would rank the same objects in the same way when asked to judge their emotional or "affective" qualities. To do this, they stripped away everything that could distract the computers. They took thousands of 3D models of everyday items like chairs, lamps, and jars and rendered them as plain, untextured grey shapes. By removing color, material, and texture, the researchers ensured the models were judging the objects based on form alone, much like a sculptor looking at a clay model before painting it.
The researchers then asked each of the six computer models to rank these grey shapes along various scales. Some scales were straightforward physical descriptions, such as "tall versus short" or "boxy versus curvy," which served as a control to see if the models could even agree on basic geometry. Others were the more complex, human-like descriptors the designers actually use, such as "modern versus traditional" or "elegant versus messy." Finally, they included a set of nonsense pairings, like "loud versus quiet," to establish a baseline for random agreement. If the models agreed on these irrelevant pairs, it would mean their rankings were just noise.
The results revealed a nuanced picture of agreement. The models were remarkably consistent when judging simple physical traits; they largely agreed on which objects were tall or short. When it came to the more abstract, emotional qualities, the models showed a moderate level of agreement that was significantly better than random chance, but not perfect. On average, the models agreed on the ranking of objects for emotional terms about 36 percent of the time, compared to a 14 percent agreement rate for the nonsense terms. This suggests that while the computers have developed a shared sense of how shape relates to feeling, they are not yet perfectly aligned.
Crucially, the study found that this agreement is not uniform across all types of objects. For some categories, like jars, the models agreed strongly on what looked "elegant" or "luxurious," with agreement rates reaching over 50 percent. For other categories, like bookshelves, the models struggled to find a common ground, with agreement rates dropping as low as 21 percent. The researchers discovered that this variation depends less on how different the objects look from one another and more on whether the specific shape variations in that category actually match the direction the word is pointing. For instance, if a group of chairs varies mostly in height, a model might easily agree on which is "tall," but if the group varies in subtle curves that don't align with the concept of "elegance," the models will disagree.
The researchers also tested whether these agreements were specific to the object type or just a general quirk of the text. They found that a descriptor like "plush versus rigid," which makes sense for sofas, did not work well when applied to bottles. This confirmed that the models are not just applying a generic rule but are actually interpreting the relationship between the specific shape of an object and the word used to describe it.
To make these findings useful for real-world design, the team built a prototype interface that visualizes this agreement. In this tool, a designer can select an object and see a list of potential controls, such as "make it more modern." Next to each control, the system displays a score indicating how much the different computer models agree on that instruction. If the score is high, the designer can be confident that the tool will behave predictably. If the score is low, the interface suggests that the control might be unreliable for that specific object and could be withheld or replaced with a better-supported alternative.
This work does not claim that the computers have learned human feelings or that their agreement proves they are right. Instead, it offers a practical method to check if a digital tool is stable before it is handed over to a user. By treating the agreement between different computer models as a form of quality control, designers can identify which emotional descriptors are safe to use and which might lead to confusion. The study suggests that while artificial intelligence is becoming a powerful partner in design, we must still carefully audit its understanding of the world, ensuring that when a designer asks for "elegance," the machine is looking at the same thing they are.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.