More Accurate, Less Human: Gestalt Grouping in Vision Models
This paper introduces a behavioral battery that evaluates 45 vision models against human perceptual data on four Gestalt grouping tasks, revealing that alignment with human visual organization is a distinct and critical metric that conventional benchmarks often fail to capture, particularly for closed foundation models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Technical Summary: More Accurate, Less Human: Gestalt Grouping in Vision Models
Problem Statement
Vision-language models (VLMs) are increasingly deployed in visualization pipelines to interpret charts, evaluate designs, and act as proxies for human viewers. While these systems demonstrate high benchmark accuracy on task completion, it remains unverified whether they organize visual information according to the same perceptual regularities that govern human vision. Conventional evaluations measure capability (e.g., visualization literacy scores) but fail to distinguish between correct answers derived from human-like perceptual grouping and those derived from fundamentally different, non-human mechanisms. The paper argues that high accuracy does not establish "human-likeness" in perceptual organization, necessitating a shift from task-performance metrics to behavioral alignment with established cognitive psychology.
Methodology
The authors introduce a behavioral evaluation battery grounded in Gestalt psychology, specifically testing the principles of Closure (perceptual completion of incomplete structures) and Similarity (grouping by shared attributes). Instead of collecting new human data, the methodology reuses stimuli and published human behavioral data from prior perception studies as evaluation targets.
The battery consists of four tasks across two tracks (chart-based and natural-image controls):
- Silhouette Recognition (Closure): A 16-way forced-choice task on filled silhouettes where texture and interior details are removed. The human anchor is the modal response of ten observers.
- Mark-Color Odd-One-Out (Similarity): Identifying the outlier among marks based on color, scored against a crowdsourced perceptual kernel.
- Color-Series Counting (Similarity): Counting distinct color series in a chart, scored against the published human capacity band (6–12 categories).
- Object Odd-One-Out (Similarity): Identifying the semantic outlier in natural images, scored against the THINGS dataset's test-retest ceiling.
The study evaluates 45 models across five training families: supervised encoders, self-supervised encoders, contrastive vision–language encoders, open-weight VLMs, and closed foundation models.
Metrics
Three complementary metrics quantify behavioral alignment:
- Behavioral Agreement (): The rate at which a model's response matches the human target (modal choice or capacity band). This measures human-likeness, not just skill.
- Behavioral Error Consistency (): A measure of whether the model makes errors on the same specific items that humans do, beyond what would be expected by chance given their respective accuracies. This is compared against a human–human consistency ceiling.
- Behavioral Effect Replication (): Whether the model's accuracy profile across varying series counts mimics the shape of the human capacity curve.
Key Results
- Dissociation of Accuracy and Human-Likeness: High benchmark accuracy does not guarantee alignment with human perceptual organization. Models often trade places in rankings depending on the specific grouping task (e.g., a top performer in color similarity may be average in semantic grouping).
- Error Consistency Variance: While 34 of 45 models exceeded human accuracy, only four achieved error consistency () at or beyond the human–human ceiling. Many models, including some closed foundation models, exhibit high accuracy but fail to replicate human error patterns.
- Family-Specific Properties:
- Open-Weight VLMs: Generally performed poorly, with most failing to clear the position-guessing floor on semantic tasks, suggesting their generated answers do not reflect the grouping capabilities of their underlying vision towers.
- Closed Foundation Models: Showed a block-level shift toward human-like grouping, with several models (e.g., Claude-Opus-4.6) achieving both above-human accuracy and human-level error consistency. However, this alignment is not uniform across the tier.
- Training Objectives: The type of grouping a model exhibits tracks its training objective rather than its architecture. Self-supervised training led in color similarity, while contrastive vision–language training led in semantic grouping and shape.
- Task Specificity: Human-like grouping is not a global trait; a model that aligns with humans on one task may diverge sharply on another.
Significance and Claims
The paper claims to provide a reusable, theory-driven yardstick for auditing vision models entering visualization pipelines. By leveraging existing psychophysics data, the methodology allows researchers to determine if a model "sees" charts the way a human audience does without conducting new user studies.
The authors assert that:
- Behavioral grounding is measurable: It is possible to quantify whether models organize visual content via human-like Gestalt principles.
- Accuracy is insufficient: Conventional metrics miss critical differences in perceptual organization; a model can be "more accurate" yet "less human" in its grouping logic.
- Task-indexed evaluation is required: No single score certifies a model as a human proxy. Audits must be specific to the task (e.g., grouping marks vs. recognizing shapes) and the training family.
- Current limitations: The study focuses on isolated marks and objects (the building blocks of charts) rather than full chart contexts (axes, legends), and covers only two Gestalt principles (Closure and Similarity).
The work concludes that while some closed foundation models have reached a level of behavioral alignment where they can serve as human proxies for specific tasks, this capability is earned through specific training recipes rather than scale alone, and it remains task-dependent.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.