← Latest papers
💻 computer science

Multi model deliberation improves the stability of automated scoring by large language models across genres and assessment contexts

This study demonstrates that a multi-model deliberation framework, where three heterogeneous Chinese large language models iteratively exchange scores and rationales, significantly improves the stability and precision of automated essay scoring across diverse genres and assessment contexts, with any resulting systematic bias being effectively correctable through standard human-anchored calibration.

Original authors: Xiaoying ZHENG, Benhui CHEN, Yuqing CHEN, Yun YANG

Published 2026-08-19
📖 1 min read☕ Coffee break read

Original authors: Xiaoying ZHENG, Benhui CHEN, Yuqing CHEN, Yun YANG

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Technical Summary: Multi-Model Deliberation for Automated Scoring Stability

Problem Statement
While Large Language Models (LLMs) demonstrate promise in automated essay scoring, their practical deployment in educational assessment is hindered by scoring instability and a lack of cross-context validation. Existing research prioritizes predictive accuracy (agreement with human raters) but often neglects reliability—specifically, the consistency of a model when repeatedly evaluating the same response. As probabilistic generators, LLMs introduce stochastic variability similar to human rater fatigue, which threatens educational equity and validity. Classical measurement theory posits that reliability is a precondition for validity; an instrument yielding different scores for the same response cannot support valid inferences. Furthermore, traditional ensemble methods (e.g., simple averaging) fail to foster "cognitive calibration," as they aggregate isolated judgments without facilitating information exchange or reasoning together.

Methodology
The study proposes a Multi-Model Deliberation Framework grounded in collective intelligence theory. This framework reconceptualizes automated scoring as a deliberative process involving three heterogeneous, open-source Chinese LLMs: Baichuan2-13B, Qwen2.5-14B, and ChatGLM4-9B. The process operates in three sequential rounds:

  1. Independent Scoring (R1): Models score a response in isolation, generating initial scores and rationales.
  2. Information Exchange (R2): Models are presented with the scores and rationales of the other two models. They are prompted to re-evaluate the response, with the option to adjust their scores or maintain their original judgment based on the new evidence.
  3. Final Confirmation (R3): Models receive the full set of adjustments and rationales from R2 and submit a definitive final score.

The framework was evaluated across two complementary datasets to test generalizability and stability:

  • ASAP-AES Corpus: A public benchmark containing approximately 13,000 K-12 student essays distributed across eight writing prompts. For this study, five essays were deliberately sampled from each of the eight prompts (covering upper, middle, and lower score segments) to create a test set of 40 essays, with 30 trials conducted per essay (7,350 trait-level observations).
  • Journalism Corpus: Authentic responses from a Chinese university journalism course (10 essays, 100 trials each, 9,000 scoring records) scored by an expert human rater.

The study compared the proposed deliberation ensemble against single-model baselines, simple mean/median ensembles, and post-deliberation single-model scoring. Metrics included Cross-Model Mean Absolute Error (MAE) for convergence, Coefficient of Variation (CV) for stability, and rank-order correlations with human scores.

Key Results

  • Significant Reduction in Disagreement: On the ASAP-AES corpus, inter-model disagreement (MAE) decreased significantly from R1 to R3 across all prompts (Holm-corrected p < 0.001). The pooled effect size was large (Cohen's dz = 1.08), with a rank-biserial correlation of 0.92, indicating that deliberation consistently drove convergence regardless of genre or rubric scale.
  • Enhanced Stability: On the journalism corpus, the Coefficient of Variation (CV) for the deliberation ensemble dropped from 7.29% (baseline) to 2.78%, representing a 61.9% relative improvement (t(9) = 7.98, p < 0.001). This improvement significantly outperformed simple ensemble strategies (which reduced CV by 40.8%) and post-deliberation single-model scoring (34.2% improvement), demonstrating a synergistic effect between deliberation and aggregation.
  • Round Decomposition: The convergence effect was partitioned into two segments. The information-exchange round (R1 to R2) accounted for approximately two-thirds (66–69%) of the total convergence, while the final confirmation round (R2 to R3) contributed the remaining one-third (31–34%). Both stages were found to be statistically significant and non-redundant.
  • Model Heterogeneity and Behavior: The three models exhibited distinct, consistent roles across datasets. Baichuan2 acted as a conservative "anchor" with minimal score changes; Qwen served as an "active adjuster," frequently revising scores upward; and ChatGLM functioned as a "robust judge," balancing adjustment with resistance to groupthink. This heterogeneity prevented premature convergence.
  • Alignment with Human Scores: While deliberation improved precision (stability), it induced a systematic upward shift in absolute scores relative to human raters (e.g., +13.46 points on the 100-point journalism scale). However, rank-order correspondence with human scores was largely preserved (Spearman ρ = 0.75 on the journalism corpus). The study demonstrated that this systematic bias could be corrected through standard linear calibration against human anchors, reducing Mean Absolute Error by 50.3%.

Key Contributions

  1. Reframing Scoring Reliability: The study shifts the focus from purely predictive accuracy to measurement reliability, proposing a deliberation framework that achieves stability through active cognitive calibration rather than numerical error cancellation.
  2. Empirical Validation: It provides robust evidence of the framework's effectiveness across diverse languages (Chinese/English), genres (argumentative/narrative), and assessment scales, validating the generalizability of multi-model deliberation.
  3. Mechanism Characterization: The research characterizes the specific convergence patterns and adjustment behaviors of heterogeneous models, revealing that information exchange drives the majority of convergence while a final confirmation round provides necessary stabilization.

Significance and Claims
The paper claims that multi-model deliberation substantially enhances the stability of automated scoring, making LLMs more viable for educational assessment where consistency is critical for fairness. The authors argue that while deliberation improves the precision (reliability) of scores, it does not automatically guarantee validity (absolute alignment with human standards). Therefore, the framework is presented as a precision-enhancing layer that must be paired with human-anchored calibration for high-stakes applications. The study concludes that such a system can support formative assessment and reduce grading workloads, provided that human professional judgment remains central to the evaluation process and that systematic biases are managed through calibration. The findings suggest that the "collective wisdom" of heterogeneous models, when structured through deliberation, offers a path toward more reliable automated scoring than simple aggregation or single-model approaches.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →