Skip to main navigation Skip to search Skip to main content

AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition

  • Yun Wang
  • , Zhaojun Ding
  • , Xuansheng Wu
  • , Siyue Sun
  • , Ninghao Liu
  • , Xiaoming Zhai

Research output: Journal article publicationConference articleAcademic researchpeer-review

Abstract

Automated scoring plays a crucial role in education by reducing the reliance on human raters and offering scalable and immediate evaluation of student work. While large language models (LLMs) have shown strong potential in this task, their use as end-to-end raters faces challenges such as low accuracy, prompt sensitivity, limited interpretability, and rubric misalignment, which hinders practical implementation. To address the limitations, we propose AutoSCORE, a multi-agent LLM framework enhancing automated scoring via rubric-aligned Structured COmponent REcognition. With two agents, AutoSCORE first extracts rubric-relevant components from student responses and encodes them into a structured representation (i.e., Scoring Rubric Component Extraction Agent), which is then used to assign final scores (i.e., Scoring Agent). This design ensures that model reasoning follows a human-like grading process, enhancing interpretability and robustness. We evaluate AutoSCORE on four benchmark datasets from the ASAP benchmark, using both proprietary and open-source LLMs (GPT-4o, LLaMA-3.1-8B, LLaMA-3.1-70B). Across diverse tasks and rubrics, AutoSCORE predominantly improves scoring accuracy, human-machine agreement (QWK, correlations), and reduces error metrics (MAE, RMSE) compared to single-agent baselines, with particularly strong benefits on complex, multidimensional rubrics, and especially large relative gains on smaller LLMs. These results demonstrate that structured component recognition combined with multi-agent design offers a scalable, reliable, and interpretable solution for automated scoring.

Original languageEnglish
Pages (from-to)40898-40906
Number of pages9
JournalProceedings of the AAAI Conference on Artificial Intelligence
Volume40
Issue number48
DOIs
Publication statusPublished - Mar 2026
Externally publishedYes
Event40th AAAI Conference on Artificial Intelligence, AAAI 2026 - Singapore, Singapore
Duration: 20 Jan 202627 Jan 2026

ASJC Scopus subject areas

  • Artificial Intelligence

Fingerprint

Dive into the research topics of 'AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition'. Together they form a unique fingerprint.

Cite this