Uncertainty-Aware Multi-metric Evaluation of Human–Machine Agreement for LLM-Based Educational Assessment

Citations

SCOPUS

0

초록

Reliable evaluation of large language models (LLMs) in educational settings remains challenging, particularly for ordinal rubrics where human disagreement and class imbalance are pervasive. Using a corpus of teacher utterances rated on a 1–5 ordinal scale, we compare human consensus labels with predictions from multiple LLM-based approaches (fine-tuning and prompting). Performance is evaluated via exact-match measures (accuracy, F1), chance-corrected agreement (κ-family, Gwet’s AC1/AC2), and a distributional distance metric (binned Jensen–Shannon divergence; JSDb). This work demonstrates that no single agreement metric adequately characterizes HM alignment in the presence of human uncertainty. We therefore argue for a robust, multi-metric, and uncertainty-aware evaluation approach as a foundation for reliable LLM-based assessment in education.

키워드

Agreement MetricsClass ImbalanceClassroom DiscourseHuman–Machine AgreementLarge Language ModelsLLM-as-a-JudgeComputational linguistics
제목
Uncertainty-Aware Multi-metric Evaluation of Human–Machine Agreement for LLM-Based Educational Assessment
저자
Yoo, Jin EunKim, Hyeong GwanKim, Taeuk
DOI
10.1007/978-3-032-29788-4_43
발행일
2026-06
유형
Conference Paper
저널명
Communications in Computer and Information Science
3031
페이지
305 ~ 311