상세 보기
Uncertainty-Aware Multi-metric Evaluation of Human–Machine Agreement for LLM-Based Educational Assessment
- Yoo, Jin Eun;
- Kim, Hyeong Gwan;
- Kim, Taeuk
SCOPUS
0초록
Reliable evaluation of large language models (LLMs) in educational settings remains challenging, particularly for ordinal rubrics where human disagreement and class imbalance are pervasive. Using a corpus of teacher utterances rated on a 1–5 ordinal scale, we compare human consensus labels with predictions from multiple LLM-based approaches (fine-tuning and prompting). Performance is evaluated via exact-match measures (accuracy, F1), chance-corrected agreement (κ-family, Gwet’s AC1/AC2), and a distributional distance metric (binned Jensen–Shannon divergence; JSDb). This work demonstrates that no single agreement metric adequately characterizes HM alignment in the presence of human uncertainty. We therefore argue for a robust, multi-metric, and uncertainty-aware evaluation approach as a foundation for reliable LLM-based assessment in education.
키워드
- 제목
- Uncertainty-Aware Multi-metric Evaluation of Human–Machine Agreement for LLM-Based Educational Assessment
- 저자
- Yoo, Jin Eun; Kim, Hyeong Gwan; Kim, Taeuk
- 발행일
- 2026-06
- 유형
- Conference Paper
- 권
- 3031
- 페이지
- 305 ~ 311