상세 보기
한국어 벤치마크를 활용한 오픈웨이트 LLM의 아첨 경향 평가: 어조와 논거 강도의 영향을 중심으로
- 김용우;
- 우상미;
- 김영민
초록
As large language models (LLMs) are increasingly being deployed in high-stakes fields such as medicine and law, the phenomenon of sycophancy—where models blindly align with user opinions regardless of factual correctness—is emerging as a major threat to reliability. This study empirically analyzed sycophancy patterns in Korean Physical Common Sense QA (Ko-PIQA) and Korean Medical QA (korMedMCQA) benchmarks, targeting the latest open-weight models (Gemma3 27B, Qwen3 Next 80B, Mistral Large 3). We generated prompts by combining user tones and argument strengths and tracked the changes in the models' responses. Within the scope of this study, no consistent relationship was observed between the models' QA accuracy and their resistance to sycophancy. Furthermore, regressive sycophancy was prominently observed even in the latest open-weight models. It was also observed that the tones triggering sycophancy varied by domain, suggesting the necessity for domain-specific evaluations. Additionally, we confirmed that sycophancy tendencies increased drastically when false authority or fabricated evidence was presented, rather than simple rebuttals. This study emphasizes that LLMs must be able to mitigate sycophancy and maintain honesty even in real-world user scenarios.
키워드
- 제목
- 한국어 벤치마크를 활용한 오픈웨이트 LLM의 아첨 경향 평가: 어조와 논거 강도의 영향을 중심으로
- 제목 (타언어)
- Evaluating Sycophancy in Open-Weights LLMs using Korean Benchmarks: The Impact of Tone and Argument Strength
- 저자
- 김용우; 우상미; 김영민
- 발행일
- 2026-04
- 유형
- Y
- 저널명
- Journal of Information Technology and Applied Engineering
- 권
- 16
- 호
- 1
- 페이지
- 11 ~ 23