Memorization or Reasoning? Exploring the Idiom Understanding of LLMs

  • Kim, Jisu
  • Shin, Youngwoo
  • Hwang, Uiji
  • Choi, Jihun
  • Xuan, Richeng
  • ... Kim, Taeuk
Citations

SCOPUS

3

초록

Idioms have long posed a challenge due to their unique linguistic properties, which set them apart from other common expressions. While recent studies have leveraged large language models (LLMs) to handle idioms across various tasks, e.g., idiom-containing sentence generation and idiomatic machine translation, little is known about the underlying mechanisms of idiom processing in LLMs, particularly in multilingual settings. To this end, we introduce MIDAS, a new large-scale dataset of idioms in six languages, each paired with its corresponding meaning. Leveraging this resource, we conduct a comprehensive evaluation of LLMs' idiom processing ability, identifying key factors that influence their performance. Our findings suggest that LLMs rely not only on memorization but also adopt a hybrid approach that integrates contextual cues and reasoning, especially when processing compositional idioms. This implies that idiom understanding in LLMs emerges from an interplay between internal knowledge retrieval and reasoning-based inference

키워드

Large datasetsMachine translationNatural language processing systems
제목
Memorization or Reasoning? Exploring the Idiom Understanding of LLMs
저자
Kim, JisuShin, YoungwooHwang, UijiChoi, JihunXuan, RichengKim, Taeuk
DOI
10.18653/v1/2025.emnlp-main.1099
발행일
2025-11
유형
Conference paper
저널명
EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
페이지
21678 ~ 21699