NuCap: A Numerically Aware Captioning Framework for Improved Numerical Reasoning

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

3

초록

Despite advances in image captioning, existing models struggle to generate captions that include accurate numerical information, especially the number of objects. One reason for this issue is that the dataset used for training has a limited number of samples with numerical information about the image. To address this issue, we propose a new framework, the Numerically Aware Captioning (NuCap) model, to enhance numerical reasoning in caption generation. We extract dual features by combining a region-attended object encoder for finer-grained object features and a spatially attended grid encoder for encoding spatially distributed global features. We also propose a number-focused cross-entropy loss component to increase sensitivity to numerical tokens, and introduce CountCOCO, a dataset for structured understanding of numerical information. Experiments show that our method achieves statistically significant counting performance improvements over state-of-the-art image captioning models while maintaining similar captioning performance. Despite the significant improvement in numerical reasoning power, our proposed approach has significantly fewer parameters and lower inference latency than large-scale vision language models, demonstrating both computational efficiency and stability. NuCap is an image captioning model that can represent specific numerical information in a given image, making it more suitable for applications that require precise object enumeration, such as automated surveillance, store monitoring, and scientific documentation.

키워드

multi-modal learningimage captioningnumerical reasoningImage codingImage enhancementMulti-task learningPhotointerpretation
제목
NuCap: A Numerically Aware Captioning Framework for Improved Numerical Reasoning
저자
Jeong, YunaChoi, Yongsuk
DOI
10.3390/app15105608
발행일
2025-05
유형
Article
저널명
APPLIED SCIENCES-BASEL
15
10
페이지
1 ~ 17