Domain-specific Large Language Model Pretraining for Construction Specification Review

Citations

SCOPUS

0

초록

With advancements in large language models (LLMs), there has been growing attention to automatic construction specification reviews, enabling faster and more standardized assessments. However, general-purpose LLMs’ effectiveness has been bounded when applied to construction contexts, because they lack necessary domain-specific vocabulary and knowledge. Therefore, we propose a domain-specialization LLM pretraining method to leverage construction terms and knowledge and quantitatively evaluate its efficacy in two specification review tasks. To this end, we generate two construction-specific corpora, i.e., close-domain and in-domain, and investigate their hybrid effects on LLM performance at varying scales. Results reveal that construction-specific contextualization led to a significant improvement, with a 6.9% increase in accuracy for extractive question answering and an 8.0% enhancement in information retrieval accuracy. We also observe that augmenting in-domain corpus with close-domain data, can achieve performance comparable to models trained with extremely large in-domain data. These will facilitate the adoption of LLMs in construction and support specification reviews. © 2026 International Association on Automation and Robotics in Construction. All Rights Reserved.

키워드

Construction specificationDomain-specificLarge language modelPretrainingComputational linguisticsDomain KnowledgeInformation retrievalKnowledge management
제목
Domain-specific Large Language Model Pretraining for Construction Specification Review
저자
Wang, ShuyiKim, Jinwoo
DOI
10.22260/ISARC2026/0172
발행일
2026-00
유형
Conference paper
저널명
Proceedings of the International Symposium on Automation and Robotics in Construction
페이지
1340 ~ 1347