Feature Selection for High-Dimensional Data: A Case Study of NFT Valuation

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

In this study, we propose hedonic models for valuing Non-Fungible Tokens (NFTs) from the Azuki collection. We first analyze the NFT’s metadata and introduce a market volatility-robust dependent variable. Specific information of Azuki attributes is encoded via Term Frequency-Inverse Document Frequency (TF-IDF) to reflect both presence and collection-wide scarcity, yielding hundreds of features for each token. Two hedonic models are considered: a linear model and a squared model. To address high dimensionality, we tailor three variable-selection procedures—forward, backward, and stepwise—and compare them with regularization benchmarks and machine-learning methods. Using actual Azuki transaction data, we evaluate performance on a train-validation partition. The squared model overfits out of sample, while the linear model generalizes better and is adopted as the baseline. Applying variable selection to the linear baseline improves both parsimony and predictive performance. Machine-learning models exhibit very high training fit but notable performance degradation on the validation set, indicating overfitting in this setting. Overall, carefully specified hedonic models combined with principled variable selection offer competitive, interpretable, and more generalizable NFT valuation.

키워드

Azukihedonic modelhigh-dimensional dataNFT valuationNon-Fungible Token (NFT)Term Frequency-Inverse Document Frequency (TF-IDF)variable selectionVARIABLE SELECTION
제목
Feature Selection for High-Dimensional Data: A Case Study of NFT Valuation
저자
Lee, Geun-cheolLee, HeejungKoo, Hoon-Young
DOI
10.12720/jait.17.1.141-152
발행일
2026-01
유형
Article
저널명
JOURNAL OF ADVANCES IN INFORMATION TECHNOLOGY
17
1
페이지
141 ~ 152