MACA: Multi-Axis Covariance Alignment for Robust Distillation of Self-Supervised Speech Models

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Self-supervised learning (SSL) has become a standard paradigm for speech representation learning, but modern SSL encoders are computationally expensive, motivating compression for resource-limited environments. Knowledge distillation (KD) has emerged as a promising solution to this; however, standard point-wise feature regression is under-constraining and fails to preserve the teacher’s representational geometry. Relation-based distillation adds self-similarity constraints, yet it still captures only a partial view of second-order structure and may amplify spurious correlations under low-resource distillation or domain shift. In this work, we revisit speech SSL distillation from a probabilistic perspective and interpret KD as distribution alignment between teacher and student representations, which decomposes into first-order mean matching and second-order covariance matching. Since full covariance alignment over vectorized sequences is intractable, we propose multi-axis covariance alignment (MACA), a principled extension of relation-based KD that tractably aligns temporal and feature covariances under a matrix-normal surrogate. Concretely, MACA aligns intra-layer covariance matrices by separately matching their diagonal variance and off-diagonal covariance components for stable conditioning and additionally matches inter-layer cross-covariances between consecutive layers to transfer depth-wise dependency structure. To stabilize training, we further introduce an adaptive min–max weighting scheme that balances the second-order alignment terms. Extensive experiments on SUPERB with multiple SSL teacher backbones show consistent gains over representative distillation baselines, especially under low-resource distillation and domain shifts. Analyses of covariance discrepancy and eigenspectra demonstrate the empirical relevance of covariance alignment, while representation autocorrelation and weight-space loss landscapes indicate reduced spurious correlation amplification and flatter, better-centered local geometry, supporting improved generalization.

키워드

covariance alignmentKnowledge distillationsecond-order statisticsself-supervised learningspeech representation learningREPRESENTATION
제목
MACA: Multi-Axis Covariance Alignment for Robust Distillation of Self-Supervised Speech Models
저자
Kim, Dong-HyunLee, Jae-HongChang, Joon-Hyuk
DOI
10.1109/TASLPRO.2026.3703215
발행일
2026-06
유형
Article in press
저널명
Ieee Transactions on Audio Speech and Language Processing
34
페이지
3328 ~ 3343