상세 보기
MACA: Multi-Axis Covariance Alignment for Robust Distillation of Self-Supervised Speech Models
- Kim, Dong-Hyun;
- Lee, Jae-Hong;
- Chang, Joon-Hyuk
WEB OF SCIENCE
0SCOPUS
0초록
Self-supervised learning (SSL) has become a standard paradigm for speech representation learning, but modern SSL encoders are computationally expensive, motivating compression for resource-limited environments. Knowledge distillation (KD) has emerged as a promising solution to this; however, standard point-wise feature regression is under-constraining and fails to preserve the teacher’s representational geometry. Relation-based distillation adds self-similarity constraints, yet it still captures only a partial view of second-order structure and may amplify spurious correlations under low-resource distillation or domain shift. In this work, we revisit speech SSL distillation from a probabilistic perspective and interpret KD as distribution alignment between teacher and student representations, which decomposes into first-order mean matching and second-order covariance matching. Since full covariance alignment over vectorized sequences is intractable, we propose multi-axis covariance alignment (MACA), a principled extension of relation-based KD that tractably aligns temporal and feature covariances under a matrix-normal surrogate. Concretely, MACA aligns intra-layer covariance matrices by separately matching their diagonal variance and off-diagonal covariance components for stable conditioning and additionally matches inter-layer cross-covariances between consecutive layers to transfer depth-wise dependency structure. To stabilize training, we further introduce an adaptive min–max weighting scheme that balances the second-order alignment terms. Extensive experiments on SUPERB with multiple SSL teacher backbones show consistent gains over representative distillation baselines, especially under low-resource distillation and domain shifts. Analyses of covariance discrepancy and eigenspectra demonstrate the empirical relevance of covariance alignment, while representation autocorrelation and weight-space loss landscapes indicate reduced spurious correlation amplification and flatter, better-centered local geometry, supporting improved generalization.
키워드
- 제목
- MACA: Multi-Axis Covariance Alignment for Robust Distillation of Self-Supervised Speech Models
- 저자
- Kim, Dong-Hyun; Lee, Jae-Hong; Chang, Joon-Hyuk
- 발행일
- 2026-06
- 유형
- Article in press
- 저널명
- Ieee Transactions on Audio Speech and Language Processing
- 권
- 34
- 페이지
- 3328 ~ 3343