論文ID: 2026TAP0003
Knowledge distillation is a widely used technique for transferring knowledge from a large teacher model to a smaller student model. While this approach has achieved remarkable empirical success across various domains, theoretical understanding of the approximation error when the student model has strictly less representational capacity than the teacher remains limited. Existing theoretical analyses have primarily focused on self-distillation settings where teacher and student share the same architecture, leaving a fundamental gap in our understanding of capacity-mismatched distillation. In this paper, we address this gap by analyzing knowledge distillation in multivariate polynomial regression, where the teacher and student models are polynomials of different degrees. We derive a closed-form expression for the optimal student model under KL divergence minimization and prove that the mean squared error between teacher and optimal student can be exactly characterized using the Schur complement of the feature covariance matrix. This characterization reveals that the Schur complement, which represents the conditional covariance of high-order features given low-order features, serves as a distillation difficulty indicator that quantifies how input distribution geometry affects knowledge transfer. Our analysis provides the first exact closed-form expression for approximation error in capacity-mismatched knowledge distillation, connecting this machine learning problem to classical results in linear statistical models. We extend our analysis to polynomial logistic regression and demonstrate that under high-temperature approximation, the expected KL divergence exhibits a specific scaling relationship with temperature. Comprehensive numerical experiments validate our theoretical predictions across various settings.