Residual-as-Teacher (RaT): Functional Gradients for
Model Distillation and Covariate Shift
Name
yamamoto-kakei-sm-eecs-2026-thesis.pdf
Size
2.15 MB
Format
Adobe PDF
Checksum (MD5)
4e343c4f12072380b1d6e7cea29d1fe6
Author(s)
Yamamoto, Kakei
Advisor(s)
Wainwright, Martin J.
Date Issued
February 2026
Publisher
Massachusetts Institute of Technology
Abstract
Self-training and pseudo-labeling (PL) have emerged as dominant paradigms for unsupervised domain adaptation and knowledge distillation. However, these methods su!er from a fundamental limitation: confirmation bias. When a teacher model is biased due to covariate shift or limited capacity, standard PL forces the student to mimic these errors, preventing it from recovering the true target function. In this paper, we propose Residual-as-Teacher (RaT), a novel framework that reformulates the adaptation process as approximate functional gradient descent. Instead of treating the teacher’s predictions as ground truth, RaT trains a teacher to estimate the residuals of the student model, thereby guiding the student to iteratively correct its errors. We provide a rigorous theoretical analysis of RaT using linear teachers, e.g., Kernel Ridge Regression and Nadaraya-Watson. Our main contribution is a theoretical separation result under ”adversarial” covariate shift and singular density ratios, a regime where the target distribution contains high-frequency components that degenerate in the source domain. We show that while PL is bottlenecked by the teacher’s bias, yielding a degraded convergence rate of O(n⁻²/³); RaT e!ectively acts as a bias-correction mechanism, recovering the fast parametric rate of O(n⁻¹). This demonstrates that learning how to correct a model is more robust to distribution shift than simply learning what to predict.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link