1 briefs tagged #knowledge distillation.
The Hugging Face post describes offline top-K logits and fused chunked KL loss for lower-memory training.