1 briefs tagged #quantization.
The method compresses LM-head storage without keeping a dense BF16 head, according to a new arXiv paper.