1 briefs tagged #awq.
The method compresses LM-head storage without keeping a dense BF16 head, according to a new arXiv paper.