1 briefs tagged #flashprefill.
New papers propose sparse-attention and neural-memory methods to cut long-context inference costs.