#kv-cache
2 briefs tagged #kv-cache.
Research
Researchers target sparse attention for long-context AI
New papers propose sparse-attention and neural-memory methods to cut long-context inference costs.
arXiv cs.CL+1 outlet·3h ago
Research
New studies probe serving systems for agentic AI
Four papers examine latency, caching, traces and edge deployment as LLM workloads become more agentic.
arXiv cs.AI+1 outlet·3d ago