Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference

arXiv:2602.08329v2 Announce Type: replace-cross Abstract: A core bottleneck in large language model (LLM) inference is the cost of attending over the ever-growing key-value (KV) cache. Although near-oracle top-k KV selection can preserve the quality of dense attention while sharply reducing…

aiscience

Sources

Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference · TechNews