Anna Cordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jesus Olivera
To address KV cache memory bottlenecks in long-context LLM inference, the paper proposes an adaptive compression method that allocates cross-layer low-rank basis factorization and residuals based on token importance.
KV caches consume large memory bandwidth and capacity during long-context inference. Existing compression applies uniform budgets across layers or tokens, degrading retrieval when lexical cues and semantic states require different preservation levels.
Key-value states are factorized across neighboring transformer layers using shared low-rank channel bases, while lightweight token-specific residuals are retained where attention is sensitive. A token-conditional depth router assigns higher reconstruction rank to instruction- and retrieval-critical tokens, and calibration-free online error tracking from attention-output probes adapts compression during generation without retraining. A fused CUDA kernel jointly performs basis lookup, residual dequantization, and attention projection to reduce decode-time memory traffic.
Across LongBench, Needle-in-a-Haystack, L-Eval, and long-form QA and summarization benchmarks, DepthWeave-KV achieves near-full-cache quality with up to 8.3x KV memory reduction and 72.8 tokens per second at 64K context, improving average score and retrieval accuracy over prior compressed caches.