科研速览 · Science Skim继续刷下去 · Keep skimming →
◆ Frontiers in artificial intelligence2026-01-01

AdaK: adaptive KV cache budget estimation framework for analyzing long-context large language model inference.

Tianjun Shao

一句话结论 · In one sentence

AdaK reveals estimated KV cache reductions of up to 17.9% relative to fixed-k = 2048 baselines across 16 settings on Qwen3-4B, Qwen3-8B, and Mistral-7B.

原始摘要(英文原文)· Original abstract
INTRODUCTION: The deployment of LLMs on resource-constrained hardware is hindered by the memory-intensive KV Cache mechanism. METHODS: We propose AdaK, an adaptive KV cache budget estimation framework with three strategies: entropy-based thresholding, task-aware lookup table, and a lightweight policy network. RESULTS: AdaK reveals estimated KV cache reductions of up to 17.9% relative to fixed-k = 2048 baselines across 16 settings on Qwen3-4B, Qwen3-8B, and Mistral-7B. DISCUSSION: AdaK's decoupled design enables safe budget estimation as a dynamic ceiling for downstream sparse attention kernels.
读原文 · Read the paper ↗

AI 追问PRO

登录后使用 AI 追问

讨论区

登录后参与讨论

相关论文 · Related

AdaK: adaptive KV cache budget estimation framework for analyzing long-context large language model inference. — 科研速览 Science Skim