Why KV cache is now your bottleneck

Two different sources point at the same pain point:

  • Agentic inference runs many observation–reasoning–action cycles and “accumulates long KV caches across iterative observation-reasoning-action loops, imposing substantial memory overhead and limiting serving throughput” ActKV.
  • In applied systems, “the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs” even when the base weights are quantized DeepEdu-v1.

So even if you tame static weights (AWQ, GPTQ, etc.), the *dynamic* footprint from KV plus prefill dominates for long-horizon agents.

DeepEdu-v1 responds by overhauling the *context* side: a long-context inference engine “amortizes token selection from per-sub-chunk to per-cluster granularity” and reduces retrieval calls 7.7× vs a selective-attention baseline, cutting time-to-first-token by ~35% while keeping or improving accuracy DeepEdu-v1. That handles how much you *feed* into the model.

ActKV attacks the other side of the problem: how much of the model’s internal history you keep.

What ActKV claims to do

The abstract positions ActKV as “the first KV cache compression framework tailored for agentic LLM inference” ActKV. The central ideas:

  • Action-centric value function: Instead of treating every token’s KV equally, it “establish[es] a compression criterion that values KV entries by their contribution to action generation and prioritizes action quality” ActKV.
  • Agent-aware challenges: Iterative execution, dynamic memory demand, and “scattered action-critical entries” create coupled problems in eviction policy design, budget allocation, and paged-memory integration ActKV.
  • Three components:
  1. *Action-oriented KV cache eviction* that “exploits stable action access patterns to retain entries critical to future actions” and aims for “reliable task progress under compression” ActKV.
  2. *Confidence-driven adaptive budget allocation* that “uses LLM’s intrinsic confidence to adapt the budget to evolving action-critical memory demands” ActKV.
  3. *Page-aware compression management* that “standardizes compression into three primitives with customized kernels, realizing practical throughput gains” ActKV.

On “long-trace tasks”, the paper reports:

  • Using only 25.98% of peak KV cache memory vs a full KV baseline.
  • Retaining 98.53% of the baseline’s accuracy on average.
  • Achieving 3.97× token throughput and 3.58× task throughput, claiming state-of-the-art performance ActKV.

So, as a system builder, what does it mean to make your KV compression *action-guided*, and how would you wire something like this into an existing stack?

The abstract leaves the implementation details out, but we can reason from the design levers it names and tie them to concrete patterns seen in other systems.

Step 1: Make actions a first-class optimization target

Most KV pruning or compression work assumes all output tokens matter about equally; you pick some distance-based or salience-based heuristic and try to preserve perplexity or generic task accuracy.

Agentic loops are not symmetric:

  • Many tokens in an agent conversation are *expository*: chain-of-thought, tool logs, paraphrases.
  • A small subset are actions that trigger tools, change environment state, or select next steps.

ActKV explicitly notes an “asymmetric importance of actions in driving task progress” and designs its compression criterion around *contribution to action generation* ActKV.

In practice, that means you’d:

  • Label action spans in your traces (e.g., tool calls, API arguments, routing decisions).
  • Train or derive a notion of which past tokens the model attends to when it generates actions, versus when it’s just narrating intermediate reasoning.
  • Treat KV entries that matter for future actions as high-value; other entries become candidates for eviction or more aggressive compression.

The abstract says eviction “exploits stable action access patterns” [ActKV](http://arxiv.org/abs/2609.31395v1], implying that across traces, the model tends to look at similar structures when deciding actions. You could imagine:

  • Tool schemas, earlier tool outputs, and specific instruction lines being recurrently attended to when emitting tool_name or arguments.
  • Ephemeral chit-chat or verbose rationale contributing less to these decisions.

From a stack perspective, this suggests a separation:

  • Action-critical KV pages: keep resident on fast memory (e.g., HBM) with minimal quantization.
  • Non-critical or low-impact pages: eligible for eviction, offloading, or aggressive compression.

ActKV’s reported 98.53% accuracy with ~26% KV peak footprint ActKV suggests that, on their long-trace benchmarks, a big fraction of the accumulated KV is indeed disposable once you optimize for action quality instead of raw text fidelity.

Step 2: Use confidence as a control signal, not just a metric

ActKV’s second component is “confidence-driven adaptive budget allocation” that “uses LLM’s intrinsic confidence to adapt the budget to evolving action-critical memory demands” ActKV.

This parallels how some decision-style systems already surface probabilities as *outputs*:

  • A Jev-like decision setup with GLM-5.3-Flash takes a prompt ending with choice_index: and then reads the distribution over option indices from a single forward pass System One from GLM-Flash.
  • The implementation uses vLLM’s logprob_token_ids to get precise log-probs for each option index, then normalizes them into a probability distribution System One from GLM-Flash.

In that work, these probabilities are used mostly downstream—for routing and human-in-the-loop decisions. ActKV flips the direction: similar “intrinsic confidence” signals become upstream control for *how much KV memory to spend*.

Operationally, this invites a control loop:

  • When the model is confident about its actions, you can likely tighten the KV budget without much risk: evict more aggressively, or accept higher compression ratios.
  • When confidence drops, you expand the budget: preserve more past KV, back off compression, or even rehydrate from slower tiers.

You could imagine, for example, using:

  • Logit margins between the top-1 and top-2 action tokens.
  • Entropy of the distribution over tool choices or discrete decisions.

as proxies for “intrinsic confidence”. That is consistent with what the ActKV abstract describes but is not spelled out, so treat it as an engineering pattern rather than a claim about the specific implementation.

Key point for builders: rather than choosing a static KV size per request, you dynamically trade memory for reliability guided by the model’s own uncertainty. This is especially attractive in long traces where memory pressure and task difficulty are non-uniform over time.

Step 3: Align your policy with your paging and kernels

The third component ActKV calls out is “page-aware compression management” that “standardizes compression into three primitives with customized kernels, realizing practical throughput gains” ActKV.

Long-context systems like DeepEdu-v1 already show that amortizing operations across larger units—“per-cluster granularity” instead of per sub-chunk—can reduce overhead sharply DeepEdu-v1. ActKV appears to make a similar move at the KV memory level:

  • Treat KV not as arbitrary tensors, but as paged memory with well-defined operations.
  • Implement compression and eviction as a small set of standard primitives with dedicated kernels.

There are a few practical reasons this matters:

  • Predictable performance: When you have a fixed palette of primitives, you can hand-optimize kernels (e.g., in CUDA or Triton) without explosion in code paths.
  • Better paging decisions: If eviction operates on page boundaries, your scheduler and allocator can reason cleanly about moving pages between GPU and host, or between different GPU pools.
  • Lower bookkeeping overhead: Instead of bespoke logic at each layer or request, you call into a unified “compression manager”.

The abstract doesn’t list the three primitives, but a typical design might include (this is illustrative, not a description of the paper):

  • *Evict*: drop or demote a KV page.
  • *Compress*: lossily shrink a page (e.g., lower precision or a reduced representation).
  • *Restore*: re-inflate a compressed page when needed.

The reported 3.97× token throughput and 3.58× task throughput vs a full-KV baseline on long traces ActKV suggest that these page-aware kernels and policies cut enough memory pressure to significantly improve concurrency and decoding speed, not just peak footprint.

How it could fit into a modern agent stack

Even though ActKV is a research prototype, you can map its ideas to existing production stacks:

  • Serving layer: If you’re on vLLM, you’re already used to features like allowed_token_ids and log-prob APIs for decision-style outputs System One from GLM-Flash. ActKV-like logic would live alongside vLLM’s KV management, wrapping or extending its paging subsystem.
  • Agent framework: Tool-using agents (for example, the self-improving agentic layer in DeepEdu’s SCALE framework DeepEdu-v1) are natural clients: they have explicit actions, traces, and task-oriented metrics that let you define “action quality”.
  • Workload types:
  • Long educational dialogues (DeepEdu’s tutoring deployments DeepEdu-v1).
  • High-volume decision calls where actions are a single token (similar to Jev-like setups System One from GLM-Flash) but interleaved with occasional longer reasoning bursts.
  • Multi-turn tool orchestration.

You’d wire it approximately as:

  1. Instrument actions and confidence:
  • Mark action segments in your prompts and outputs.
  • Record per-action confidence measures from token logits.
  1. Attach a KV policy engine:
  • For each request, maintain an *action-criticality state* derived from past attention patterns or action labels (following the ActKV notion of “stable action access patterns” ActKV).
  • For each step, decide (a) which pages are critical, (b) how much total KV budget the request gets (confidence-driven), and (c) which compression primitive to apply to each page.
  1. Integrate with paged KV allocator:
  • Extend your serving layer’s KV allocator to respect these per-page decisions.
  • Ensure you can account for both peak and average KV footprint per request, to see if you’re moving toward the ~26% peak figure ActKV reports ActKV.
  1. Evaluate on long-trace tasks:
  • Use tasks with many agent steps—similar in spirit to DeepEdu’s complex agent benchmarks, where they report lifting “agentic accuracy from 70.0% to 79.5% on complex tasks” with their own framework DeepEdu-v1.
  • Compare action success and throughput versus a full-KV baseline.

Why this matters now

Two trends collide here:

  • Agentic patterns are moving from toy demos to deployed systems. DeepEdu-v1 describes a deployed AI-tutoring system running on consumer GPUs for Vietnamese education, where long-context tutoring sessions must obey data-sovereignty laws and still perform well DeepEdu-v1. That’s a concrete example of long-horizon, regulation-constrained agents.
  • Decision-style LLM usage is rising. The GLM-5.3-Flash work shows how to turn a standard LLM into a Jev-like “System One” model, making typed decisions in a single forward pass, with accuracy on par with Jev and much better than a smaller alternative model, across 29 datasets System One from GLM-Flash.

In both cases, the per-request cost and latency dominate viability:

  • DeepEdu reports that their long-context engine, in deployed form, “achieves a nearly ×2 TTFT speedup over standard vLLM serving” DeepEdu-v1.
  • The GLM-5.3-Flash decision setup yields per-decision latencies typically “within a few hundred milliseconds” and, at list prices, about €62 per million decisions vs €16 for Jev System One from GLM-Flash.

ActKV claims another large factor—roughly 3–4× throughput uplift on long agent traces—without retraining the base model, by touching only KV management and kernels ActKV. If these numbers hold up under independent evaluation and across more workloads, this is one of the highest-leverage places you can optimize an agentic stack.

What to watch and how to experiment

If you’re running or planning long-lived agents, the ActKV paper suggests a few concrete lines of experimentation:

  • Action-aware retention: Even without full ActKV, you can start by *not* treating all tokens equally—prefer retaining instructions, tool specs, and recent tool I/O over verbose reasoning logs.
  • Confidence-gated KV budgets: Use your model’s log-probs to decide when to be stingy vs generous with KV retention, especially around high-stakes actions.
  • Page-aligned compression: If your KV manager isn’t already page-based, consider aligning it that way so you can later plug in eviction/compression policies with predictable kernel performance.

At the research level, key open questions include:

  • How robust “stable action access patterns” really are across different models and tasks.
  • How well confidence-derived budgets generalize vs hand-tuned KV limits.
  • Interactions with other optimizations like the long-context clustering used in DeepEdu’s SCALE engine DeepEdu-v1.

For production teams, the right move is likely incremental: steal the design ideas (action-centric metrics, confidence feedback, page-level operations) and implement a simple policy in your existing serving framework, then iterate toward something ActKV-like as you observe behavior.

What is not documented

The available sources do not document:

  • The exact algorithmic details of ActKV’s eviction policy, including how “stable action access patterns” are detected or estimated.
  • How “LLM’s intrinsic confidence” is computed in ActKV (specific logits, margins, or calibration methods are not specified).
  • The precise definition of the three compression primitives or the layout and implementation details of the “customized kernels” in the page-aware manager.
  • The models, datasets, or hardware setups used to obtain the 98.53% accuracy, 25.98% peak KV memory, and 3.97× / 3.58× throughput figures reported for ActKV.
  • Any comparative results between ActKV and other specific KV compression baselines beyond the statement that it delivers state-of-the-art performance.