PIN-0001 — Persistent restore for hybrid GDN models

Scientific report. Hybrid GDN standing-place state stores to disk and restores across a full vLLM restart with the exact answer intact. Published 2026-08-22; revised 2026-09-04.

Executive summary

ProblemHybrid GDN state was not restored correctly from persistent LMCache storage.
ContributionA four-file LMCache 0.5.3 fix for hybrid full-attention page geometry and async metadata integrity.
Strongest resultUp to 7.6× lower end-to-end wall time and 9.9× lower TTFT after restart.
Validation12/12 hard-suite pass with exact buried-key recovery and 97.0–98.8% cache hit.
Current statusPIN-0001 green; patch and public explainer available.
Key limitationExact-prefix restore only; changed prefixes must be compiled again.
Download PDF

PIN-0001 result — measured 2026-08-22 (patch rev 2)

Hybrid GDN-class models (verified on Qwen3.8, 27B), vLLM + LMCache multiprocess, TP=2, speculative decoding on, CPU L1 + disk L2. A random 24-hex key is buried inside a long document; the context is stored, the engine is fully restarted, then the same prefix is restored from disk and the key is asked for.

ContextRestore wallCold wallE2E ×Restore TTFTCold TTFTTTFT ×Hit
5k2.13 s5.62 s2.6×1.00 s4.95 s5.0×97.0%
13k2.31 s13.08 s5.7×1.58 s12.62 s8.0×98.8%
26k3.60 s27.41 s7.6×2.72 s26.97 s9.9×98.7%

E2E figures are total wall including answer generation. TTFT is the prefill the restore actually replaces, and is the fairer measure of what was eliminated.

Two root causes were in LMCache 0.5.3, not vLLM: a layout misread that persisted only 1 of 25 kernel pages per chunk, and async host-buffer corruption from mutating shared metadata. A third failure mode — a variable-shadowing TypeError that disguised failures as worker-timeout hangs — lived in our own diagnostic layer, not upstream, and hid the real bugs until it was removed. Upstream ships an intended fix for the layout bug, but it never activates on vLLM ≥ 0.26's KV layout — verified broken through the current development branch. As of 2026-08-21 the 4-file patch — verified byte-for-byte against a pristine wheel — is the only verified working fix on vLLM ≥ 0.26.

Patch: lmcache-053-hybrid-fa.patch (sha256 ec14d5012…, rev 2) · full explainer · repo

The 12/12 figures above were measured on rev 2 (sha256 ec14d5012…) on 2026-08-22, applied to a pristine lmcache==0.5.3 install. Rev 1 (2026-08-21) passed the same suite with the same exact keys and the same hit rates; rev 2 is faster in both arms — including the cold-prefill baseline the restore path does not touch — so the difference is run-to-run variation plus the removal of rev 1's diagnostic instrumentation, not a change in restore behaviour. Rev 1 reference: restore 2.84 / 4.11 / 6.99 s vs cold 6.91 / 16.29 / 33.92 s.

  1. Problem. Hosted prefix caches are best-effort and evaporate. Hybrid GDN models do not have a growing token-indexed KV for most layers — they have a recurrent state. CacheBlend-style recombination does not apply to that state.
  2. Standing place. After a prefix is read, the useful object to persist is the GDN recurrent snapshot plus the remaining full-attention KV — one place you stand, not a trail of footprints.
  3. Planned product order — not yet production functionality. First: lossless session still-points (PIN-0002) — the agent's working set is observed at runtime, frozen into a canonical prefix at a turn boundary, reused every following turn, and evicted when the task is done. Keyed by exact token hash, salted per project; chronically hot docs graduate into a pre-compiled base pack. Second: LinearKV for shuffled RAG on hybrids (PIN-0003). Third: cross-model still-point transfer for agent cascades (PIN-0004, research). Not CacheBlend on GDN pages.
  4. Why it beats re-ranked RAG. Retrieval re-ranking reshuffles search results every turn, so the token sequence changes and even native prefix caching misses on the knowledge portion — every single turn. Freezing the working set into one canonical layout is one honest prefill; every later turn hits it. Text-memory layers (Mem0-style) solve recall, not compute: each remembered fact is re-prefilled at full token cost per session.
  5. Isolation. Dedicated hardware. Verified today: a never-stored context gets 0% hit and a clean cold prefill. This negative control does not establish tenant isolation. Planned for PIN-0002: a per-tenant salt in every cache key. Knowledge stays in the customer’s Git. Flaiwheel is a source of curated chunks, not a writer of vLLM KV.
  6. What we will measure. Exact-answer restore vs cold prefill, latency after disk restore, survival across a vLLM restart, and failure modes (NaN logits when MTP and LMCache disagree). Done for PIN-0001 — numbers above. Still to measure: PIN-0002 freeze economics in real agent sessions (hit rate, TTFT saved, cost of a wrong freeze) and still-point residency under concurrency, then PIN-0003 LinearKV quality on shuffled RAG. No number on this site until it is measured on the pin hardware.

The public artifact — patch, explainer, verification steps — lives in the fix repo. Pin status is tracked on the lab page.