PIN-0001 — Persistent restore for hybrid GDN models
Scientific report. Hybrid GDN standing-place state stores to disk and restores across a full vLLM restart with the exact answer intact. Published 2026-08-22; revised 2026-09-04.
Executive summary
PIN-0001 result — measured 2026-08-22 (patch rev 2)
Hybrid GDN-class models (verified on Qwen3.8, 27B), vLLM + LMCache multiprocess, TP=2, speculative decoding on, CPU L1 + disk L2. A random 24-hex key is buried inside a long document; the context is stored, the engine is fully restarted, then the same prefix is restored from disk and the key is asked for.
| Context | Restore wall | Cold wall | E2E × | Restore TTFT | Cold TTFT | TTFT × | Hit |
|---|---|---|---|---|---|---|---|
| 5k | 2.13 s | 5.62 s | 2.6× | 1.00 s | 4.95 s | 5.0× | 97.0% |
| 13k | 2.31 s | 13.08 s | 5.7× | 1.58 s | 12.62 s | 8.0× | 98.8% |
| 26k | 3.60 s | 27.41 s | 7.6× | 2.72 s | 26.97 s | 9.9× | 98.7% |
- Correctness: exact buried key, character-for-character, on every restore run.
- Negative control: a never-stored context gets 0% hit and a clean cold prefill. This is not a general tenant-isolation guarantee.
- Stability: 12/12 pass — multi-length restore, warm re-ask idempotence, cold control, and two independent repeatability cycles.
- TP-agnostic: the same patch verified at TP=2 (25:1 page geometry) and TP=1 on a different hybrid (17:1) — exact key, 95.4% hit, 0.72 s vs 1.53 s. The fix is geometry-driven, not configuration-driven.
E2E figures are total wall including answer generation. TTFT is the prefill the restore actually replaces, and is the fairer measure of what was eliminated.
Two root causes were in LMCache 0.5.3, not vLLM: a layout misread that
persisted only 1 of 25 kernel pages per chunk, and async host-buffer
corruption from mutating shared metadata. A third failure mode — a
variable-shadowing TypeError that disguised failures as
worker-timeout hangs — lived in our own diagnostic layer, not upstream,
and hid the real bugs until it was removed. Upstream ships an intended
fix for the layout bug, but it never activates on vLLM ≥ 0.26's KV
layout — verified broken through the current development branch. As of
2026-08-21 the 4-file patch — verified byte-for-byte against a pristine
wheel — is the only verified working fix on vLLM ≥ 0.26.
Patch: lmcache-053-hybrid-fa.patch
(sha256 ec14d5012…, rev 2) ·
full explainer ·
repo
The 12/12 figures above were measured on rev 2
(sha256 ec14d5012…) on 2026-08-22, applied to a pristine
lmcache==0.5.3 install. Rev 1 (2026-08-21) passed the same
suite with the same exact keys and the same hit rates; rev 2 is faster in
both arms — including the cold-prefill baseline the restore path does not
touch — so the difference is run-to-run variation plus the removal of
rev 1's diagnostic instrumentation, not a change in restore behaviour.
Rev 1 reference: restore 2.84 / 4.11 / 6.99 s vs cold 6.91 / 16.29 / 33.92 s.
- Problem. Hosted prefix caches are best-effort and evaporate. Hybrid GDN models do not have a growing token-indexed KV for most layers — they have a recurrent state. CacheBlend-style recombination does not apply to that state.
- Standing place. After a prefix is read, the useful object to persist is the GDN recurrent snapshot plus the remaining full-attention KV — one place you stand, not a trail of footprints.
- Planned product order — not yet production functionality. First: lossless session still-points (PIN-0002) — the agent's working set is observed at runtime, frozen into a canonical prefix at a turn boundary, reused every following turn, and evicted when the task is done. Keyed by exact token hash, salted per project; chronically hot docs graduate into a pre-compiled base pack. Second: LinearKV for shuffled RAG on hybrids (PIN-0003). Third: cross-model still-point transfer for agent cascades (PIN-0004, research). Not CacheBlend on GDN pages.
- Why it beats re-ranked RAG. Retrieval re-ranking reshuffles search results every turn, so the token sequence changes and even native prefix caching misses on the knowledge portion — every single turn. Freezing the working set into one canonical layout is one honest prefill; every later turn hits it. Text-memory layers (Mem0-style) solve recall, not compute: each remembered fact is re-prefilled at full token cost per session.
- Isolation. Dedicated hardware. Verified today: a never-stored context gets 0% hit and a clean cold prefill. This negative control does not establish tenant isolation. Planned for PIN-0002: a per-tenant salt in every cache key. Knowledge stays in the customer’s Git. Flaiwheel is a source of curated chunks, not a writer of vLLM KV.
- What we will measure. Exact-answer restore vs cold prefill, latency after disk restore, survival across a vLLM restart, and failure modes (NaN logits when MTP and LMCache disagree). Done for PIN-0001 — numbers above. Still to measure: PIN-0002 freeze economics in real agent sessions (hit rate, TTFT saved, cost of a wrong freeze) and still-point residency under concurrency, then PIN-0003 LinearKV quality on shuffled RAG. No number on this site until it is measured on the pin hardware.
The public artifact — patch, explainer, verification steps — lives in the fix repo. Pin status is tracked on the lab page.