Engine A
Eligible · not selectedAvailable candidate
Stillpoint Lab / Persistent inference state
Knowledge is more useful when you don’t have to compute it again. Stillpoint preserves a hybrid model’s standing place — recurrent state and full-attention KV — for an exact-prefix restore from disk.
Next: one context, many eligible engines Explore the proposed architecturePIN-0001 · 12/12 hard-suite pass.
Validated exact-prefix restore after a full vLLM process restart. A research result, not a general-purpose memory product.
Next / Beyond one engine
Proposed architecture · not deployedA standing place belongs to an exact prefix, not to one machine. The proposal: find an eligible engine, reuse both forms of state from the nearest useful tier, and keep shared storage separate from execution.
Available candidate
NAS → CPU → GPU · restored
Available candidate
Manifest + GDN snapshot + FA KV prefix
Durable derived state. No inference runs here.Scenario assumptions, not measurements. Replica availability, cache residency and relative costs are chosen to explain decisions, not queried from infrastructure. No deployment, timings, hardware sizing or performance claims. A GPU hit is directly resident; a CPU hit still needs a host-to-GPU transfer. Either is different from SSD or NAS restore. Authorization, freshness and compatibility always precede cost ranking.
The task discovers useful files and knowledge. At a turn boundary, freeze a canonical token prefix and record its manifest. A changed prefix invalidates reuse for the new request: compile a new artifact, never splice opaque recurrent snapshots.
llm-d ↗ is a candidate routing foundation to investigate, not a deployed Stillpoint integration. A busy warm replica can lose to an idle eligible replica if transfer and restore still cost less than waiting. The llm-d tiered sample approximates GPU/CPU residency; precise event-based tracking is a separate path, and this proposed NAS location directory is not a validated turnkey upstream feature.
SkillZip Pro · arXiv:2608.30785 ↗ researches progressively loaded skill-bundle compression and preservation of file-routing paths. It is an optional upstream direction, not engine-replica routing. It does not establish this distributed design, generic lossless compression, or safe arbitrary composition of saved hybrid state.
01 / The restart experiment
Bury a random key in a long document. Store the state. Fully restart vLLM. Restore the same prefix from disk and ask for the key. Compare against a cold prefill on the matched run.
26k-token context / End-to-end wall time
cold-to-restore wall-time ratio
at 26k tokens, patch rev 2
Measured 2026-08-22 · PIN-0001 rev 2. Qwen3.8-27B, vLLM + LMCache multiprocess, TP=2, speculative decoding on, CPU L1 + disk L2. End-to-end ratios: 2.6× at 5k, 5.7× at 13k, 7.6× at 26k. TTFT ratios: 5.0× / 8.0× / 9.9× (26k: 26.97 s cold / 2.72 s restored). These are scoped lab results, not universal speedups or host-reboot evidence.
The cached object is inference state for a matching prefix. The model still generates the answer; this is not a response cache.
A never-stored context returned 0% cache hits and took a cold prefill. This does not establish tenant isolation. Per-tenant salted keys remain planned work.
02 / What a standing place contains
The footprints metaphor only goes so far. A hybrid model keeps both recurrent state and full-attention KV. Saving one without the other is not a correct restore.
Gated DeltaNet / Recurrent layers
Tokens update a recurrent snapshot. Its shape is fixed for a given model configuration; its contents change with the prefix. This does not mean unlimited lossless document memory.
Full attention / Remaining layers
Full-attention layers retain key/value state that grows with the prefix. Stillpoint persists that KV alongside the recurrent snapshot. The entire hybrid state is not fixed-size.
Exact prefix, exact boundary. A changed token sequence requires recompilation. No blending or concatenating opaque GDN snapshots; no claim that arbitrary retrieved documents can be swapped into a saved state.
03 / From documented knowledge to runtime state
Knowledge stays readable and versioned in Git. Model state is a separate, derived artifact — useful only with the matching prefix and compatible runtime.
ADRs, bugfixes and runbooks. Flaiwheel is one curated source of context, not the only one.
Human-readable source of truthRecurrent-state snapshot + full-attention KV prefix, keyed by exact token hash and written to the disk tier.
Disk-backed inference statevLLM + LMCache restore the matching state. PIN-0001 validates recovery across a full engine process restart.
Matching prefix requiredWhat comes next: selecting a useful working set, freezing its canonical prefix, deciding when to recompile, and graceful eviction are PIN-0002 development goals. They are not production capabilities demonstrated by the restore experiment.
04 / The knowledge layer behind the lab
Stillpoint preserves the model’s standing place. Flaiwheel preserves the engineering knowledge that gets us there.
Experiments are only useful if their lessons survive. We use Flaiwheel to retain architecture decisions, failure analyses, test evidence and session handovers as structured, Git-backed documentation.
Semantic retrieval through the Model Context Protocol (MCP) brings relevant findings back to our coding agents. The next investigation starts from documented evidence, rather than rediscovering the same constraints.
Correctness gates, failed approaches and the boundaries of every measured result remain part of the research record.
Complementary layers, distinct responsibilities. Flaiwheel stores and retrieves documented knowledge; it does not write model KV caches. Stillpoint researches controlled runtime context and persistent model state. Exact-prefix restore is validated; selective context compilation, graceful eviction and protection against stale or polluted context remain development goals — not completed functionality.
05 / Evidence, artifacts, next questions
Stillpoint is the open-source lab of 4rce.com. Every finding is a numbered pin — a reproducible result on the path to persistent state for hybrid GDN models. Published papers separate measured results from the work still ahead.
We traced the failure to two bugs in LMCache 0.5.3: a layout misread that silently dropped 24 of 25 kernel pages, and async host-buffer corruption from mutating shared metadata. A third failure mode — a shadowed variable disguising errors as worker timeouts — lived in our own diagnostic layer and hid the real bugs. Both LMCache bugs are fixed by the patch; vLLM needed no changes.
At the 2026-08-21 investigation, upstream’s intended layout fix did not activate on vLLM ≥ 0.26’s KV layout, including the development branch tested then. The 4-file patch was verified byte-for-byte against a pristine wheel. The fix is geometry-driven: verified at TP=2 (25:1 page geometry) and TP=1 (17:1), rather than tied to one GPU count.
lmcache-053-hybrid-fa.patch· sha256 ec14d5012… (rev 2) ·Full explainer·PIN-0001 paper
12/12 verified on rev 2 on 2026-08-22, applied to a pristine lmcache==0.5.3 install.
The adaptive scheduler completed every correctness, capability, long-context, LMCache restore, MTP-3 and API gate, followed by 40/40 successful repeated candidate workloads. The validated profile used Qwen3.8-27B, vLLM 0.28.0, FlashInfer XQA, PIECEWISE CUDA graphs, FP8 KV and persistent LMCache.
In a preliminary paired two-repetition canary, P-PAS+ improved decode throughput by 4.90%, decode wall time by 4.52% and makespan by 4.64% when one 100k cold prefill competed with three active decodes. The dual-prefill case was effectively neutral.
GOLDEN stopped with CUDA OOM in the extended comparison; an unchanged focused retry reproduced it. Reducing the shared GPU reservation from 0.93 to 0.91 for both profiles preserved LMCache and MTP-3. GOLDEN and P-PAS+ then each completed 5/5 focused 8k-agent runs without OOM. This identifies a shared memory-envelope defect, not proof that P-PAS+ prevents OOM.
Measured on CT112 · 2026-09-04 · candidate gates passed · 40/40 repeated workloadsMethods, upstream credit & run history →