Stillpoint Lab / Persistent inference state

Keep the context.
Not the cold start.

Knowledge is more useful when you don’t have to compute it again. Stillpoint preserves a hybrid model’s standing place — recurrent state and full-attention KV — for an exact-prefix restore from disk.

Next: one context, many eligible engines Explore the proposed architecture

PIN-0001 · 12/12 hard-suite pass.
Validated exact-prefix restore after a full vLLM process restart. A research result, not a general-purpose memory product.

Living contextIllustrative · not a live backend
GDN · fixed recurrent shapeFA KV · grows with tokensDisk · persistent checkpointRestored · exact prefix
Persist both forms of stateThe fixed-shape GDN snapshot and token-growing full-attention KV are stored together. A process restart clears the running engine, not the disk checkpoint. This is a conceptual diagram, not a live cache or a memory-size benchmark.

Next / Beyond one engine

Proposed architecture · not deployed

Let the context travel.
Not the cold start.

A standing place belongs to an exact prefix, not to one machine. The proposal: find an eligible engine, reuse both forms of state from the nearest useful tier, and keep shared storage separate from execution.

The context fabricSix illustrative decisions · no live telemetry
01 / RequestAgent taskFrozen canonical prefix
+ tenant identity
02 / Manifest gatesCorrect before fast
  • Tenant authorization
  • Fresh knowledge + exact token hash
  • Compatible model, layout & boundary
Matching artifact admitted
03 / Eligible replicas onlyContext-aware routerRank estimated queue + transfer + restore + remaining prefill cost
04 / Execution fleetSame compatible model · illustrative replicas, not specified hardware

Engine A

Eligible · not selected
GPU memoryExecution stateGDNFA KV
CPU memoryHost staging / cached state
Local SSDNode-local persisted artifacts

Available candidate

Engine B

Eligible · selected
GPU memoryExecution stateGDNFA KV
CPU memoryHost staging / cached state
Local SSDNode-local persisted artifacts

NAS → CPU → GPU · restored

Engine C

Eligible · not selected
GPU memoryExecution stateGDNFA KV
CPU memoryHost staging / cached state
Local SSDNode-local persisted artifacts

Available candidate

Shared persistence / not an execution node

NAS artifact vault

Manifest + GDN snapshot + FA KV prefix

Durable derived state. No inference runs here.
Matching artifact available
● GDN recurrent state● Full-attention KV● SSD / NAS persistence● Ready / restored● Cold compile / invalidation
03 / Restore from the shared vaultAssume no eligible engine has a local copy, and the NAS has a current, authorized, compatible artifact. Engine B is assumed to have the lowest estimated total cost. Both GDN and FA KV move from NAS through CPU staging into GPU memory; execution stays on B.
Static route summary. No animation is needed to understand the decision.Global animation control ↑

Scenario assumptions, not measurements. Replica availability, cache residency and relative costs are chosen to explain decisions, not queried from infrastructure. No deployment, timings, hardware sizing or performance claims. A GPU hit is directly resident; a CPU hit still needs a host-to-GPU transfer. Either is different from SSD or NAS restore. Authorization, freshness and compatibility always precede cost ranking.

Roadmap / Select the working set

Discover. Freeze. Then reuse.

The task discovers useful files and knowledge. At a turn boundary, freeze a canonical token prefix and record its manifest. A changed prefix invalidates reuse for the new request: compile a new artifact, never splice opaque recurrent snapshots.

Candidate / Route the replica

Locality is an input, not a verdict.

llm-d ↗ is a candidate routing foundation to investigate, not a deployed Stillpoint integration. A busy warm replica can lose to an idle eligible replica if transfer and restore still cost less than waiting. The llm-d tiered sample approximates GPU/CPU residency; precise event-based tracking is a separate path, and this proposed NAS location directory is not a validated turnkey upstream feature.

Optional upstream research

File selection ≠ replica routing.

SkillZip Pro · arXiv:2608.30785 ↗ researches progressively loaded skill-bundle compression and preservation of file-routing paths. It is an optional upstream direction, not engine-replica routing. It does not establish this distributed design, generic lossless compression, or safe arbitrary composition of saved hybrid state.

01 / The restart experiment

An engine restart.
The same exact answer.

Bury a random key in a long document. Store the state. Fully restart vLLM. Restore the same prefix from disk and ask for the key. Compare against a cold prefill on the matched run.

26k-token context / End-to-end wall time

Cold prefill + answer · measured27.41 s
Disk restore + answer · measured3.60 s
Published PIN-0001 rev 2 measurements, not a live benchmark. Playback accelerated 4×: 6.8525 s cold / 0.90 s restored. Both bars start at zero and advance at the same pixels per measured second; restored stops first at 13.1339% of the cold width. Wall time includes answer generation, not engine startup. No inference is performed. Reduced motion or no JavaScript shows the final comparison.
7.6×

cold-to-restore wall-time ratio
at 26k tokens, patch rev 2

  • Exact buried-key recovery after restart
  • 98.7% cache hit at 26k
  • 12/12 hard-suite pass, including cold control and repeatability cycles
Read methods, all context lengths & limitations →

Measured 2026-08-22 · PIN-0001 rev 2. Qwen3.8-27B, vLLM + LMCache multiprocess, TP=2, speculative decoding on, CPU L1 + disk L2. End-to-end ratios: 2.6× at 5k, 5.7× at 13k, 7.6× at 26k. TTFT ratios: 5.0× / 8.0× / 9.9× (26k: 26.97 s cold / 2.72 s restored). These are scoped lab results, not universal speedups or host-reboot evidence.

Restore compute, not replies

The cached object is inference state for a matching prefix. The model still generates the answer; this is not a response cache.

A negative control, not an isolation guarantee

A never-stored context returned 0% cache hits and took a cold prefill. This does not establish tenant isolation. Per-tenant salted keys remain planned work.

02 / What a standing place contains

One hybrid. Two kinds of memory.

The footprints metaphor only goes so far. A hybrid model keeps both recurrent state and full-attention KV. Saving one without the other is not a correct restore.

Gated DeltaNet / Recurrent layers

A fixed-shape standing place

Tokens update a recurrent snapshot. Its shape is fixed for a given model configuration; its contents change with the prefix. This does not mean unlimited lossless document memory.

Full attention / Remaining layers

A token-indexed trail

Full-attention layers retain key/value state that grows with the prefix. Stillpoint persists that KV alongside the recurrent snapshot. The entire hybrid state is not fixed-size.

Exact prefix, exact boundary. A changed token sequence requires recompilation. No blending or concatenating opaque GDN snapshots; no claim that arbitrary retrieved documents can be swapped into a saved state.

03 / From documented knowledge to runtime state

Keep the source. Reuse the computation.

Knowledge stays readable and versioned in Git. Model state is a separate, derived artifact — useful only with the matching prefix and compatible runtime.

01 / SOURCE

Knowledge in Git

ADRs, bugfixes and runbooks. Flaiwheel is one curated source of context, not the only one.

Human-readable source of truth
02 / COMPILE + PERSIST

The standing place

Recurrent-state snapshot + full-attention KV prefix, keyed by exact token hash and written to the disk tier.

Disk-backed inference state
03 / RESTORE

A warm starting point

vLLM + LMCache restore the matching state. PIN-0001 validates recovery across a full engine process restart.

Matching prefix required

What comes next: selecting a useful working set, freezing its canonical prefix, deciding when to recompile, and graceful eviction are PIN-0002 development goals. They are not production capabilities demonstrated by the restore experiment.

04 / The knowledge layer behind the lab

Research that remembers.

Stillpoint preserves the model’s standing place. Flaiwheel preserves the engineering knowledge that gets us there.

A visual metaphor for accumulated research knowledge — not a diagram of KV-cache transfer. Motion runs only while visible; use Pause animations above to stop both sculptures and evidence playback.

Build on the last experiment.

Experiments are only useful if their lessons survive. We use Flaiwheel to retain architecture decisions, failure analyses, test evidence and session handovers as structured, Git-backed documentation.

Semantic retrieval through the Model Context Protocol (MCP) brings relevant findings back to our coding agents. The next investigation starts from documented evidence, rather than rediscovering the same constraints.

  1. Experiment
  2. Validate
  3. Document
  4. Retrieve
  5. Improve

Correctness gates, failed approaches and the boundaries of every measured result remain part of the research record.

Complementary layers, distinct responsibilities. Flaiwheel stores and retrieves documented knowledge; it does not write model KV caches. Stillpoint researches controlled runtime context and persistent model state. Exact-prefix restore is validated; selective context compilation, graceful eviction and protection against stale or polluted context remain development goals — not completed functionality.

05 / Evidence, artifacts, next questions

An open lab, not a black box.

Stillpoint is the open-source lab of 4rce.com. Every finding is a numbered pin — a reproducible result on the path to persistent state for hybrid GDN models. Published papers separate measured results from the work still ahead.

PIN-0001 / The restore artifact

We traced the failure to two bugs in LMCache 0.5.3: a layout misread that silently dropped 24 of 25 kernel pages, and async host-buffer corruption from mutating shared metadata. A third failure mode — a shadowed variable disguising errors as worker timeouts — lived in our own diagnostic layer and hid the real bugs. Both LMCache bugs are fixed by the patch; vLLM needed no changes.

At the 2026-08-21 investigation, upstream’s intended layout fix did not activate on vLLM ≥ 0.26’s KV layout, including the development branch tested then. The 4-file patch was verified byte-for-byte against a pristine wheel. The fix is geometry-driven: verified at TP=2 (25:1 page geometry) and TP=1 (17:1), rather than tied to one GPU count.

lmcache-053-hybrid-fa.patch· sha256 ec14d5012… (rev 2) ·Full explainer·PIN-0001 paper

12/12 verified on rev 2 on 2026-08-22, applied to a pristine lmcache==0.5.3 install.

The adaptive scheduler completed every correctness, capability, long-context, LMCache restore, MTP-3 and API gate, followed by 40/40 successful repeated candidate workloads. The validated profile used Qwen3.8-27B, vLLM 0.28.0, FlashInfer XQA, PIECEWISE CUDA graphs, FP8 KV and persistent LMCache.

In a preliminary paired two-repetition canary, P-PAS+ improved decode throughput by 4.90%, decode wall time by 4.52% and makespan by 4.64% when one 100k cold prefill competed with three active decodes. The dual-prefill case was effectively neutral.

GOLDEN stopped with CUDA OOM in the extended comparison; an unchanged focused retry reproduced it. Reducing the shared GPU reservation from 0.93 to 0.91 for both profiles preserved LMCache and MTP-3. GOLDEN and P-PAS+ then each completed 5/5 focused 8k-agent runs without OOM. This identifies a shared memory-envelope defect, not proof that P-PAS+ prevents OOM.

Measured on CT112 · 2026-09-04 · candidate gates passed · 40/40 repeated workloadsMethods, upstream credit & run history →

PIN-0001vLLM + LMCache hybrid state store/restoreValidated · 12/12
PIN-0002Session still-points: select, freeze, reuse and evict a working set; per-tenant salted keysNext · development goal
PIN-0003LinearKV on Qwen3.8 — shuffled RAG for hybridsResearch
PIN-0004Cross-model still-point transfer for agent cascadesResearch
PIN-0005P-PAS+ adaptive scheduling with persistent restoreValidated profile