P-PAS+ Adaptive Scheduling With Persistent Stillpoint Restore
Executive summary
1. Problem statement
A long cold prefill can monopolize the scheduler token budget while active requests are decoding. Persistent Stillpoint restore eliminates prefill work for exact cached prefixes, but unavoidable cold or partially cached prefills still occur. P-PAS+ complements LMCache: restore avoids work when possible; adaptive scheduling limits interference when the work cannot be avoided.
2. Research questions and hypotheses
- RQ1: Can P-PAS+ be ported to vLLM 0.28 and Qwen3.8 without violating the LMCache hybrid-alignment scheduler invariant?
- RQ2: Does the candidate preserve deterministic output, long-context retrieval, LMCache restart restore, MTP-3, API compatibility, tool calling, reasoning and image processing?
- RQ3: Does adaptive pressure capping protect active decode work when one very large cold prefill overlaps multiple decodes?
- RQ4: Does the candidate avoid material regression when two medium-long prefills overlap two decodes?
H1: the candidate passes every declared feature-equivalence and restore gate. H2: decode wall time and throughput improve in the one-large-prefill pressure case. H3: the dual-prefill case remains within ordinary run variance. Any correctness failure, cache corruption, engine death, or material non-target regression rejects the candidate.
3. Related work
Sämann introduced Prefill-Pressure Adaptive Scheduling and the P-PAS+ heterogeneous-pressure extension for long-context serving [1, 2]. Stillpoint’s persistent restore work addresses a complementary layer: exact cached prefixes avoid prefill computation entirely, while P-PAS+ controls interference from prefills that remain unavoidable. The present work is an independent port and validation, not the original scheduling contribution.
4. Stillpoint adaptation
Upstream values could not be copied directly. Qwen3.8 with LMCache hybrid alignment uses a
unified block size of 1600 and requires the configured max_num_batched_tokens
to remain in [1600, 3200). We retained the validated configured maximum and
applied only a runtime pressure cap:
B_max = 3199B_cap = 2048large_prefill_threshold = 65536remaining tokens- Pressure rule A: at least two active prefills and one running decode
- Pressure rule B: one large running prefill and at least two running decodes
The configured maximum stays unchanged so MTP draft-slot accounting and LMCache input-budget validation retain their established behavior.
5. System under test and lab setup
6. Test suite and experimental methodology
The candidate had to preserve the complete production feature set, not merely start or improve one synthetic trace.
- Deterministic generation: 20 identical answers and five identical 200-token generations
- Buried-key retrieval at 13k, 30k, 60k, 66.7k, 112k and 130k
- LMCache store and restore from 5k through 112k, including restart recovery
- Model discovery, non-streaming completion and SSE streaming
- German output, reasoning parser, native tool calling and deterministic coding
- Real image processing path and eight-request queue submission
- Repeated short, agent, long-context, concurrent, heterogeneous and mixed-pressure workloads
7. Results: paired GOLDEN canary
Two identical repetitions were run per scenario. These are preliminary medians, not universal performance claims.
| Scenario / metric | GOLDEN | P-PAS+ | Change |
|---|---|---|---|
| 1 × 100k prefill + 3 decodes · decode wall | 456.19 s | 435.59 s | 4.52% better |
| Same · decode throughput | 1.751 tok/s | 1.837 tok/s | 4.90% better |
| Same · makespan | 457.07 s | 435.88 s | 4.64% better |
| 2 × 56k prefills + 2 decodes · decode wall | 134.13 s | 134.44 s | 0.23% worse |
| Same · decode throughput | 5.957 tok/s | 5.944 tok/s | 0.23% worse |
| Same · makespan | 134.12 s | 134.65 s | 0.40% worse |
8. Results: clean candidate validation
Each workload below completed five repetitions without a candidate request error.
| Workload | Median wall | Median rate | Result |
|---|---|---|---|
| Short decode · 512 output | 9.522 s | 54.491 tok/s | 5/5 |
| 8k agent · 512 output | 17.698 s | 49.501 tok/s | 5/5 |
| 111k prompt · 256 output | 142.437 s | 43.521 tok/s | 5/5 |
| 4 concurrent short | 12.519 s | 163.597 aggregate tok/s | 5/5 |
| 4 concurrent 111k | 569.211 s | 0.899 aggregate tok/s | 5/5 |
| Heterogeneous 16k–130k | 304.698 s | 2.310 aggregate tok/s | 5/5 |
| 1 large prefill + 3 decodes | 474.043 s | 1.730 decode tok/s | 5/5 |
| 2 prefills + 2 decodes | 140.252 s | 5.698 decode tok/s | 5/5 |
MTP acceptance was 59.15% in the steady suite and 61.75% under mixed pressure.
9. Failure analysis and current interpretation
The extended GOLDEN arm started successfully, passed its real health check and completed
five short-decode repetitions. During the second 8k-agent repetition, both workers stopped
on CUDA OOM in the MTP proposal path and vLLM raised EngineDeadError. P-PAS+
completed the full candidate sequence.
gpu_memory_utilization symmetrically to 0.91 preserved
LMCache, MTP-3, FP8 KV and both schedulers. GOLDEN and P-PAS+ then each completed 5/5 focused
8k/512 runs without OOM. This isolates a shared memory-envelope defect; it does not prove that
P-PAS+ intrinsically uses less memory or prevents OOM.
10. Threats to validity, caveats and non-claims
- The paired A/B figures are medians from two repetitions per scenario and do not establish production-wide performance.
- The strongest improvement was workload-specific; the dual-prefill result was effectively neutral and slightly favored GOLDEN numerically.
- The GOLDEN OOM does not prove that P-PAS+ uses less memory, prevents OOM, or caused the difference.
- Results apply to the declared Qwen3.8, vLLM 0.28, LMCache, MTP-3, FlashInfer and dual-SM120 setup.
- Historical champion numbers were not substituted for the incomplete contemporaneous extended control arm.
11. Reproducibility and artifact availability
- Upstream implementation: TimoSaemann/ppas-vllm
- Upstream paper: arXiv:2608.15171
- Stillpoint adaptation: vLLM 0.28 scheduler port with a switchable isolated runtime profile
- Evidence set: raw JSON gates, benchmark results, service state and OOM journal, all SHA-256 verified after transfer
The engineering repository retains the exact patch, pinned upstream snapshot, raw evidence and SHA-256 manifests. Public artifact links will be added when that repository’s disclosure review is complete.
12. Conclusion
P-PAS+ is a validated active Stillpoint profile: it preserves the full production capability set and completed the entire repeated candidate campaign. The preliminary paired evidence supports the intended scheduling benefit when one very large prefill competes with active decoding, while showing no meaningful benefit in the dual-prefill case. The reproduced GOLDEN OOM was corrected by adding equal activation headroom to both profiles; the corrected focused A/B completed 5/5 runs per profile without OOM.
13. Acknowledgements
We thank Timo Sämann for publishing P-PAS/P-PAS+ and its implementation openly, enabling independent adaptation and validation. We also acknowledge the vLLM, LMCache and FlashInfer open-source communities whose software forms the evaluated stack.
14. References
- T. Sämann, “P-PAS: Prefill-Pressure Adaptive Scheduling for Long-Context LLM Serving,” arXiv:2608.15171, 2026. https://arxiv.org/abs/2608.15171.
- T. Sämann, “ppas-vllm,”
ppas-plusbranch, commit50340d20a2da5ab6fec704dceb83cecf8734dbb5, Apache-2.0. GitHub repository. Accessed 2026-09-04. - vLLM Project, “vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention,” software version 0.28.0. GitHub repository. Accessed 2026-09-04.
- LMCache Project, “LMCache,” persistent KV-cache software. GitHub repository. Accessed 2026-09-04.
- FlashInfer Project, “FlashInfer,” attention kernels for LLM serving. GitHub repository. Accessed 2026-09-04.
15. Revision and run history
- 2026-09-04 · Preliminary paired canary: target case improved by approximately 5%; dual-prefill case neutral.
- 2026-09-04 · Clean candidate validation: all gates and 40/40 repeated candidate workloads passed.
- 2026-09-04 · Extended GOLDEN arm: CUDA OOM in the MTP proposal path; focused unchanged retry reproduced the failure.
- 2026-09-04 · Corrective focused A/B: shared GPU reservation reduced from 0.93 to 0.91; GOLDEN and P-PAS+ each completed 5/5 8k/512 runs with LMCache and MTP-3 retained.