PIN-0005 · LIVING WHITEPAPER · 2026-09-04

P-PAS+ Adaptive Scheduling With Persistent Stillpoint Restore

candidate gates passed 40/40 repeated workloads validated active profile corrected focused A/B passed 5/5 + 5/5

Executive summary

ProblemLarge cold prefills interfere with active decoding under scheduler pressure.
ContributionP-PAS+ ported to vLLM 0.28 without removing LMCache, MTP-3 or long-context support.
Strongest result+4.90% decode throughput and −4.52% decode wall time in the target paired scenario.
ValidationAll candidate gates and 40/40 repeated candidate workloads passed.
Current statusValidated active profile; corrected memory-headroom A/B passed 5/5 focused runs on each profile.
Key limitationThe paired gain is a two-repetition median for one workload shape, not a universal speed claim.
Download PDF
Abstract. We ported prefill-pressure adaptive scheduling to a production-like Qwen3.8-27B serving profile while preserving persistent LMCache restore, MTP-3, FlashInfer XQA, FP8 KV cache, 131k context and the OpenAI-compatible capability surface. P-PAS+ passed every candidate correctness and restore gate and completed 40/40 repeated workloads. In a preliminary paired canary, it improved decode throughput by 4.90% in the intended one-large-prefill pressure case and was effectively neutral in the dual-prefill case.
Prior work and thanks. P-PAS and P-PAS+ were developed by Timo Sämann. We thank him for publishing the implementation and research openly. Original sources: P-PAS+ GitHub repository and P-PAS: Prefill-Pressure Adaptive Scheduling for Long-Context LLM Serving. Stillpoint’s contribution is an independent adaptation and validation on a different model, vLLM version, GPU class and cache architecture; it is not the original P-PAS+ work.

1. Problem statement

A long cold prefill can monopolize the scheduler token budget while active requests are decoding. Persistent Stillpoint restore eliminates prefill work for exact cached prefixes, but unavoidable cold or partially cached prefills still occur. P-PAS+ complements LMCache: restore avoids work when possible; adaptive scheduling limits interference when the work cannot be avoided.

2. Research questions and hypotheses

  1. RQ1: Can P-PAS+ be ported to vLLM 0.28 and Qwen3.8 without violating the LMCache hybrid-alignment scheduler invariant?
  2. RQ2: Does the candidate preserve deterministic output, long-context retrieval, LMCache restart restore, MTP-3, API compatibility, tool calling, reasoning and image processing?
  3. RQ3: Does adaptive pressure capping protect active decode work when one very large cold prefill overlaps multiple decodes?
  4. RQ4: Does the candidate avoid material regression when two medium-long prefills overlap two decodes?

H1: the candidate passes every declared feature-equivalence and restore gate. H2: decode wall time and throughput improve in the one-large-prefill pressure case. H3: the dual-prefill case remains within ordinary run variance. Any correctness failure, cache corruption, engine death, or material non-target regression rejects the candidate.

3. Related work

Sämann introduced Prefill-Pressure Adaptive Scheduling and the P-PAS+ heterogeneous-pressure extension for long-context serving [1, 2]. Stillpoint’s persistent restore work addresses a complementary layer: exact cached prefixes avoid prefill computation entirely, while P-PAS+ controls interference from prefills that remain unavoidable. The present work is an independent port and validation, not the original scheduling contribution.

4. Stillpoint adaptation

Upstream values could not be copied directly. Qwen3.8 with LMCache hybrid alignment uses a unified block size of 1600 and requires the configured max_num_batched_tokens to remain in [1600, 3200). We retained the validated configured maximum and applied only a runtime pressure cap:

The configured maximum stays unchanged so MTP draft-slot accounting and LMCache input-budget validation retain their established behavior.

5. System under test and lab setup

ModelQwen3.8-27B W4A16 AWQ
RuntimevLLM 0.28.0 · tensor parallel 2
GPU2 × NVIDIA RTX PRO 2000 Blackwell 16 GB · SM120
AttentionFlashInfer XQA · PIECEWISE CUDA graphs
KV cacheFP8 · 160,799-token GPU capacity
Speculative decodingMTP-3
Persistent cacheLMCache multiprocess · disk-backed L2
Hybrid invariantMamba align · unified block size 1600

6. Test suite and experimental methodology

The candidate had to preserve the complete production feature set, not merely start or improve one synthetic trace.

7. Results: paired GOLDEN canary

Two identical repetitions were run per scenario. These are preliminary medians, not universal performance claims.

Scenario / metricGOLDENP-PAS+Change
1 × 100k prefill + 3 decodes · decode wall456.19 s435.59 s4.52% better
Same · decode throughput1.751 tok/s1.837 tok/s4.90% better
Same · makespan457.07 s435.88 s4.64% better
2 × 56k prefills + 2 decodes · decode wall134.13 s134.44 s0.23% worse
Same · decode throughput5.957 tok/s5.944 tok/s0.23% worse
Same · makespan134.12 s134.65 s0.40% worse

8. Results: clean candidate validation

Each workload below completed five repetitions without a candidate request error.

WorkloadMedian wallMedian rateResult
Short decode · 512 output9.522 s54.491 tok/s5/5
8k agent · 512 output17.698 s49.501 tok/s5/5
111k prompt · 256 output142.437 s43.521 tok/s5/5
4 concurrent short12.519 s163.597 aggregate tok/s5/5
4 concurrent 111k569.211 s0.899 aggregate tok/s5/5
Heterogeneous 16k–130k304.698 s2.310 aggregate tok/s5/5
1 large prefill + 3 decodes474.043 s1.730 decode tok/s5/5
2 prefills + 2 decodes140.252 s5.698 decode tok/s5/5

MTP acceptance was 59.15% in the steady suite and 61.75% under mixed pressure.

9. Failure analysis and current interpretation

The extended GOLDEN arm started successfully, passed its real health check and completed five short-decode repetitions. During the second 8k-agent repetition, both workers stopped on CUDA OOM in the MTP proposal path and vLLM raised EngineDeadError. P-PAS+ completed the full candidate sequence.

A focused unchanged GOLDEN retry reproduced the same CUDA OOM. The shared configuration had reserved 93% of GPU memory for vLLM, leaving insufficient activation headroom for a 54 MiB MTP proposal allocation. Reducing gpu_memory_utilization symmetrically to 0.91 preserved LMCache, MTP-3, FP8 KV and both schedulers. GOLDEN and P-PAS+ then each completed 5/5 focused 8k/512 runs without OOM. This isolates a shared memory-envelope defect; it does not prove that P-PAS+ intrinsically uses less memory or prevents OOM.

10. Threats to validity, caveats and non-claims

11. Reproducibility and artifact availability

The engineering repository retains the exact patch, pinned upstream snapshot, raw evidence and SHA-256 manifests. Public artifact links will be added when that repository’s disclosure review is complete.

12. Conclusion

P-PAS+ is a validated active Stillpoint profile: it preserves the full production capability set and completed the entire repeated candidate campaign. The preliminary paired evidence supports the intended scheduling benefit when one very large prefill competes with active decoding, while showing no meaningful benefit in the dual-prefill case. The reproduced GOLDEN OOM was corrected by adding equal activation headroom to both profiles; the corrected focused A/B completed 5/5 runs per profile without OOM.

13. Acknowledgements

We thank Timo Sämann for publishing P-PAS/P-PAS+ and its implementation openly, enabling independent adaptation and validation. We also acknowledge the vLLM, LMCache and FlashInfer open-source communities whose software forms the evaluated stack.

14. References

  1. T. Sämann, “P-PAS: Prefill-Pressure Adaptive Scheduling for Long-Context LLM Serving,” arXiv:2608.15171, 2026. https://arxiv.org/abs/2608.15171.
  2. T. Sämann, “ppas-vllm,” ppas-plus branch, commit 50340d20a2da5ab6fec704dceb83cecf8734dbb5, Apache-2.0. GitHub repository. Accessed 2026-09-04.
  3. vLLM Project, “vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention,” software version 0.28.0. GitHub repository. Accessed 2026-09-04.
  4. LMCache Project, “LMCache,” persistent KV-cache software. GitHub repository. Accessed 2026-09-04.
  5. FlashInfer Project, “FlashInfer,” attention kernels for LLM serving. GitHub repository. Accessed 2026-09-04.

15. Revision and run history

  1. 2026-09-04 · Preliminary paired canary: target case improved by approximately 5%; dual-prefill case neutral.
  2. 2026-09-04 · Clean candidate validation: all gates and 40/40 repeated candidate workloads passed.
  3. 2026-09-04 · Extended GOLDEN arm: CUDA OOM in the MTP proposal path; focused unchanged retry reproduced the failure.
  4. 2026-09-04 · Corrective focused A/B: shared GPU reservation reduced from 0.93 to 0.91; GOLDEN and P-PAS+ each completed 5/5 8k/512 runs with LMCache and MTP-3 retained.