Spaces:
Running
Running
File size: 11,776 Bytes
400054a 0c6c82c 400054a 0c6c82c 400054a 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c 44745f2 0c6c82c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 | ---
title: InferScale-Sim
emoji: 📈
colorFrom: indigo
colorTo: blue
sdk: static
app_file: index.html
pinned: false
license: mit
short_description: Interactive LLM serving simulator and SLO planner
---
# InferScale-Sim v0.3.0
**Interactive LLM serving simulator and SLO-aware design laboratory - written in Python and executed entirely in the browser.**
InferScale-Sim explores a systems question:
> How do workload shape, batching, scheduling, KV-cache pressure, reusable prefixes, and prefill/decode disaggregation change tail latency and sustainable LLM-serving capacity?
The public Hugging Face Space uses **no server CPU, no GPU, no API key, and no paid inference provider**. Hugging Face serves static files; Pyodide executes the same Python package used by the local test suite inside a Web Worker on the visitor's ordinary CPU.
> [!IMPORTANT]
> v0.3 ships with **analytical reference latency profiles**, not measured GPU calibration data. It is designed for systems behavior, what-if studies, and comparative exploration. Absolute milliseconds must not be presented as empirical L4/A10G/A100 benchmark results.
## Why simulation?
Exhaustively testing serving configurations on real GPU clusters is expensive. **Vidur** reported finding a LLaMA2-70B deployment configuration in roughly one CPU-hour while estimating that deployment-based exploration would require about **42,000 GPU-hours (~$218K)**. Its evaluated latency predictions were within 9% error.
That research direction has continued. **Revati (2026)** explores GPU-free time-warp emulation of real serving control paths; **LLMServingSim 2.0 (2026)** models heterogeneous and disaggregated serving; **Frontier (May 2026)** models P/D disaggregation, communication, runtime optimizations, stateful workloads, and SLA-dependent Pareto exploration; and **HeteroPanacea (Aug 2026)** studies heterogeneous specialization across prefill/decode/attention/FFN serving stages.
InferScale-Sim is deliberately smaller: a dependency-light, inspectable Python implementation intended to make the core serving trade-offs interactive and understandable.
## v0.3 capabilities
### Workload and serving loop
- deterministic constant, Poisson, and bursty arrival processes
- log-normal prompt/output-length distributions
- static batching baseline
- continuous batching with FCFS
- shortest-job-first scheduling
- least-slack/SLO-aware scheduling
- chunked prefill
- paged KV-cache accounting and VRAM admission control
- analytical roofline-style prefill/decode latency proxy
- FP16 / INT8 / INT4 weight-footprint scenarios
### Metrics and SLOs
- TTFT, TPOT, E2E and queue-latency percentiles
- request and output-token throughput
- **goodput**: completed requests that satisfy both configured SLOs per simulated second
- component TTFT/E2E SLO attainment
- virtual resource utilization
- peak KV usage
- heuristic bottleneck diagnosis with explicit simulator provenance
### Prefix reuse - new in v0.3
v0.3 adds a controlled shared-prefix scenario:
- configurable shared-prefix length
- configurable request reuse fraction
- deterministic cache-hit assignment independent of the generated workload trace
- cached prefill tokens skipped on a hit
- one persistent shared KV allocation instead of per-request duplication
- cache-hit rate and saved-prefill-token telemetry
This is intentionally **not** a complete RadixAttention implementation. It isolates the compute/memory effect of exact reusable prefixes while keeping the simulator small enough to inspect.
### Prefill/decode disaggregation - new in v0.3
The simulator now supports a separate P/D topology with:
- independent prefill/decode accelerator profiles
- configurable prefill and decode worker counts
- independent role utilization
- continuous decode admission between iterations
- explicit KV transfer after prefill
- serialized analytical interconnect model
- configurable interconnect GB/s and base transfer latency
- p95 KV-transfer latency, total transfer volume, and link utilization
- P/D-specific bottleneck diagnoses such as prefill-pool, decode-pool, and transfer pressure
### Interactive tools
**Serving Lab** - run one colocated or P/D scenario and inspect live metrics, timelines, cache behavior, and diagnostics.
**Scheduler Arena** - run every colocated scheduler against the identical deterministic workload.
**Capacity Planner** - robust binary search for the highest request rate where **every repetition** satisfies the target and drains. The trace exposes mean SLO attainment, worst repetition, target, and min/max range.
**Modern Serving Lab** - compare four controlled variants on one workload:
1. colocated
2. colocated + prefix reuse
3. P/D disaggregated
4. P/D disaggregated + prefix reuse
**Design Explorer** - runs a bounded live sweep over scheduler, batch size, prefix caching, and P/D worker splits, then reports two non-dominated frontiers: **raw goodput vs p95 TTFT** and **goodput per accelerator vs p95 TTFT**. The second view makes the resource cost of multi-worker P/D layouts explicit.
All charts can be expanded and exported as meaningfully named PNG files. Result tables support clipboard copy and CSV export.
## Architecture
```text
Hugging Face Static Space
|
| serves files only
v
+----------------------------------------------------------------+
| Browser |
| |
| UI / Chart.js Pyodide Web Worker |
| | | |
| +------------------------------->| |
| v |
| Python inferscale package |
| | |
| +--------------------------------+-------------------+ |
| | | | | |
| workload schedulers KV cache profiles |
| | | | | |
| +--------------------+-----------------+-------------+ |
| | |
| +--------------------+------------------+ |
| | | |
| colocated simulator P/D simulator |
| | |
| prefill -> transfer -> decode|
| | | |
| +--------------------+------------------+ |
| v |
| metrics / SLO / search |
+----------------------------------------------------------------+
```
The discrete-event engine never sleeps for simulated compute time. A predicted 40 ms decode step advances virtual time by 0.040 seconds immediately.
## Run locally
The Python simulator has **zero runtime dependencies** beyond Python 3.10+.
```bash
python -m venv .venv
source .venv/bin/activate
pip install -e '.[dev]'
pytest -q
ruff check src tests scripts
python scripts/release_check.py
```
Run one simulation:
```bash
python scripts/run_simulation.py
```
Serve the browser app:
```bash
python -m http.server 8000
```
Open `http://localhost:8000`. The first page load downloads Pyodide and Chart.js from the CDN; simulation runs then execute locally in the worker.
## Deploy to Hugging Face
Create or reuse a **Static Space** and push this repository. Root metadata already contains:
```yaml
sdk: static
app_file: index.html
```
No Hugging Face secret is required.
Before deployment:
```bash
python scripts/sync_web_python.py
pytest -q
python scripts/release_check.py
```
`sync_web_python.py` mirrors the canonical `src/inferscale/` package into `py/inferscale/`, which the Pyodide worker imports. CI fails if the mirror is stale.
## Core modeling choices
### Goodput
```text
goodput = requests satisfying TTFT and E2E SLOs / simulated makespan
```
Raw throughput can reward overload. Goodput penalizes requests that finish too late to be useful.
### KV-cache model
Per-token KV bytes are approximated as:
```text
2 x layers x KV heads x head dimension x 2 bytes
```
for K and V with FP16 KV state. Paged allocation rounds live per-request state to configurable token blocks.
With prefix reuse enabled, the shared prefix is represented once as persistent KV state and request allocations contain only the uncached suffix plus generated tokens.
### P/D transfer model
After prefill, newly computed prompt KV is moved to the decode pool through a serialized reference link:
```text
transfer_time = base_latency + KV_bytes / interconnect_bandwidth
```
The link is intentionally simple and explicit. It does not claim to reproduce NCCL, NIXL, RDMA, PCIe, or NVLink behavior.
### Latency-profile honesty
`AnalyticalLatencyModel` estimates operation duration from model architecture, accelerator peak FP16 compute, memory bandwidth, quantization footprint, context length, and conservative efficiency factors. It is an analytical proxy.
A future empirical interpolator can replace that backend without rewriting workload generation, scheduling, KV logic, P/D orchestration, or metrics.
## Research lineage
- **Vidur: A Large-Scale Simulation Framework for LLM Inference** (MLSys 2024) - profiling + simulation + deployment search. https://arxiv.org/abs/2405.05465
- **SGLang: Efficient Execution of Structured Language Model Programs** (NeurIPS 2024) - RadixAttention motivates exact shared-prefix reuse. https://arxiv.org/abs/2312.07104
- **TokenSim** (2025) - extensible scheduling and memory-management simulation. https://arxiv.org/abs/2503.08415
- **Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving** (2026) - real control logic with virtualized GPU time. https://arxiv.org/abs/2601.00397
- **LLMServingSim 2.0** (2026) - heterogeneous/disaggregated infrastructure and runtime interactions. https://arxiv.org/abs/2602.23036
- **Frontier: Towards Comprehensive and Accurate LLM Inference Simulation** (2026) - P/D and Attention-FFN disaggregation, runtime optimizations, stateful workloads, and Pareto exploration. https://arxiv.org/abs/2605.21312
- **When Does Disaggregation Pay? / HeteroPanacea** (Aug 2026) - heterogeneous P/D/attention/FFN specialization and cross-stack design exploration. https://arxiv.org/abs/2608.03741
See `docs/research.md`, `docs/methodology.md`, and `docs/validation.md` for scope and limitations.
## Repository
```text
.
|-- src/inferscale/ canonical Python simulator
|-- py/inferscale/ generated browser mirror
|-- tests/ deterministic unit tests
|-- scripts/ local runner + release tooling
|-- docs/ architecture / methodology / research notes
|-- index.html HF Static Space entry point
|-- app.js UI + charts + export tooling
|-- worker.mjs Pyodide Web Worker bridge
`-- styles.css
```
## Roadmap
v0.3 deliberately stops short of pretending to be a production-runtime emulator. Strong next steps would be:
- empirical latency-profile import and held-out calibration report
- multi-turn / agentic session traces
- speculative decoding
- multi-replica routing and tenant fairness
- richer prefix-tree eviction/scheduling
- attention/FFN disaggregation
## License
MIT.
|