InferScale-Sim / docs /validation.md
ArchitSharma's picture
Deepen InferScale simulation research workflow
e5c4ee4
|
Raw
History Blame
2.94 kB

Validation checklist

Release checks prevent software, deployment, and provenance mistakes; they do not pretend the analytical profile has measured-hardware fidelity.

Automated software checks

  • deterministic constant/Poisson/bursty workload generation
  • exact trace replay and trace validation
  • prefill latency monotonicity
  • quantization footprint ordering
  • latency sensitivity-scale behavior
  • colocated end-to-end completion
  • static vs continuous behavioral difference
  • component TTFT/E2E SLO accounting
  • robust repetition-aware capacity search
  • bottleneck-diagnosis provenance
  • prefix hits do not alter the underlying generated request trace
  • prefix reuse reduces modeled prefill work
  • P/D pipeline completion and transfer telemetry
  • configurable P/D worker counts
  • four-scenario topology/cache comparison
  • bounded design-space sweep and two Pareto objectives
  • paired common-seed research study
  • bootstrap paired-effect intervals
  • analytical-profile perturbation study
  • external-measurement validation report generation
  • deterministic stateful agent-session trace generation
  • zero cross-turn hits under immediate eviction
  • KV reuse and reduced recomputation under retention + affinity
  • common-trace four-policy agent comparison
  • TTL latency/memory Pareto frontier
  • ASCII-only public UI labels
  • chart export controls present
  • explicit planner worst-repetition and target columns
  • no product-style footer
  • Hugging Face short_description <= 60 characters
  • sdk: static metadata
  • canonical Python source equals browser mirror
  • every browser Python module is included by the worker
  • provenance remains analytical-reference
  • JavaScript syntax parse
  • Python compilation

Empirical validation protocol

A future hardware run should be stored as external validation cases, not copied into analytical profiles without provenance. At minimum each case should record:

  • model / revision
  • serving topology and scheduler
  • accelerator / count
  • precision
  • workload arrival process or trace ID
  • prompt/output distribution or exact trace
  • SLO definition
  • p95 TTFT / p95 E2E / goodput
  • runtime/software versions

Run:

python scripts/validate_measurements.py measured_cases.json --output validation_report.json

The report provides residuals and aggregate percentage errors. Calibration should be fit on a training subset and reported on held-out cases if empirical fitting is added later.

Not claimed

  • empirical L4/A10G/A100 latency accuracy
  • exact vLLM/SGLang scheduler equivalence
  • CUDA-kernel modeling
  • real network-protocol fidelity
  • radix-tree prefix-cache eviction/scheduling
  • speculative decoding
  • Attention-FFN disaggregation
  • production-fidelity agentic serving or dynamic batching inside Agent Sessions
  • learned workflow prediction, proactive prefetch, host/CXL KV tiers, or real tool execution

These are explicit scope boundaries, not hidden assumptions.