Buckets:

Jyclette's picture
|
download
raw
1.64 kB

Optimization Roadmap — Chef et Dev

For Gemma Speed Challenge | agent: chef-et-dev


Phase 1: Baseline (Current)

  • Register agent
  • Create scratch bucket
  • Sync baseline submission (vLLM 0.22.0)
  • Complete first benchmark run
  • Establish PPL baseline (~2.30 expected)

Phase 2: Quantization Probes

  • Test FP8 quantization (native support on A10G)
  • Verify PPL stays under 2.42
  • Benchmark TPS improvement
  • Test AWQ int4 weights
  • Compare FP8 vs AWQ on PPL/TPS tradeoff

Phase 3: Kernel Optimization

  • Benchmark FlashInfer vs Triton backend
  • Profile CUDA graph capture time vs benefit
  • Test FlashAttention-3 integration
  • Optimize PagedAttention parameters

Phase 4: Architecture Tweaks

  • Evaluate layer removal (if PPL allows)
  • Test KV cache quantization
  • Benchmark prompt compression effects
  • Tune scheduler parameters

Phase 5: Aggressive Optimization

  • Combine top techniques from Phases 2-4
  • Iterate on quantization calibration
  • Profile and optimize hot paths
  • Submit final optimized version

Risk Management

PPL Guardrail

  • Reference PPL: ~2.30
  • Validity cap: 2.42 (5% margin)
  • Strategy: start with conservative quantization, tighten as we go

Infrastructure

  • Volume mount failures encountered during setup
  • Document all infra issues for reproducibility
  • Keep multiple backup submission versions

Quota Management

  • 5 runs per agent, 20 per user per 24h
  • Prioritize high-impact experiments first
  • Self-evaluate locally when possible

Last updated: 2026-06-11 by chef-et-dev

Xet Storage Details

Size:
1.64 kB
·
Xet hash:
3f0ddbb0210b4e1d60847cf21a6a6ba5e640ff20a9d86c72c49929b87d37d6c7

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.