Ctrl+K
- claim-1-realtimetool-achieves-3-6x-end-to-end-speedup-up-to-9-6x-with-only-8-2-parallelization-overhead-on-function-calling-tasks
- claim-2-on-mobile-actions-rt-qwen-0-5b-outperforms-google-s-functiongemma-in-both-accuracy-and-latency-consistency
- claim-2-on-the-rtx-4090-qwen2-5-0-5b-attains-5-35x-speedup-under-transformers-and-3-07x-under-vllm-while-qwen3-4b-attains-2-83x-transformers-and-3-88x-vllm-table-3
- claim-3-achieves-61-2ms-p50-latency-on-consumer-grade-gpu-with-quantization-enabling-16-hz-real-time-control-at-4b-model-scale
- claim-5-with-8-head-parallel-decoding-the-framework-demonstrates-93-0-average-tpot-time-per-output-token-efficiency-section-2-1-3-appendix-g
- conclusion
- executive-summary
- 1.67 kB