Examples
vLLM's examples are organized into the following categories:
basic/β Minimal examples for offline inference and online serving.generate/β Text generation examples, including multimodal models.pooling/β Examples for embedding, classification, scoring, reward, etc.speech_to_text/β Speech transcription, translation and real-time audio examples.features/β Demonstrations of individual vLLM features: automatic prefix caching, speculative decoding, LoRA, structured outputs, prompt embedding, pause/resume, batch invariance, KV events, data parallelism, and more.reasoning/β Examples for reasoning with vLLM.tool_calling/β Examples for function/tool calling with vLLM.applications/β Application examples such as chatbots and RAG (Retrieval-Augmented Generation).rl/β Reinforcement learning examples.deployment/β Examples for deploying vLLM in production.ray_serving/β Scalable serving using Ray.disaggregated/β Examples for disaggregated serving (separate prefill and decode), including various kv cache connectors (LMCache, Mooncake, FlexKV, P2P NCCL) and failure recovery.observability/β Metrics, logging, tracing (OpenTelemetry), and dashboards (Grafana, Perses).