| # Launch vLLM using CLI | |
| See if we can move this snippet to the request.py | |
| ```bash | |
| module load anaconda3/2024.2 | |
| conda activate ecir2026 | |
| # Initialize vllm server | |
| MODEL=meta-llama/Llama-3.3-70B-Instruct | |
| NCCL_P2P_DISABLE=1 VLLM_SKIP_P2P_CHECK=1 vllm serve $MODEL \ | |
| --max-model-len 8192 \ | |
| --port 8000 \ | |
| --dtype bfloat16 \ | |
| --disable-custom-all-reduce \ | |
| --tensor-parallel-size 2 > vllm_server.log 2>&1 & | |
| PID=$! | |
| # Wait until server responds | |
| echo "Waiting for vLLM server (PID=$PID) to start..." | |
| until curl -s http://localhost:8000/v1/models >/dev/null; do | |
| sleep 10 | |
| done | |
| ``` | |
Xet Storage Details
- Size:
- 594 Bytes
- Xet hash:
- e5e96c13cf16b6184d835840cfff5592e3bb377f8a37146970f772aadf21b885
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.