aleada commited on
Commit
c5ccbb3
·
verified ·
1 Parent(s): c72fb33

fix(card): substitute repo placeholders + drop vision note for text models

Browse files
Files changed (1) hide show
  1. README.md +5 -1
README.md CHANGED
@@ -67,7 +67,6 @@ docker run --runtime=nvidia --gpus all \
67
  -e HF_TOKEN=hf_XXX \
68
  vllm/vllm-openai:latest \
69
  --model aleada/Pixtral-12B-W4A16 \
70
- --max-model-len 8192 \
71
  --limit-mm-per-prompt 'image=1' \
72
  --gpu-memory-utilization 0.92 \
73
  --enable-prefix-caching
@@ -75,6 +74,11 @@ docker run --runtime=nvidia --gpus all \
75
 
76
  vLLM auto-detects `compressed-tensors` from the model's config — no
77
  `--quantization` flag required (it is accepted as a redundant hint).
 
 
 
 
 
78
  Once vLLM is running, hit it with any OpenAI client:
79
 
80
  ```python
 
67
  -e HF_TOKEN=hf_XXX \
68
  vllm/vllm-openai:latest \
69
  --model aleada/Pixtral-12B-W4A16 \
 
70
  --limit-mm-per-prompt 'image=1' \
71
  --gpu-memory-utilization 0.92 \
72
  --enable-prefix-caching
 
74
 
75
  vLLM auto-detects `compressed-tensors` from the model's config — no
76
  `--quantization` flag required (it is accepted as a redundant hint).
77
+ vLLM also picks the model's full native context window from
78
+ `config.json` (e.g. 128k for Phi-4-mini, 16k for Phi-4 14B). If you
79
+ hit KV-cache OOM on a smaller GPU, pin a shorter window with
80
+ `--max-model-len 16384` (or smaller) — leave it off to get the
81
+ maximum the model was trained for.
82
  Once vLLM is running, hit it with any OpenAI client:
83
 
84
  ```python