Instructions to use Cloudflare/clef-flash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Cloudflare/clef-flash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Cloudflare/clef-flash") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Cloudflare/clef-flash") model = AutoModelForMultimodalLM.from_pretrained("Cloudflare/clef-flash", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Cloudflare/clef-flash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Cloudflare/clef-flash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Cloudflare/clef-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Cloudflare/clef-flash
- SGLang
How to use Cloudflare/clef-flash with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Cloudflare/clef-flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Cloudflare/clef-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Cloudflare/clef-flash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Cloudflare/clef-flash", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Cloudflare/clef-flash with Docker Model Runner:
docker model run hf.co/Cloudflare/clef-flash
Support / Roadmap for serving Clef-Flash via vLLM?
Hi @Cloudflare team,
Thanks for releasing Clef-Flash! The joint schema evaluation approach and single forward-pass probability output design are exceptional for high-throughput decision and classification pipelines.
Since standard deployment in production environments relies heavily on vLLM (for PagedAttention, continuous batching, and high concurrency), I wanted to check on vLLM compatibility:
- Custom Architecture Support: Since
clef-flashrelies on custom code (joint_schema_model.py/systemone) and evaluates typed questions via direct hidden-state readouts rather than token-by-token autoregressive generation, is there any current pathway or plan to registerClefForConditionalGeneration/ its joint classification heads as an out-of-the-box model architecture in vLLM? - Forward-Pass / Embedding API: In vLLM, classification and embedding models typically use the pooling/forward-pass endpoints rather than
/v1/completions. Has the team explored integratingsystemoneschema scoring on top of vLLM’s forward/pooling engine? - Current Serving Recommendations: Outside of running native Hugging Face Transformers + FastAPI/Triton, what is your recommended high-throughput serving stack for serving Clef-Flash at scale?
Appreciate any guidance or pointers from the team or community!
@Vishva007 If stock vLLM support is the immediate requirement, there is another implementation approach: score the next-token logits restricted to the decision's option keys, rather than serving a custom joint head. I built Seb-9B around that approach: https://huggingface.co/ironbcc/seb-9b
The card includes a vLLM chat API example with constrained choices and option logprobs. It handles one question per request, with up to 20 choice options; that differs from Clef's joint multi-question design. Released-checkpoint GPU throughput measurements are still pending, so I can't claim it solves your high-concurrency requirement yet. It may be useful as a stock-vLLM baseline while the Clef integration question is worked through.