Instructions to use suryatmodulus/apex-flash-1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use suryatmodulus/apex-flash-1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="suryatmodulus/apex-flash-1") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("suryatmodulus/apex-flash-1") model = AutoModelForMultimodalLM.from_pretrained("suryatmodulus/apex-flash-1", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use suryatmodulus/apex-flash-1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "suryatmodulus/apex-flash-1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "suryatmodulus/apex-flash-1", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/suryatmodulus/apex-flash-1
- SGLang
How to use suryatmodulus/apex-flash-1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "suryatmodulus/apex-flash-1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "suryatmodulus/apex-flash-1", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "suryatmodulus/apex-flash-1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "suryatmodulus/apex-flash-1", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use suryatmodulus/apex-flash-1 with Docker Model Runner:
docker model run hf.co/suryatmodulus/apex-flash-1
apex-flash-1
apex-flash-1 is Cantina Security's first open-weights security model, developed in partnership with Yeta (@yetalabs on X). It is a reinforcement learning fine-tune of GLM-5.3-Flash for focused investigations: reading code, using tools, pursuing an exploit, and verifying its effect in a running target.
Held-out case evaluation
We evaluated 60 tasks from 20 held-out vulnerability cases. Each case has guided whitebox, focused whitebox, and focused blackbox views. Targets ran in isolated environments, and verifiers checked the final target state. The table reports adjudicated first-draw pass@1.
| Model | Tasks solved | Pass@1 | Estimated cost for 60 tasks |
|---|---|---|---|
| Claude Opus 5 High | 43/60 | 71.7% | $74.68 (provider pricing) |
| apex-flash-1 | 40/60 | 66.7% | $2.38 |
| GLM-5.3-Flash | 36/60 | 60.0% | $4.56 (provider pricing) |
Capabilities
The checkpoint retains the base model's image-text-to-text architecture. Our reported evaluation covers text-based security tasks; we have not evaluated image or video performance.
Training and use
Training used production-like software and protocol environments with the Codex agent harness. The model is intended as a focused worker under a larger agent's direction. We recommend the Codex harness for this checkpoint.
What comes next
We are building harder, multi-step investigations as the model improves, broadening the data mix, and training across agent harnesses. We plan to publish more held-out and public benchmark results as they are validated. Explore Apex to see how this work supports production security.
Related links
- Apex Flash release post
- apex-flash-1-abliterated, an experimental derivative with modified refusal behavior. The evaluation above applies to the standard checkpoint.
- Downloads last month
- -
Model tree for suryatmodulus/apex-flash-1
Base model
zai-org/GLM-5.3-Flash