Image-Text-to-Text
Transformers
Safetensors
English
qwen3_5
piko
piko-9b
multimodal
vision-language
hybrid-attention
linear-attention
ocr
document-understanding
conversational
Instructions to use Dexy2/Piko-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Dexy2/Piko-9b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Dexy2/Piko-9b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Dexy2/Piko-9b") model = AutoModelForMultimodalLM.from_pretrained("Dexy2/Piko-9b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Dexy2/Piko-9b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Dexy2/Piko-9b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Dexy2/Piko-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Dexy2/Piko-9b
- SGLang
How to use Dexy2/Piko-9b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Dexy2/Piko-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Dexy2/Piko-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Dexy2/Piko-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Dexy2/Piko-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Dexy2/Piko-9b with Docker Model Runner:
docker model run hf.co/Dexy2/Piko-9b
| # Piko-9b — Performance Report | |
| Every number here was measured on the hardware listed below. Nothing is extrapolated to other | |
| GPUs, other precisions, or other context lengths. Configurations that failed are recorded as | |
| failures, not omitted. | |
| Raw data: [`benchmarks/results/inference_4bit.json`](../benchmarks/results/inference_4bit.json), | |
| [`benchmarks/results/memory_4bit.json`](../benchmarks/results/memory_4bit.json). | |
| ## Test environment | |
| | | | | |
| |---|---| | |
| | GPU | NVIDIA GeForce RTX 5070 Ti, 15.92 GB usable | | |
| | CPU RAM | 33.9 GB | | |
| | OS | Windows 11 + WSL2 (Linux 6.6.87.2-microsoft-standard) | | |
| | Python | 3.12.3 | | |
| | torch | 2.10.0+cu128 | | |
| | transformers | 5.5.0 | | |
| | Precision | bfloat16 compute, **4-bit NF4** weights (bitsandbytes, double quant) | | |
| | Placement | `device_map={"": 0}` — fully resident, no CPU offload | | |
| | Decoding | greedy (`do_sample=False`) | | |
| | Linear-attention kernels | **not installed** — pure PyTorch fallback | | |
| | Weights location | internal NVMe | | |
| Because `flash-linear-attention` and `causal-conv1d` were not installed, every throughput figure | |
| below is a **floor**. Installing them should improve it; by how much was not measured. | |
| ## Cold load | |
| | Measurement | Value | | |
| |---|---| | |
| | Load time, NVMe | **101–119 s** | | |
| | Load time, external USB via WSL 9P | **10–25 min** (observed repeatedly) | | |
| | Resident VRAM after load | **7.37 GB** | | |
| | Host RSS before load | 0.55–0.81 GB | | |
| | Host RSS after load | **1.28 GB** | | |
| The 10× difference between NVMe and USB dominates every other cost in this report. Copy the | |
| checkpoint to internal storage before doing repeated work. | |
| `reported_parameters` from the loaded model is 5,724,972,272 — this is the **packed** element | |
| count under 4-bit quantization, not the logical parameter count. The true figure, read from the | |
| safetensors headers, is **9,653,104,368**. | |
| ## Latency and throughput | |
| `max_new_tokens=128`, greedy. TTFT is a full `generate()` capped at one new token; the decode rate | |
| excludes that first token. Both are measured after a warm-up pass. | |
| | Context | Batch | TTFT (s) | Prefill (tok/s) | Decode (tok/s) | Peak VRAM (GB) | | |
| |---:|---:|---:|---:|---:|---:| | |
| | 512 | 1 | 0.139 | 3,689 | **29.1** | 7.53 | | |
| | 512 | 2 | 0.230 | 4,461 | **42.8** | 7.62 | | |
| | 512 | 4 | 0.411 | 4,986 | **84.9** | 7.83 | | |
| | 2,048 | 1 | 0.407 | 5,036 | **35.8** | 7.75 | | |
| | 2,048 | 2 | 0.750 | 5,466 | **43.2** | 8.15 | | |
| | 2,048 | 4 | 1.470 | 5,576 | **82.6** | 8.95 | | |
| | 8,192 | 1 | 1.475 | 5,553 | **34.8** | 8.86 | | |
| | 8,192 | 2 | 2.925 | 5,601 | **42.4** | 10.38 | | |
| | 8,192 | 4 | 5.866 | 5,584 | **81.2** | 13.40 | | |
| All nine configurations completed. No failures. | |
| Three things worth noting: | |
| **Decode rate is nearly flat across context length.** 29–36 tok/s at batch 1 whether the prompt is | |
| 512 or 8,192 tokens. That is the hybrid architecture working as designed: 24 of 32 layers keep a | |
| fixed-size recurrent state, so only the 8 full-attention layers pay a growing KV cost. | |
| **Batching scales well.** Batch 4 gives ~2.4–2.9× the aggregate throughput of batch 1 at a modest | |
| memory cost. | |
| **Prefill throughput is stable at ~5,000–5,600 tok/s** once past the shortest prompts. | |
| ## Memory by context length | |
| Single forward pass with `use_cache=True`. "Added" is peak minus the resident weights. | |
| | Context | Peak VRAM (GB) | Added over weights (GB) | Prefill (s) | | |
| |---:|---:|---:|---:| | |
| | 512 | 7.63 | +0.29 | 0.805 | | |
| | 2,048 | 8.40 | +1.05 | 0.402 | | |
| | 8,192 | 11.48 | +4.13 | 1.590 | | |
| | 32,768 | **failed** | — | — | | |
| **32,768 tokens failed** on this GPU with `RuntimeError: CUDA driver error: device not ready` — an | |
| out-of-memory condition surfacing as a driver fault. Extrapolating the 8K figure (+4.13 GB) puts a | |
| 32K single forward pass well beyond the 15.92 GB available. | |
| This is a limit of **this GPU**, not of the model. It is also narrower than what generation can | |
| do: the custom suite's long-context cases at ~32,000 filler tokens **passed**, because `generate()` | |
| builds its cache incrementally rather than materialising a full-length forward pass at once. | |
| So, precisely: | |
| * 32K-token retrieval through `generate()`: **works** on 15.92 GB. | |
| * 32K-token single forward with full cache retention: **does not fit** on 15.92 GB. | |
| ## Image preprocessing | |
| | Measurement | Value | | |
| |---|---| | |
| | Median | **13 ms** | | |
| | Min / max over 5 repeats | 10.1 ms / 22.3 ms | | |
| | Prompt tokens for a 520×300 receipt | 214 | | |
| Negligible against generation time. A 520×300 document consumes ~200 tokens of context. | |
| Measured end-to-end image + text answers during validation: **~4 s**, against 0.9–2.9 s for a | |
| comparable text-only answer. | |
| ## Real-world latencies from the validation run | |
| Taken from [`reports/inference_validation.json`](inference_validation.json), 4-bit, NVMe: | |
| | Task | Time | | |
| |---|---| | |
| | Short text answer (48 tokens) | 0.85–1.5 s | | |
| | Code generation (96 tokens) | 2.9 s | | |
| | Image OCR / document JSON (128 tokens) | ~4.0–4.3 s | | |
| | Multi-turn follow-up (32 tokens) | 0.91 s | | |
| | Batch of 2, short answers | 1.19 s | | |
| | 14,429-token prompt + 32-token answer | 3.53 s | | |
| ## Configurations not measured | |
| | Configuration | Status | Why | | |
| |---|---|---| | |
| | bfloat16 unquantized | **Not run** | Needs ~22 GB; this GPU has 15.92 GB | | |
| | 8-bit | **Not run** | Would fit, but was not exercised | | |
| | Context 32K single forward | **Failed** | OOM on this GPU; recorded above | | |
| | Context beyond 32K | **Not run** | Hardware limit | | |
| | Multi-GPU | **Not run** | Splitting risks the same corruption as CPU offload | | |
| | CPU-only | **Not run** | ~38.6 GB float32 and no fast path; not viable | | |
| | With linear-attention kernels | **Not run** | All figures are therefore a floor | | |
| | vLLM / SGLang serving | **Not run** | Engine support for `qwen3_5` unverified | | |
| ## Reproducing | |
| ```bash | |
| python benchmarks/profile_inference.py \ | |
| --model <local-path> --quantization 4bit \ | |
| --context-lengths 512 2048 8192 --batch-sizes 1 2 4 --max-new-tokens 128 \ | |
| --image evaluation/custom_suite/assets/receipt.png \ | |
| --output benchmarks/results/inference_4bit.json | |
| python benchmarks/profile_memory.py \ | |
| --model <local-path> --quantization 4bit \ | |
| --context-lengths 512 2048 8192 32768 \ | |
| --output benchmarks/results/memory_4bit.json | |
| ``` | |
| Roughly 20 minutes total on comparable hardware, plus load time. | |
| ## A methodology note | |
| An earlier run of `profile_inference.py` reported decode rates of 33 and 164 billion tokens per | |
| second. The cause was measuring prefill as a bare forward pass and subtracting it from the full | |
| `generate()` time: when the subtraction went negative it was clamped to a floor, and the division | |
| exploded. The script now measures TTFT as `generate(max_new_tokens=1)`, warms up first, and | |
| reports `null` with a note when the decode window is too short to measure rather than emitting a | |
| number. The bad figures were never published; they are recorded here because a benchmark harness | |
| that can produce absurd values silently is worth documenting. | |