Image-Text-to-Text
Transformers
Safetensors
English
qwen3_5
piko
piko-9b
multimodal
vision-language
hybrid-attention
linear-attention
ocr
document-understanding
conversational
Instructions to use Dexy2/Piko-9b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Dexy2/Piko-9b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Dexy2/Piko-9b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Dexy2/Piko-9b") model = AutoModelForMultimodalLM.from_pretrained("Dexy2/Piko-9b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Dexy2/Piko-9b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Dexy2/Piko-9b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Dexy2/Piko-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Dexy2/Piko-9b
- SGLang
How to use Dexy2/Piko-9b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Dexy2/Piko-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Dexy2/Piko-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Dexy2/Piko-9b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Dexy2/Piko-9b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Dexy2/Piko-9b with Docker Model Runner:
docker model run hf.co/Dexy2/Piko-9b
| # Installation | |
| ## Requirements | |
| | Component | Minimum | Notes | | |
| |---|---|---| | |
| | Python | 3.10 – 3.12 | 3.13+ works if wheels exist for your torch build; 3.15 currently has no `torchvision` wheel | | |
| | `transformers` | **5.5** | `AutoModelForMultimodalLM` does not exist in 4.x | | |
| | `torch` | 2.6 | CUDA build; validated on 2.10.0+cu128 | | |
| | `torchvision` | any matching build | **Mandatory** — `AutoProcessor` fails to construct without it | | |
| | `accelerate` | 0.30 | device placement | | |
| | `bitsandbytes` | 0.43 | only for 4-bit / 8-bit | | |
| | `pillow` | 10.0 | image input | | |
| | GPU | 8 GB (4-bit) / 22 GB (bf16) | CUDA required; see [hardware.md](hardware.md) | | |
| `trust_remote_code` is **not** required. The repository ships no Python files. | |
| ## Quick install | |
| ```bash | |
| python -m venv .venv | |
| source .venv/bin/activate # Windows: .venv\Scripts\activate | |
| pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128 | |
| pip install -r requirements.txt | |
| ``` | |
| For evaluation and development: | |
| ```bash | |
| pip install -r requirements-dev.txt | |
| ``` | |
| ## Verifying the install | |
| ```bash | |
| python - <<'PY' | |
| import torch, transformers, torchvision | |
| assert tuple(int(x) for x in transformers.__version__.split(".")[:2]) >= (5, 5), transformers.__version__ | |
| print("torch", torch.__version__, "cuda", torch.cuda.is_available()) | |
| print("transformers", transformers.__version__) | |
| print("torchvision", torchvision.__version__) | |
| print("gpu", torch.cuda.get_device_name(0) if torch.cuda.is_available() else "NONE") | |
| PY | |
| ``` | |
| All four lines must print, and `cuda` must be `True`. CPU-only inference is not a supported | |
| configuration for this model — see [hardware.md](hardware.md). | |
| ## Getting the weights | |
| ### From the Hub | |
| ```python | |
| from transformers import AutoModelForMultimodalLM | |
| model = AutoModelForMultimodalLM.from_pretrained("Dexy2/Piko-9b") | |
| ``` | |
| Pin a revision for reproducible work: | |
| ```python | |
| model = AutoModelForMultimodalLM.from_pretrained("Dexy2/Piko-9b", revision="<commit-sha>") | |
| ``` | |
| The repository is public, so no authentication is needed. If you are behind a proxy or working | |
| with a private mirror: | |
| ```bash | |
| hf auth login | |
| ``` | |
| ### Downloading ahead of time | |
| ```bash | |
| hf download Dexy2/Piko-9b --local-dir ./piko-9b | |
| ``` | |
| ≈ 21 GB across 11 safetensors shards, plus a 20 MB tokenizer. | |
| ### From a local directory | |
| Every script in this repository accepts a path anywhere a repo id is accepted: | |
| ```bash | |
| python examples/inference_transformers.py --model ./piko-9b --prompt "Hello" | |
| ``` | |
| > Load the checkpoint from **internal NVMe**. Loading 21 GB from an external USB disk is I/O | |
| > bound and takes 10–20 minutes per load; from NVMe it takes seconds. | |
| ## CUDA compatibility | |
| | torch build | Driver | Status | | |
| |---|---|---| | |
| | `2.10.0+cu128` | ≥ 525 | Validated for every result in this repository | | |
| | `cu121` / `cu124` builds | ≥ 525 | Expected to work; not tested here | | |
| | ROCm | — | Not tested | | |
| | CPU-only | — | Loads, but see [hardware.md](hardware.md) before trying | | |
| Blackwell cards (RTX 50-series) need a cu128 or newer build. | |
| ## Optional: linear-attention kernels | |
| ```bash | |
| pip install flash-linear-attention causal-conv1d | |
| ``` | |
| 24 of the 32 layers are linear-attention. Without these kernels `transformers` logs *"The fast | |
| path is not available"* and falls back to pure PyTorch — correct, but slower. Every measurement in | |
| this repository was taken **without** these kernels, so treat published throughput as a floor. | |
| ## Reproducible environment | |
| ```bash | |
| pip install -r requirements-lock.txt # exact versions used for the published results | |
| ``` | |
| If that file is absent, the environment behind every measured number is recorded in the | |
| `environment` block of each JSON file under `evaluation/results/` and `benchmarks/results/`. | |