# Installation ## Requirements | Component | Minimum | Notes | |---|---|---| | Python | 3.10 – 3.12 | 3.13+ works if wheels exist for your torch build; 3.15 currently has no `torchvision` wheel | | `transformers` | **5.5** | `AutoModelForMultimodalLM` does not exist in 4.x | | `torch` | 2.6 | CUDA build; validated on 2.10.0+cu128 | | `torchvision` | any matching build | **Mandatory** — `AutoProcessor` fails to construct without it | | `accelerate` | 0.30 | device placement | | `bitsandbytes` | 0.43 | only for 4-bit / 8-bit | | `pillow` | 10.0 | image input | | GPU | 8 GB (4-bit) / 22 GB (bf16) | CUDA required; see [hardware.md](hardware.md) | `trust_remote_code` is **not** required. The repository ships no Python files. ## Quick install ```bash python -m venv .venv source .venv/bin/activate # Windows: .venv\Scripts\activate pip install torch torchvision --index-url https://download.pytorch.org/whl/cu128 pip install -r requirements.txt ``` For evaluation and development: ```bash pip install -r requirements-dev.txt ``` ## Verifying the install ```bash python - <<'PY' import torch, transformers, torchvision assert tuple(int(x) for x in transformers.__version__.split(".")[:2]) >= (5, 5), transformers.__version__ print("torch", torch.__version__, "cuda", torch.cuda.is_available()) print("transformers", transformers.__version__) print("torchvision", torchvision.__version__) print("gpu", torch.cuda.get_device_name(0) if torch.cuda.is_available() else "NONE") PY ``` All four lines must print, and `cuda` must be `True`. CPU-only inference is not a supported configuration for this model — see [hardware.md](hardware.md). ## Getting the weights ### From the Hub ```python from transformers import AutoModelForMultimodalLM model = AutoModelForMultimodalLM.from_pretrained("Dexy2/Piko-9b") ``` Pin a revision for reproducible work: ```python model = AutoModelForMultimodalLM.from_pretrained("Dexy2/Piko-9b", revision="") ``` The repository is public, so no authentication is needed. If you are behind a proxy or working with a private mirror: ```bash hf auth login ``` ### Downloading ahead of time ```bash hf download Dexy2/Piko-9b --local-dir ./piko-9b ``` ≈ 21 GB across 11 safetensors shards, plus a 20 MB tokenizer. ### From a local directory Every script in this repository accepts a path anywhere a repo id is accepted: ```bash python examples/inference_transformers.py --model ./piko-9b --prompt "Hello" ``` > Load the checkpoint from **internal NVMe**. Loading 21 GB from an external USB disk is I/O > bound and takes 10–20 minutes per load; from NVMe it takes seconds. ## CUDA compatibility | torch build | Driver | Status | |---|---|---| | `2.10.0+cu128` | ≥ 525 | Validated for every result in this repository | | `cu121` / `cu124` builds | ≥ 525 | Expected to work; not tested here | | ROCm | — | Not tested | | CPU-only | — | Loads, but see [hardware.md](hardware.md) before trying | Blackwell cards (RTX 50-series) need a cu128 or newer build. ## Optional: linear-attention kernels ```bash pip install flash-linear-attention causal-conv1d ``` 24 of the 32 layers are linear-attention. Without these kernels `transformers` logs *"The fast path is not available"* and falls back to pure PyTorch — correct, but slower. Every measurement in this repository was taken **without** these kernels, so treat published throughput as a floor. ## Reproducible environment ```bash pip install -r requirements-lock.txt # exact versions used for the published results ``` If that file is absent, the environment behind every measured number is recorded in the `environment` block of each JSON file under `evaluation/results/` and `benchmarks/results/`.