nimble / README.md
mertcobanov's picture
Ollaya package for bespokelabs/Bespoke-Nimble-9B-v2
2d630ba verified
|
Raw History Blame Contribute Delete
2.02 kB
metadata
license: apache-2.0
base_model:
  - bespokelabs/Bespoke-Nimble-9B-v2
  - Qwen/Qwen3.5-9B
library_name: onnx
tags:
  - ollaya
  - onnx
  - decision-model
  - system-one
pipeline_tag: text-classification

nimble for Ollaya

Ollaya package of bespokelabs/Bespoke-Nimble-9B-v2 and Qwen/Qwen3.5-9B by Bespoke Labs (adapter) and the Qwen team (base model). Ollaya runs open decision models locally, the way Ollama runs LLMs: typed questions in, calibrated answers out, behind a TypeSafe-compatible API.

ollaya run nimble

What is in this repository

This repository holds only the files Ollaya derives, with no weights. Each graph is an ONNX export of the original model whose weights reference the authors' own weight files by byte offset, so ollaya pull downloads the weights from the upstream repositories, unmodified and pinned to a commit, and verifies their sha256.

Tag Upstream Files
nimble:9b bespokelabs/Bespoke-Nimble-9B-v2@4b8c04d, Qwen/Qwen3.5-9B@c202236 9b/model-fp32.onnx, 9b/decision.json, 9b/calibration.json

Each tag has an fp32 graph, used on CPU and GPU. Each tag also has decision.json (sequence layout, special tokens) and calibration.json (temperatures).

Parity

Ollaya's Rust runtime matches the author's reference code (serving_schema.prepare_prompts and inference.candidate_logits, PyTorch fp32 with the LoRA unmerged) on 492 questions from 104 requests: identical token rows, the same 4 rejected requests, the same decision on every question, option logits within 1.1e-4 and probabilities within 6.5e-6 (CUDA).

License

Same as the upstream model (Apache-2.0). Ollaya itself is Apache-2.0.