Osaurus AI

OsaurusAI/Diplodocus-preview-JANG_6M

Diplodocus preview 3, the Osaurus assistant for finance, documents, spreadsheets and tool-assisted work, built on IFM/K2-Horizon-7B. This calibrated JANG affine mixed-precision bundle is approximately 6.65 GB. It contains the target model only; no DFlash or Uno drafter is included or required.

Osaurus 0.25.20 or newer is required. osaurus.json identifies this replacement as model version 2. Download the complete new bundle when upgrading from the former preview-2 JANGH7 repository.

Quantization

Component Format
Attention, embeddings and untied output head Affine 8-bit, group size 128
Layer 1 down projection Affine 8-bit, group size 64
Layers 3 and 5: gate, up and down projections Affine 5-bit, group size 128
Remaining MLP projections Affine 4-bit, group size 128
Norms BF16

Layer numbers are zero-based. The packed weights total 6,627,366,288 bytes (6.17 GiB); tokenizer and metadata bring the download to about 6.65 GB. All 254 quantized module layouts were cross-checked against their packed tensor shapes. Important projections retain higher precision; this is not a uniform 4-bit quant.

Calibration uses importance statistics, AWQ and GPTQ over 2,195,646 tokens, including approximately 25% generic text. The bundle uses MLX-compatible affine packing with per-module precision overrides. A plain uniform-bit MLX loader is insufficient: use Osaurus or a compatible JANG-aware vmlx-swift runtime. No matched quality or speed comparison against a separately built uniform MLX quant is claimed.

Reasoning and tools

  • High: reasoning on. Low: reasoning off, using the supplied template's empty prefilled reasoning block. Do not send medium.
  • Preserve native assistant reasoning and tool history across turns, and use the included tokenizer and chat template.
  • Tool calls use the native k2_horizon envelope. Osaurus handles its parsing.
  • Default sampling is temperature 1.0, top_p 0.95.
  • For money and exact arithmetic, provide and execute a calculator tool, then return its result to the model.

Validation and limits

The target was exercised in the real local vmlx-swift runtime with reasoning on/off, native tool calls and continuations. After removing speculative decoding, the bundle was checked again using the runtime's default strategy and an actual calculator-backed invoice workflow. See evaluation/target-only.json for the release checks.

An earlier paired local Release-build run measured target-only decoding at 83.63 tokens/s on one greedy invoice prompt on an M5 Max with 128 GB memory. This is a single-case measurement, not a general speed guarantee or an installed-app benchmark.

On a 128-row quantization holdout, the final allocation reduced weighted restricted-top-token KL from 0.11522 to 0.10690 relative to the initial preview-3 affine candidate. These are training-distribution quantization diagnostics, not an independent task-accuracy benchmark. Historical preview-2 evaluation numbers do not describe these weights.

Unaided arithmetic remains unreliable: one short invoice produced $283.70 from this quant and $283.68 from BF16; the calculator result is $283.71. The calculator-backed workflow returns the correct total. The model may identify itself as K2 or attempt an unavailable tool from older conversation history. Use current tool definitions and validate consequential outputs.

Version 2 changes

Replaces preview-2 JANGH7 with preview-3 calibrated affine JANG_6M weights, updates the tokenizer/template and reasoning toggle metadata, and ships without a speculative drafter. The repository and model ID now match the new quantization format.

한국어 안내

Diplodocus 프리뷰 3의 JANG 혼합 정밀도 affine 양자화 모델입니다. 크기는 약 6.65 GB이며, 어텐션·임베딩·출력 헤드는 8비트를 유지합니다. DFlash 또는 Uno 드래프터는 포함하지 않습니다. Osaurus 0.25.20 이상을 사용하고, 추론은 high로 켜거나 low로 끌 수 있습니다. 금액 계산에는 계산기 도구를 사용하세요.

Quantized and validated by Jinho Jang — eric@osaurus.ai

Downloads last month
48
Safetensors
Model size
9B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OsaurusAI/Diplodocus-preview-JANG_6M

Quantized
(40)
this model