Instructions to use OsaurusAI/Diplodocus-preview-JANG_6M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use OsaurusAI/Diplodocus-preview-JANG_6M with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("OsaurusAI/Diplodocus-preview-JANG_6M") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use OsaurusAI/Diplodocus-preview-JANG_6M with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "OsaurusAI/Diplodocus-preview-JANG_6M"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "OsaurusAI/Diplodocus-preview-JANG_6M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use OsaurusAI/Diplodocus-preview-JANG_6M with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "OsaurusAI/Diplodocus-preview-JANG_6M"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "OsaurusAI/Diplodocus-preview-JANG_6M" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OsaurusAI/Diplodocus-preview-JANG_6M", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use OsaurusAI/Diplodocus-preview-JANG_6M with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "OsaurusAI/Diplodocus-preview-JANG_6M"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default OsaurusAI/Diplodocus-preview-JANG_6M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use OsaurusAI/Diplodocus-preview-JANG_6M with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "OsaurusAI/Diplodocus-preview-JANG_6M"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "OsaurusAI/Diplodocus-preview-JANG_6M" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
OsaurusAI/Diplodocus-preview-JANG_6M
Diplodocus preview 3, the Osaurus assistant for finance, documents, spreadsheets and tool-assisted work, built on IFM/K2-Horizon-7B. This calibrated JANG affine mixed-precision bundle is approximately 6.65 GB. It contains the target model only; no DFlash or Uno drafter is included or required.
Osaurus 0.25.20 or newer is required. osaurus.json identifies this replacement as model version 2. Download the complete new bundle when upgrading from the former preview-2 JANGH7 repository.
Quantization
| Component | Format |
|---|---|
| Attention, embeddings and untied output head | Affine 8-bit, group size 128 |
| Layer 1 down projection | Affine 8-bit, group size 64 |
| Layers 3 and 5: gate, up and down projections | Affine 5-bit, group size 128 |
| Remaining MLP projections | Affine 4-bit, group size 128 |
| Norms | BF16 |
Layer numbers are zero-based. The packed weights total 6,627,366,288 bytes (6.17 GiB); tokenizer and metadata bring the download to about 6.65 GB. All 254 quantized module layouts were cross-checked against their packed tensor shapes. Important projections retain higher precision; this is not a uniform 4-bit quant.
Calibration uses importance statistics, AWQ and GPTQ over 2,195,646 tokens, including approximately 25% generic text. The bundle uses MLX-compatible affine packing with per-module precision overrides. A plain uniform-bit MLX loader is insufficient: use Osaurus or a compatible JANG-aware vmlx-swift runtime. No matched quality or speed comparison against a separately built uniform MLX quant is claimed.
Reasoning and tools
- High: reasoning on. Low: reasoning off, using the supplied template's empty prefilled reasoning block. Do not send
medium. - Preserve native assistant reasoning and tool history across turns, and use the included tokenizer and chat template.
- Tool calls use the native
k2_horizonenvelope. Osaurus handles its parsing. - Default sampling is temperature 1.0, top_p 0.95.
- For money and exact arithmetic, provide and execute a calculator tool, then return its result to the model.
Validation and limits
The target was exercised in the real local vmlx-swift runtime with reasoning on/off, native tool calls and continuations. After removing speculative decoding, the bundle was checked again using the runtime's default strategy and an actual calculator-backed invoice workflow. See evaluation/target-only.json for the release checks.
An earlier paired local Release-build run measured target-only decoding at 83.63 tokens/s on one greedy invoice prompt on an M5 Max with 128 GB memory. This is a single-case measurement, not a general speed guarantee or an installed-app benchmark.
On a 128-row quantization holdout, the final allocation reduced weighted restricted-top-token KL from 0.11522 to 0.10690 relative to the initial preview-3 affine candidate. These are training-distribution quantization diagnostics, not an independent task-accuracy benchmark. Historical preview-2 evaluation numbers do not describe these weights.
Unaided arithmetic remains unreliable: one short invoice produced $283.70 from this quant and $283.68 from BF16; the calculator result is $283.71. The calculator-backed workflow returns the correct total. The model may identify itself as K2 or attempt an unavailable tool from older conversation history. Use current tool definitions and validate consequential outputs.
Version 2 changes
Replaces preview-2 JANGH7 with preview-3 calibrated affine JANG_6M weights, updates the tokenizer/template and reasoning toggle metadata, and ships without a speculative drafter. The repository and model ID now match the new quantization format.
한국어 안내
Diplodocus 프리뷰 3의 JANG 혼합 정밀도 affine 양자화 모델입니다. 크기는 약 6.65 GB이며, 어텐션·임베딩·출력 헤드는 8비트를 유지합니다. DFlash 또는 Uno 드래프터는 포함하지 않습니다. Osaurus 0.25.20 이상을 사용하고, 추론은 high로 켜거나 low로 끌 수 있습니다. 금액 계산에는 계산기 도구를 사용하세요.
Quantized and validated by Jinho Jang — eric@osaurus.ai
- Downloads last month
- 48
4-bit
Model tree for OsaurusAI/Diplodocus-preview-JANG_6M
Base model
IFM/K2-Horizon-7B