Instructions to use ToTo-40417/EXLLM-ONI5M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ToTo-40417/EXLLM-ONI5M with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM-ONI5M:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM-ONI5M:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM-ONI5M:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM-ONI5M:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ToTo-40417/EXLLM-ONI5M:F16 # Run inference directly in the terminal: ./llama-cli -hf ToTo-40417/EXLLM-ONI5M:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ToTo-40417/EXLLM-ONI5M:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ToTo-40417/EXLLM-ONI5M:F16
Use Docker
docker model run hf.co/ToTo-40417/EXLLM-ONI5M:F16
- LM Studio
- Jan
- vLLM
How to use ToTo-40417/EXLLM-ONI5M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ToTo-40417/EXLLM-ONI5M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ToTo-40417/EXLLM-ONI5M", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ToTo-40417/EXLLM-ONI5M:F16
- Ollama
How to use ToTo-40417/EXLLM-ONI5M with Ollama:
ollama run hf.co/ToTo-40417/EXLLM-ONI5M:F16
- Unsloth Desktop
- Docker Model Runner
How to use ToTo-40417/EXLLM-ONI5M with Docker Model Runner:
docker model run hf.co/ToTo-40417/EXLLM-ONI5M:F16
- Lemonade
How to use ToTo-40417/EXLLM-ONI5M with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ToTo-40417/EXLLM-ONI5M:F16
Run and chat with the model
lemonade run user.EXLLM-ONI5M-F16
List all available models
lemonade list
- Atomic Chat
EXLLM-ONI5M
EXLLM-ONI5M is an experimental 0.005377824B-parameter Little Language Model trained to produce short Japanese replies in a deliberately impatient, old-school teacher persona. It is a playful, narrow template-based model—not a general assistant or a reliable source of facts.
EXQ12 inference has been verified on a CASIO EX-word XD-B4800 (DATAPLUS 6). The included GGUF is a separately trained Llama-compatible companion, not a conversion or quantization of the embedded checkpoint.
Highlights
- 0.005377824B-parameter embedded checkpoint with a 128-token context and 868-token vocabulary
- Short kana-oriented prompts and an intentionally gruff fictional persona
- Physical-device inference verified at approximately 0.57 token/s in three recorded XD-B4800 runs
- Separately trained 0.005441184B-parameter GGUF companion for LM Studio and llama.cpp
- Project-authored templates and synthetic combinations; no external pretrained checkpoint
- Apache-2.0
Intended Behavior
The training objective favors a short answer followed by a teacher-like admonition. Some unknown-input examples ask the user to rephrase, but this is learned text behavior—not a rejection mechanism or a guarantee against hallucination.
The persona is fictional. It may produce rude or inaccurate text and has not been evaluated as a safety system. It is intended for embedded inference experiments and entertainment, not factual, professional, or safety-critical use.
Physical EX-word Results
Readback SHA-256 verification, model loading, inference, screenshots, and PERFLOG capture were completed on an XD-B4800. The three recorded runs decoded at approximately 0.57 token/s. One prompt containing kanji did not produce the expected identity response and instead reached the 40-token limit, so these measurements demonstrate execution rather than general response quality.
The machine-readable observations are in benchmarks/xd-b4800-20260930.json.
Usage
LM Studio / llama.cpp
EXLLM-ONI5M-F16.gguf was separately initialized and trained from the ONI5M public training data. It has different weights, tokenizer, architecture details, and runtime context from the EX-word checkpoint.
llama-cli -m EXLLM-ONI5M-F16.gguf -c 512 --single-turn -p "こんにちは"
An RTX 3060 llama.cpp smoke test observed approximately 443 token/s and produced:
話を聞け!あいさつはいい!質問を言え! 一つずつ確実にやれ!
This is a PC companion smoke-test measurement, not a directly comparable benchmark against the EX-word model. Training details and hashes are in lmstudio/training_manifest.json.
CASIO EX-word
Use exllm-exword v1.3.1 or later with weights/oni5m.q12. The file can be checked before deployment with exllm-model-check.
Training and Provenance
The embedded checkpoint was trained from random initialization for 18 epochs, recording 6,692,919 processed tokens and 4,532,281 loss-bearing tokens. Its data consists of project-authored templates and synthetic combinations. No third-party pretrained checkpoint or real-person quotation corpus was used.
See DATA_PROVENANCE.md, training/RECIPE.md, and release-manifest.json for the release boundary and reproducibility details.
Release Artifacts
| Artifact | Size | SHA-256 | Purpose |
|---|---|---|---|
weights/EXLLM-ONI5M-v1.0.pt |
21,528,721 bytes (20.53 MiB; 0.021528721 GB) | 661857ffc19dbca63f738173d1ef5e6c36e110413028fe88e0e24ad0799d6e7b |
FP32 PyTorch checkpoint |
weights/EXLLM-ONI5M-v1.0-int8.bin |
5,443,105 bytes (5.19 MiB; 0.005443105 GB) | 5a6bb0b7c9b15bfcda2cdd50c734b428488e9994a53ef47f9fa49653cf9b9adf |
EXLLM8 artifact |
weights/oni5m.q12 |
5,443,105 bytes (5.19 MiB; 0.005443105 GB) | 64c8ce542fc251e26e9e10d756ef0bd33cbdf9e0d7336d1a7048538aca42c5ee |
EX-word EXQ12 artifact |
EXLLM-ONI5M-F16.gguf |
10,918,368 bytes (10.41 MiB; 0.010918368 GB) | 6c78fa17860062730bdcfa45835528dc6372b47c0fc28667bd0bc9e4d18483a8 |
Separately trained LM Studio / llama.cpp companion |
tokenizer.json |
6,864 bytes | 4c2a583b554e8165936cb3b11727f79d2ede43b73d416a48681f8d38385f5964 |
Embedded-model tokenizer |
Limitations
- This is a narrow persona experiment trained on finite templates and synthetic combinations.
- It may answer incorrectly, ignore the intended persona, repeat text, or consume the full generation limit.
- Unknown-input rephrasing is not guaranteed and does not prevent hallucination.
- Kana variation, typos, kanji, long prompts, arithmetic, and open-ended conversation are unreliable.
- The embedded model context is 128 tokens; the GGUF companion's 512-token runtime allocation does not establish 512-token task quality.
- Physical-device and PC companion results belong to different models and numerical runtimes.
- Do not use it for medical, legal, financial, safety-critical, or other professional decisions.
Links
- Shared architecture, conversion, and evaluation project: ToTo-40417/exllm
- EX-word runtime: ToTo-40417/exllm-exword
- Related Japanese-to-English model: ToTo-40417/EXLLM-JPTOEN
- Base EXLLM model: ToTo-40417/EXLLM
日本語
EXLLM-ONI5Mは、短気な昔気質の先生を演じる、5,377,824(0.005377824B)パラメータの遊び寄りのLittle Language Modelです。短いかな中心の入力に対し、短い回答と小言を返すよう学習していますが、汎用会話モデルでも安全機構でもありません。誤答や想定外の応答も発生します。
CASIO XD-B4800(DATAPLUS 6)でEXQ12推論を確認しています。同梱GGUFはEX-word用重みの変換ではなく、公開学習データから別途学習したPC用companionです。詳しい来歴、実行方法、制限事項は上の英語本文と付属資料を参照してください。
- Downloads last month
- 141
16-bit