Instructions to use ToTo-40417/EXLLM-JPTOEN with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ToTo-40417/EXLLM-JPTOEN with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM-JPTOEN:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM-JPTOEN:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ToTo-40417/EXLLM-JPTOEN:F16 # Run inference directly in the terminal: llama cli -hf ToTo-40417/EXLLM-JPTOEN:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ToTo-40417/EXLLM-JPTOEN:F16 # Run inference directly in the terminal: ./llama-cli -hf ToTo-40417/EXLLM-JPTOEN:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ToTo-40417/EXLLM-JPTOEN:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ToTo-40417/EXLLM-JPTOEN:F16
Use Docker
docker model run hf.co/ToTo-40417/EXLLM-JPTOEN:F16
- LM Studio
- Jan
- vLLM
How to use ToTo-40417/EXLLM-JPTOEN with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ToTo-40417/EXLLM-JPTOEN" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ToTo-40417/EXLLM-JPTOEN", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ToTo-40417/EXLLM-JPTOEN:F16
- Ollama
How to use ToTo-40417/EXLLM-JPTOEN with Ollama:
ollama run hf.co/ToTo-40417/EXLLM-JPTOEN:F16
- Unsloth Desktop
- Docker Model Runner
How to use ToTo-40417/EXLLM-JPTOEN with Docker Model Runner:
docker model run hf.co/ToTo-40417/EXLLM-JPTOEN:F16
- Lemonade
How to use ToTo-40417/EXLLM-JPTOEN with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ToTo-40417/EXLLM-JPTOEN:F16
Run and chat with the model
lemonade run user.EXLLM-JPTOEN-F16
List all available models
lemonade list
- Atomic Chat
EXLLM-JPTOEN
EXLLM-JPTOEN is an experimental 0.005377824B-parameter Little Language Model for generating short English expressions from kana-oriented Japanese input. EXQ12 integer inference has been verified on a CASIO EX-word XD-B4800 (DATAPLUS 6).
This is a narrow task-specific model centered on 266 lexical entries, not a general translation model. The included GGUF is a separately trained Llama-compatible companion—not a conversion or quantization of the embedded EXLLM checkpoint.
Highlights
- 0.005377824B parameters; 0.000128M-token context; 0.000868M-token vocabulary
- 0.100003620B processed non-padding training tokens, including extensive repetition and template-derived examples
- EX-word integer inference verified at approximately 0.57–0.58 token/s in the recorded XD-B4800 runs
- Project-specific evaluation: 266/266 bare known words, but only 34/266 fully unseen query templates
- Separately trained 0.005441184B-parameter GGUF companion for LM Studio and llama.cpp
- Apache-2.0
Model Details
| Field | Value |
|---|---|
| Model type | decoder-only Transformer |
| Parameters | 5,377,824 (0.005377824B) |
| Layers | 6 |
| Hidden size | 288 |
| Attention heads | 9 |
| FFN size | 896 |
| Context | 128 tokens (0.000128M) |
| Vocabulary | 868 tokens (0.000868M) |
| Tokenizer | frequent Japanese characters + UTF-8 byte fallback, NFC |
| Embeddings | tied token embedding and output head |
| License | Apache-2.0 |
The parameter count excludes duplication from tied weights.
Training and Provenance
EXLLM-JPTOEN continues from an unpublished in-project checkpoint of the same architecture. No external pretrained checkpoint was used. The final checkpoint records 100,003,620 (0.100003620B) processed non-padding input tokens, including 97,102,074 continuation tokens over 11,488 steps with seed 40417001.
This count is not unique corpus size. Training repeatedly sampled deterministic examples derived from a project-authored 266-row lexicon, finite question and noise templates, and unknown-calibration lists. Most processed tokens are repeated or template-derived. The project owner assembled and edited these resources with generative-AI assistance; no third-party dictionary corpus or web crawl is recorded in this checkpoint lineage.
Exact hashes, token-count definitions, data rows, training mix, and lineage are recorded in DATA_PROVENANCE.md, the GitHub provenance document, and the machine-readable case-study manifest.
Evaluation
These are project-specific greedy-decoding regression suites, not standard translation benchmarks.
| Suite | Result |
|---|---|
| Bare known words | 266 / 266 |
| Learned-family query forms | 262 / 266 |
| Fully unseen query templates | 34 / 266 |
| Real-word unknown holdout | 22 / 40 |
| Ambiguous unknown | 12 / 12 |
| Nonce unknown | 62 / 64 |
| General OOD | 5 / 5 |
| EX-word regression set | 6 / 6 |
| Visible ASCII output | 925 / 925 |
FP32 and dequantized INT8 produced the same results on these suites. The 34/266 fully unseen-template result is the clearest limitation: 0.100003620B processed tokens did not establish broad grammatical or general translation ability.
Real-word unknown holdout measures rejection of words outside the registered vocabulary contract, not translation accuracy. Semantically correct outputs such as たまご → egg or にく → meat therefore count as failures in that suite. Full cases and input hashes are available in common-eval-int8.json.
Physical EX-word Results
On the XD-B4800, model identification, loading, inference, and log capture were verified from both internal storage and the microSD registration cache. Recorded decode throughput was approximately 0.57–0.58 token/s. Correct answers, appropriate unknown responses, and incorrect English outputs all occurred; for example, ぎんこう produced temperature.
The complete five-prompt timing record and screenshots are referenced by the fixed-5M case study. Host-side regression results must not be interpreted as general free-form translation quality.
Usage
LM Studio / llama.cpp
EXLLM-JPTOEN-F16.gguf is a separately trained Llama-compatible companion created from the same JPTOEN data contract. It has different weights, tokenizer, architecture details, and runtime context from the embedded checkpoint.
llama-cli \
-m EXLLM-JPTOEN-F16.gguf \
-c 512 --single-turn \
-p "おはようをえいごでいうと?"
An RTX 3060 llama.cpp smoke test generated good morning at approximately 415 token/s. This is a PC companion smoke-test measurement, not a directly comparable benchmark against the EX-word model. Its training record is in lmstudio/training_manifest.json and lmstudio/dataset-manifest.json.
PyTorch reference runtime
git clone https://github.com/ToTo-40417/exllm
cd exllm
python -m venv .venv
. .venv/bin/activate
pip install -r requirements.txt
Load weights/EXLLM-JPTOEN.pt with this repository's config.json and tokenizer.json through the EXLLM reference runtime.
CASIO EX-word
Use exllm-exword v1.3.1 or later with weights/EXLLM-JPTOEN.q12. The file can be checked before deployment with exllm-model-check.
Release Artifacts
| Artifact | Size | SHA-256 | Purpose |
|---|---|---|---|
EXLLM-JPTOEN.pt |
21,536,780 bytes (20.54 MiB; 0.021536780 GB) | 942ec9f1610a5566585a7597a1bac67b4128814532669dffc334c6629bd21fe1 |
FP32 PyTorch checkpoint |
EXLLM-JPTOEN-int8.bin |
5,443,105 bytes (5.19 MiB; 0.005443105 GB) | 6f0d77b498e5e9f3bda8480832b6afdc66306205ab4fbdae2911203f935a0675 |
EXLLM8 INT8 artifact |
EXLLM-JPTOEN.q12 |
5,443,105 bytes (5.19 MiB; 0.005443105 GB) | ea5b59785e00b2a28b3c4eaa1e24785238f64ba3d45e300da7510aabf5d284eb |
EX-word EXQ12 artifact |
EXLLM-JPTOEN-F16.gguf |
10,916,864 bytes (10.41 MiB; 0.010916864 GB) | 362a237ff01fec440472ceb239fe65b41195e971b55e2fd8609fae3b467f053c |
Separately trained LM Studio / llama.cpp companion |
tokenizer.json |
6,864 bytes | 4c2a583b554e8165936cb3b11727f79d2ede43b73d416a48681f8d38385f5964 |
Embedded-model tokenizer |
EXLLM8 and EXQ12 have the same file size but are not interchangeable formats.
Intended Use
- registered kana vocabulary to short English expressions;
- prompt forms close to learned templates;
- narrow unknown-rejection experiments;
- EX-word and extremely small task-specific LM research.
Limitations
- The task is centered on 266 lexical entries.
- Unseen syntax, long-form text, free translation, and open-ended conversation are unreliable.
- Unknown calibration is imperfect; the model may emit a known but unrelated English word.
- Robustness to kana spelling variation, long vowels, small kana, and typos is not guaranteed.
- The embedded model context is 128 tokens.
- Results come from a single continuation seed; no multi-seed comparison was performed.
- Host evaluation and the SH-4A fixed/integer runtime use different numerical paths.
- Do not use it for medical, legal, financial, safety-critical, or professional translation.
- The custom embedded checkpoint is not directly supported by Hugging Face Inference Providers.
Links
- Source, architecture, training, evaluation, and conversion: ToTo-40417/exllm
- EX-word runtime: ToTo-40417/exllm-exword
- EXQ12 validator: ToTo-40417/exllm-model-check
- Base EXLLM model: ToTo-40417/EXLLM
- Data provenance:
EXLLM_JPTOEN_DATA_PROVENANCE.md - Detailed case study:
fixed-5m-training-and-generalization.md - Project naming policy:
NAMING_POLICY.md
日本語
EXLLM-JPTOENは、かな中心の日本語入力から短い英語表現を生成する、5,377,824(0.005377824B)パラメータの実験的なLittle Language Modelです。266語を中心とする狭いモデルであり、一般翻訳モデルではありません。CASIO XD-B4800(DATAPLUS 6)上でEXQ12整数推論を確認しています。
学習量は100,003,620(0.100003620B)processed non-padding tokensですが、その大部分は反復または有限templateからの派生例です。未学習の質問形式は34/266に留まっており、広い文法理解や一般翻訳能力を示す結果ではありません。
同梱GGUFはEX-word用重みの変換ではなく、同じJPTOENデータ契約から別途学習したPC用companionです。詳しい評価、来歴、実行方法、制限事項は上の英語本文と関連リンクを参照してください。
Citation
@software{toto_exllm_jptoen_2026,
author = {ToTo},
title = {EXLLM-JPTOEN},
year = {2026},
url = {https://huggingface.co/ToTo-40417/EXLLM-JPTOEN}
}
- Downloads last month
- 149
16-bit