VimeML Tiny Japanese GPT v2.1
A 12,537,920-parameter Japanese decoder-only language model for local contextual candidate reranking and short within-sentence continuation. This release contains the selected V2.1 extend5 best / step40000 FP32 inference bundle and its INT8 block32, FP32-compute Core ML variant. V2 experiments completed on 2026-10-08; the current quantization loss and observed input latency are accepted for this deployment.
This is a custom PyTorch model, not a Transformers AutoModel integration or an instruction/chat model. SentencePiece, kana retrieval, candidate scoring and search run outside the Core ML graph. The original architecture, loader and scoring code are included under source/; no trust_remote_code or training checkpoint is required.
Files and identity
| Path | Contents |
|---|---|
inference/ |
Unique FP32 state_dict, config, original tokenizer, token manifest and provenance; no optimizer |
coreml/ios18-int8-block32/ |
Original Core ML .mlpackage and manifest |
source/ |
Frozen VimeML Python source, scripts, configurations, templates and requirements |
infer.py |
FP32 generation and candidate-scoring wrapper using the original V2 loader |
evaluation/summary.json |
Aggregate quality, alignment and device results; no benchmark texts or traces |
RELEASE.json, SHA256SUMS.txt, verify_release.py |
Release inventory and optional byte verification |
- Source: Voltline/VimeML, commit
4f8c54314e511208b695e8a938a764467546fdb3. - FP32 inference manifest:
9a4f97bba87cef2d44b67bb8b8d8bd8df1d95a160191a1873102944ce3adeac1. - Core ML manifest:
1dda48c76eacfa3ba99764098fab721eed4ada413130b99c0324620365aca0b7. - Tokenizer:
cdb0c60529300619fb0d7b01370c87b3f1028f2687ff67d97c778b15bfa2d177. - Original selected training checkpoint:
fa202f18d00cf5e77469bd3be5b8b264117fb73233ec9966fa3213257a470104; the optimizer-bearing checkpoint is not distributed.
These identities are SHA256 digests. Original model manifests remain unchanged. Their historical paths describe provenance rather than required installation locations. The release inventory uses the same format as V1; the inference bundle explicitly declares the V2 architecture.
Architecture and tokenization
Pre-Norm RMSNorm/SwiGLU, bias-free Linear layers, learned absolute positions and causal attention. Vocabulary 16,384; context 128; width 320; 6 layers; 5 heads of dimension 64; SwiGLU width 832; dropout 0; RMSNorm epsilon 1e-5. Input embedding and LM head share one parameter. There is no KV cache or sliding context; generation recomputes the prefix each step.
SentencePiece 0.2.1 Unigram, identity normalization, whitespace preservation, byte fallback and add_dummy_prefix=false. PAD/UNK/BOS/EOS IDs are 0/1/2/3. The tokenizer was learned from 2 million training sentences. V2 token IDs differ from V1, despite the same vocabulary size; model and tokenizer must be updated together.
Quick start: downloaded files
Python≥3.11. Install the included requirements after selecting a suitable PyTorch build. The Darwin branch retains the Core ML experiment environment; FP32 inference does not invoke Apple runtime services.
python -m pip install -r source/requirements.txt
python verify_release.py
python infer.py --prompt "今日は雨が降っているので、" --max-new-tokens 8
python infer.py --prompt "雨が降っているので、" --candidate "傘を持っていく" --candidate "橋を渡っていく"
The candidate command ranks supplied strings; it does not retrieve dictionary candidates or perform kana conversion. The diagnostic wrapper raises on invalid or overlong inputs; the application retains dictionary order when a whole candidate pool cannot be scored.
The V2 loader restores the shared head using torch.load(..., weights_only=True) and checks file sizes, model configuration and tokenizer contract. verify_release.py checks the complete inventory and sizes without running a model; --hashes additionally checks all file bytes against the recorded digests. Checksums are not a digital signature. Keep source/, the tokenizer and inference files together.
Candidate scoring
Encode context + candidate jointly with BOS. Find the common token prefix of the context and all candidate sequences, then sum full-vocabulary log-softmax over the remaining suffix. Do not append candidate-final EOS; ties keep original order. This suffix likelihood is a scoring proxy, not exact string conditional probability. Oversized or invalid pools fall back as a whole, and recall failures stay in evaluation denominators.
Deployment uses pure LM scores, with AzooKey retrieving kana candidates. V1's historical AzooKey score + 2 × LM sum experiment is not the V2.1 policy; no inherited fusion coefficient is claimed.
Training and evaluation
The corpus is the frozen V1 Japanese sentence corpus from nine local shards of FineWeb2-Edu Japanese and Tatoeba, cleaned, exactly deduplicated and grouped into 98%/1%/1% splits. It contains 25,713,003 sentences, including 25,185,368 training sentences. Semantic near-duplicates and public evaluation overlap have not been comprehensively audited.
V2 uses BF16 AdamW and a 30% random prefix crop on eligible first windows; validation is uncropped. V2.0 trained four epochs and selected step375000. V2.1 restart1 loaded those weights with a new AdamW, batch512 and a compiled backbone, then scanned two further epochs. Extend5 continued from restart1 step98426 with AdamW state: a maximum five-epoch budget, stopped after 3.273 epochs at step161095. The selected step40000 precedes later epochs whose full-validation BPC did not improve.
The selected FP32 model's full validation BPC is 3.0802686334, versus V1's 3.4736559. BPC uses total NLL, including EOS, divided by Unicode characters×ln2; token loss/PPL across different tokenizers is not used as a direct comparison. INT8 BPC and test-split loss were not measured. The test split was not used for checkpoint selection or tuning.
Fixed N-best 20 AzooKey pools are reranked without adding correct answers. AJIMEE JWTD_v2/v1 has 200 cases; the original synthetic development set has 137 model-reviewed cases; the expanded development set has 2000 real candidate pools with draft labels.
| Pool | AzooKey Top-1 | FP32 Top-1 / Top-5 | INT8 Top-1 / Top-5 | Pool coverage |
|---|---|---|---|---|
| AJIMEE | 87/200 | 144/200 / 160/200 | 144/200 / 160/200 | 162/200 |
| Original development | 111/137 | 122/137 / 136/137 | 122/137 / 136/137 | 136/137 |
| Expanded draft development | 1224/2000 | 1487/2000 / 1748/2000 | 1481/2000 / 1748/2000 | 1756/2000 |
V1 FP32 Top-1 on these pools was 124/200, 122/137 and 1453/2000. V2.1 FP32 improves expanded draft Top-1 by 34 cases; exploratory paired p=.00648, without multiple-comparison correction. These development results do not establish application-wide or formally adjudicated blind accuracy. The separate 1000-case blind pool was not scored or used for model/quantization selection.
INT8 vs FP32 changes 7/1/30 first choices and 163/121/404 complete rankings on AJIMEE/original/expanded development. AJIMEE has 3 correct→wrong and 3 wrong→correct changes; the original development change is between acceptable answers. Expanded drafts have 8 improvements and 14 regressions, a net loss of6 cases (-0.3 percentage points, paired p=.28628). No significant decrease was detected; statistical equivalence was not established. Labels and post-hoc aliases were not changed to select the model.
Core ML variant
Weight-only symmetric INT8, block size 32; FP32 computation; CPU_ONLY; minimum iOS18. No activation quantization or INT8 arithmetic is claimed. 32 learned matrices are compressed; RMSNorm parameters and structural constants/masks remain FP32. Logical .mlpackage size is 14,330,856 bytes, versus 50,358,014 bytes for the uncompressed FP32 package, about 71.5% smaller. Package, compiled resources, IPA and resident memory are different measures.
Input input_ids: INT32 [1,T], 1≤T≤128. Output logits: FLOAT32 [1,T,16384], raw full-vocabulary logits. Use right-side PAD and .cpuOnly. Mac conversion/device observations do not establish GPU or ANE compatibility for this release.
xcrun coremlcompiler compile coreml/ios18-int8-block32/model.mlpackage compiled-int8 --platform ios --deployment-target 18.0
Uncompressed FP32 conversion passed frozen logits alignment, with maximum absolute error 0.0000457764. INT8 fails strict FP32 logits alignment, with maximum error 1.00949144 at unchanged atol/rtol 3e-4. Right-PAD and future-token prefix invariance checks have maximum difference 0. The strict failure is retained as accepted lossy quantization drift, not reported as numerical equivalence.
Expanded draft first-choice changes concentrate on close competitors: the FP32 top-two score margin median is 0.070 in 30 changed cases, versus 3.072 in 1910 unchanged multi-candidate cases. 12 fixed prompts preserve next-token Top-1, but only 8/12 preserve the full 32-token greedy continuation. Reranking acceptance does not imply identical long generation.
Performance evidence and limits
iPhone16 Pro Max / iOS27.2, Release, CPU_ONLY. The test-host App benchmark records candidate scoring p50/p95 of 21.35/26.46ms, next-word 13.79/18.87ms and beam 99.20/102.11ms. Its lifecycle footprint peak 59.64MiB is separate from the keyboard extension. First process load 185.54ms may include warm OS/Core ML caches, so fully cold startup is not certified.
A separate 201.297-second real keyboard-extension recording has 164 approximately 1-second Instruments samples and 1920 in-process 100ms samples. Settings and model version were recorded; scoring and next-word suggestions were active. Extension kernel lifetime footprint peaked at 34.72MiB, while private resident sampling peaked at 73.78MiB. These measures cannot be substituted for each other.
Real-extension LM scoring p95 is 53.92ms; next-word p95 is 38.11ms. Result-ready→UI publication p95 is 8.79ms but maximum 1926.09ms, with no per-event attribution. Actual manual input did not show a noticeable abnormality and the current tail is accepted. Late-window footprint growth and a second model load remain unattributed; this bounded workload does not certify long-term memory stability, system termination behavior or energy use. Measurement was disabled after collection and input text was not logged.
Intended use and limitations
Japanese within-sentence candidate reranking and short continuation, primarily for local input-method experiments. This small non-instruction-tuned model can repeat, drift, truncate or invent facts. It is not a standalone kana-to-kanji converter, general assistant or factual reference. Fixed generation and candidate-pool metrics do not cover all real inputs.
AJIMEE has been used for error analysis; it is not a fresh blind test. Original development labels were generated and reviewed with model assistance, while expanded labels remain drafts without formal native-speaker adjudication. Public training overlap and near-duplicate risks remain. Benchmark context can cross sentences while the client uses current-sentence context; candidate coverage also limits achievable reranking accuracy.
License and attribution
Weights and included VimeML code follow V1's GNU GPL v2.0 release license. See LICENSE and source/LICENSE. Training sources and third-party dependencies retain their own terms.
FineWeb2-Edu Japanese's dataset card declares ODC-BY; Tatoeba text terms default to CC-BY 2.0 FR with author attribution. AJIMEE source data is CC-BY-SA3.0. This repository distributes aggregate evaluation metrics rather than corpus or benchmark texts; SentencePiece runtime is installed as a dependency rather than vendored here. Dataset copies, API responses, optimizer/RNG state, user inputs, raw device traces, credentials and the separate Vime client are excluded.
- Downloads last month
- 1