MICA Ember v0.3a

MICA family: View the collection. This release uses the bundled MICA integer engine; standard Transformers loaders do not load its .mica checkpoint.

A one-of-a-kind language model. Not a transformer, not a neural network.

MICA models language with learned integer rewrite rules on a ring of cells, a cellular automaton. It is an original architecture, designed and trained from scratch. Ember is the byte-level MICA model: it reads text one UTF-8 byte at a time. Version 0.3a adds word assistance: it suggests the next word, or finishes a word after its first two letters are typed.

Its sibling MICA Flame-W 0.3.1 predicts whole words and continues sentences. Ember is for word suggestions and completion. It is not a sentence generator.

Why MICA is different

  • No transformer inside. There is no attention and no stack of neural layers. Each byte is written onto a ring of integer cells, learned local rules rewrite the ring, and probes read the result.
  • Integers only. The rules, the probes and the word scores are all integer arithmetic.
  • Deterministic. The same text always gives the same suggestions.
  • Tiny and CPU-only. 5.7 MB in total, about 6.4 million stored integers (mostly int8). It runs on any CPU with numpy alone.
File What it is
model.mica The automaton, unchanged from Ember v0.2A: int8 rules and probes. 4,467,396 bytes, SHA-256 3b2a94e2…8a126e72
word_heads.npz Two integer word readouts over the automaton's 240 probes, for a 4,339-word vocabulary: one for the next word, one for completing a word from two typed letters. 1,233,028 bytes, SHA-256 cc8e0135…dadaa5d5
word_model.py EmberWord, the exact integer runtime. It checks both SHA-256 hashes when it loads
test_word_model.py Conformance test: hashes, one exact probe state and known suggestions
manifest.json, package.json Hashes, provenance and the engine geometry
mica_r1/ The exact integer engine. numpy only
results/ The paired comparisons with v0.2A, as JSON

Use it

pip install numpy huggingface_hub
hf download vynly/mica-ember-0.3a --local-dir mica-ember-0.3a
cd mica-ember-0.3a
python test_word_model.py
python word_model.py "Thank you for "
python word_model.py --complete "Are you coming to the pa"
Ember v0.3a model hashes, scalar probe and word suggestions: OK
['being', 'your', 'helping', 'taking', 'understanding']
['park', 'party', 'paper', 'past', 'parking']

From Python:

from word_model import EmberWord

ember = EmberWord()
ember.suggest("Thank you for ")                   # next word, top 5
ember.suggest("Thank you for be", prefix2=True)   # complete the word from its first two letters

End the text with a space to get the next word. For completion, end it with the first two letters of the word; only vocabulary words that start with them are ranked.

Results

Ember v0.3a was compared with Ember v0.2A, the same automaton without the word readouts. v0.2A suggests a word by spelling it byte by byte (beam search of width 4 over a 5,000-word vocabulary from the training text). Both were scored on the same prompts. Top-1 means the first suggestion is exactly the right word.

Each set has 200 chat prompts and 200 everyday prompts, and the two count equally. Gains are in percentage points, with paired 95% bootstrap intervals. The development prompts were used to choose the readouts. The held-out prompts were checked only after that choice was fixed.

Prompts Task v0.2A v0.3a Gain [95% CI]
Held out next word 6.75% 14.25% +7.50 [+4.75, +10.50]
Held out completion from 2 letters 42.75% 48.00% +5.25 [+1.50, +9.00]
Development next word 5.25% 13.50% +8.25 [+5.50, +11.25]
Development completion from 2 letters 46.00% 52.75% +6.75 [+3.00, +10.50]

Both tasks together: +6.38 points [+4.00, +8.75] on the held-out prompts and +7.50 [+5.13, +9.88] on development.

Held out, by domain v0.2A v0.3a
chat, next word 9.0% 20.0%
everyday, next word 4.5% 8.5%
chat, completion 49.0% 52.5%
everyday, completion 36.5% 43.5%

The byte automaton itself is unchanged from v0.2A: 1.861 bits per byte on held-out chat text and 1.876 on held-out everyday text (counting an end-of-text symbol). The project's sealed test set was not used.

How it works

  • Automaton. Each UTF-8 byte is written onto a ring of 768 cells × 112 integer channels. Then 16 ticks of learned integer rules rewrite the newest cell, reading cells up to 8 bytes back. In each tick, the cell's own values pick a page of 256 candidate rules, the candidates score the neighbouring cells with integer weights, and the winner writes 6 new values.
  • Probes. 240 probes read single cell values: 48 from the bytes themselves (up to 64 bytes back) and 192 from the rule results.
  • Word readouts (new in 0.3a). Two integer linear readouts score all 4,339 words from the same 240 probe values: score = W · f + B, with int8 weights, int32 biases and int64 sums. The next-word readout is used after a space. The completion readout is used after two typed letters and ranks only the words that start with them. Both were trained on word labels from the training text only, then quantized to integers.

Training data

English text collected by the project, balanced by bytes: chat dialogue 45%, everyday sentences 40% and short stories (TinyStories) 15%. The vocabulary is 4,339 common words from the training text. Sources and builders are in the GitHub repository.

Limits

  • This is a research model, not an assistant. On held-out prompts, its first suggestion is the right next word about 1 time in 7, and the right completion from two letters about half the time.
  • Only the 4,339 vocabulary words can be suggested.
  • The tests count exact top-1 matches only. They do not measure whether the other suggestions make sense, when to make no suggestion, or whole sentences.
  • Completion needs the first two letters of the word, as ASCII letters. Otherwise it returns an empty list.
  • The automaton can also write text byte by byte, but its sentences reuse stock phrases and contain grammar errors. Flame-W is the MICA model for continuing text.

Links

License

MICA AI Non-Commercial License: personal and research use only. Commercial use is forbidden. Code, docs and history: https://github.com/Vovala14/Mica-Ai.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including vynly/mica-ember-0.3a