Spaces:
Runtime error
title: ECHO-1-ultranano
emoji: 🐣
colorFrom: indigo
colorTo: gray
sdk: docker
app_port: 7860
pinned: false
license: mit
ECHO-1-ultranano
A from scratch educational language model of about 9 million parameters, trained on
Gleam, with a live compile check. Pick a prompt, the model writes Gleam, and the
real gleam build decides whether it is green or red.
What this is, and is not
This is a research toy, not a production code model. It learned the shape of Gleam and writes small self contained snippets that compile some of the time, roughly 18 percent from an open prompt. If you want a model that writes real programs, this is not it.
The point is the method. Every training example was generated and filtered by the real Gleam compiler, so the model was taught to write self contained code that actually builds, not code that merely looks plausible. The demo keeps that honesty: output is green only when it truly compiles.
How it was built
Two stages. First, pretraining on real Gleam gathered from the public package ecosystem, which teaches the syntax but also teaches the model to reference symbols from a wider project context that a single snippet does not have. Second, a finetune on self contained units produced by a type directed generator and filtered by the compiler, which teaches the model to define what it uses and stand on its own.
Architecture
Built from scratch with the modern transformer toolkit, each piece a toggle for ablation: RMSNorm, RoPE, grouped query attention, multi head latent attention, SwiGLU, a sparse mixture of experts with shared experts and load balancing, multi token prediction, QK normalization, and a KV cache. This checkpoint uses latent attention with multi token prediction.
Limits
It is tiny and it overfits a redundant corpus quickly. It writes trivial functions, often ignores its own parameters, and still makes type errors. None of that is hidden here. The value is the pipeline from a verifier to a dataset to a model that writes code which compiles.