irex

English in, JavaScript regex out. A 2.6M-parameter transformer trained from random weights, with no pretrained components.

$ python infer.py "a hex color code like #fff or #a1b2c3"
/^#(?:[0-9a-fA-F]{3}){1,2}$/

Usage

Requires Apple Silicon (MLX).

hf download ntedvs/irex --local-dir irex && cd irex
pip install -r requirements.txt

python infer.py "iso date like 2024-01-31"      # /^\d{4}-\d{2}-\d{2}$/
python infer.py -k 5 "an email address"         # top 5 candidates
python infer.py                                 # interactive

This is a custom MLX model. It does not load with transformers.

Model

  • Parameters: 2.6M (d192, 4 layers, 4 heads)
  • Architecture: decoder-only, RMSNorm, RoPE, SwiGLU, tied embeddings
  • Tokenizer: BPE (4096) trained from scratch for the English, one token per digit; raw bytes for the regex
  • Decoding: beam search (8), reranked by: compiles, matches example strings in the request, in scope
  • Training: 60k steps, batch 128, about 9 passes over the data, 40 minutes on an M4 Pro

Scope

JS regex at an intermediate level: literals, ., classes, \d \w \s \b and negations, quantifiers (including lazy), groups, |, ^ $, and the i flag. No lookarounds, backrefs, named groups, \p{}, or other flags.

Training data

About 815k English-to-regex pairs:

  • Regexes: 159k, sampled from a grammar plus realistic templates. Each is labeled with matching and non-matching strings using real JS semantics (QuickJS).
  • English: generated by DeepSeek-V4.1-Flash in five styles: casual, terse, precise, purpose, sloppy. Plus short everyday requests.
  • Rule-based: 60k everyday pairs: passwords, lengths, starts/ends/contains, and similar.

The split is by regex, so test regexes never appear in training.

Results

Eval Score
Held-out test, 15,058 pairs, behavioral match 67.9%
Same, exact string match 39.3%
Compiles 100%
Everyday bench A (40 requests) 82.5%
Everyday bench B (20 requests, never tuned on) 65%

Behavioral match means the output accepts and rejects the same strings as the reference regex. It's checked on sampled strings, so it isn't a proof of equivalence.

By style: precise 90%, casual 79%, terse 66%, sloppy 62%, short 57%, purpose 47%. Most misses come from underspecified requests. For example, "zip code" doesn't say whether to allow +4, and many requests don't say whether to match the whole string or anywhere in it.

Larger models (5.9M, 12M, 28M) trained on the same data scored the same. The data is the limit, not the model size.

Limitations

  • Weak on numeric ranges (years 1900 to 2099), MAC addresses, and full UUIDs.
  • Anchoring is often a guess when the request doesn't specify it.
  • Can copy request words literally. For example, "password" may become pass.
  • Always test the output before using it.

License

MIT. Training text was generated with the DeepSeek API, whose terms assign outputs to the user and permit training other models on them.

Downloads last month
278
Safetensors
Model size
2.61M params
Tensor type
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support