directional-llm-v1

A micro language model trained from scratch on a single task: answering the same arithmetic problem in two different formats โ€” exact and directional.

Model Description

This model is a proof of concept for the format-resolution hypothesis: a model's measured capability on a benchmark depends not only on its underlying reasoning but on the resolution at which the answer is requested. A model can reason correctly about a quantity while failing to produce its exact value, and can succeed at producing a direction while failing the exact-value version of the same question.

directional-llm-v1 is trained on paired data where the same underlying problem is presented in two formats:

  • Exact format: 12+34= โ†’ 46
  • Directional format: 12+34>40? โ†’ y (is 12+34 greater than 40?)
  • Directional format: 12+34<40? โ†’ n (is 12+34 less than 40?)

The model is not pretrained. There is no distillation. It is a small transformer trained from random initialization on synthetic arithmetic data. The purpose is to demonstrate at the smallest possible scale that format matters and that a model can be trained to answer at both resolutions.

Architecture

Component Value
Parameters ~20,000
Layers 2
Attention heads 2
d_model 32
d_ff 64
Vocab 19 characters
Max sequence length 12
Activation ReLU
Normalization LayerNorm (pre-norm)
Optimizer AdamW
Framework Pure numpy, no PyTorch

Training Data

The training data is synthetic and generated on the fly. Each example is a two-digit arithmetic problem with a randomly chosen format. The distribution includes:

  • Exact sums: a+b= for a, b โˆˆ [10, 50]
  • Greater-than queries: a+b>c? with the correct answer y or n
  • Less-than queries: a+b<c? with the correct answer y or n

There is no held-out distinction between training and evaluation content (the arithmetic is always the same), but the evaluation set is fixed to prevent overfitting to specific sequences.

Intended Uses

  • Research on format-sensitivity: study how the resolution of the requested answer affects measured capability.
  • Benchmark auditing: pair with band-edge-detector to identify whether a benchmark's low score is a reasoning failure or a format failure.
  • Mode-agnostic evaluation: provide a capability estimate that separates reasoning from answer-format handling.
  • Education: a minimal, reproducible example of how the same underlying computation can be expressed at different resolutions.

Evaluation

The model is evaluated on three metrics:

Metric Description
Exact-format accuracy Fraction of a+b= queries with the correct sum.
Directional-format accuracy Fraction of a+b>c? and a+b<c? queries with the correct y/n.
Per-example rank correlation Pearson correlation between per-example correctness in exact and directional formats.

The rank correlation is the key metric. If it is high (~0.7+), the model's underlying arithmetic is the bottleneck and both formats measure the same capability. If it is low (<0.3), the model's format handling is the bottleneck and the two formats measure different capabilities.

How to Use

This model is not intended for production use. It is a research artifact. To run it:

# Download the standalone script from the repository
# The script trains from scratch in a few minutes on CPU
python directional_llm_v1.py

There is no from_pretrained method. The model is defined and trained entirely within the script.

Limitations

  • Tiny scale. The model has ~20k parameters. It is a demonstration, not a serious language model.
  • Narrow domain. Only two-digit arithmetic addition and comparison. No transfer to other tasks is expected.
  • No pretraining. The model has no general language ability. It only knows what it was trained on.
  • English-only. Character-level tokenization over digits and symbols.
  • No safety guarantees. The model has no alignment, no filtering, no refusal behaviour. It is a research tool.
  • Single format per example. Each training example is presented in exactly one format. The model does not see the same problem in both formats simultaneously.

Citation

If you use this model in your research, please cite:

@software{directional_llm_v1,
  title = {directional-llm-v1: A From-Scratch Micro Language Model for Format-Resolution Evaluation},
  author = {zeechimp},
  year = {2026},
  url = {https://huggingface.co/zeechimp/directional-llm-v1}
}

License

Apache 2.0. The model weights, training script, and documentation are all released under the same license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support