MicroMamba

MicroMamba is a small-compute test of input-dependent state-space dynamics. Each sequence contains distracting symbols, a few marked symbols, and a final query asking for one marked item by ordinal position. Solving the task requires selective storage and retrieval rather than ordinary next-token statistics.

The model uses a compact Mamba-inspired block with:

  • a causal depthwise convolution;
  • learned stable diagonal state dynamics;
  • input-dependent discretization, input, and readout terms;
  • gated residual output.

The benchmark retains two controls: a state-space model whose dynamics do not depend on the current input and a GRU with comparable scale. This is a pedagogical Mamba-inspired implementation, not a bit-exact reproduction of the official Mamba kernel or its large-scale language-model results.

Verified results

All variants trained on the same 12,000 length-48 sequences and were evaluated on 4,000 independently generated sequences at each length.

Variant Parameters Length 48 Length 96 zero-shot
Selective SSM 4,594 87.85% 87.23%
Fixed-dynamics SSM 3,314 43.23% 31.05%
GRU control 7,146 69.38% 70.00%

On this controlled task, input-dependent state dynamics improved in-distribution accuracy by 44.63 points over fixed dynamics and 18.48 points over the larger GRU. The result is specific to this synthetic selective-memory benchmark.

Reproduce

uv run python projects/micro-mamba/train.py
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support