Llama-3.2-1B — OASIS (softmax1)
OASIS softmax1 变体(eager attention + history propagation)(attn_softmax=softmax1, attn_res_softmax_fn=softmax1).
- Base model: meta-llama/Llama-3.2-1B
- Training data: BookCorpus + Wiki40B (en)
- Steps: 1000 (warmup 100), block_size 512, effective batch 192 (6 × 4 GPU × 8 grad-accum)
- Optimizer: AdamW, lr 4e-4, linear, weight_decay 0.1
- Precision: fp16 mixed (fp32 master weights)
Research checkpoint from an outlier-efficiency study; only 1000 steps, not production-ready.
Metrics (final eval)
| metric | value |
|---|---|
| perplexity | 72.83 |
| model.norm inf-norm | 25.1 |
| max per-layer inf-norm | 250.3 |
Companion run: robinzixuan/llama-3.2-1b-vanilla.
- Downloads last month
- 73
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for robinzixuan/llama-3.2-1b-oasis
Base model
meta-llama/Llama-3.2-1B