You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Llama-3.2-1B — vanilla

标准 softmax 基线(attn_softmax=vanilla, attn_res_softmax_fn=vanilla).

  • Base model: meta-llama/Llama-3.2-1B
  • Training data: BookCorpus + Wiki40B (en)
  • Steps: 1000 (warmup 100), block_size 512, effective batch 192 (6 × 4 GPU × 8 grad-accum)
  • Optimizer: AdamW, lr 4e-4, linear, weight_decay 0.1
  • Precision: fp16 mixed (fp32 master weights)

Research checkpoint from an outlier-efficiency study; only 1000 steps, not production-ready.

Metrics (final eval)

metric value
perplexity 143.82
model.norm inf-norm 14.8
max per-layer inf-norm 311.9

Companion run: robinzixuan/llama-3.2-1b-oasis.

Downloads last month
66
Safetensors
Model size
1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for robinzixuan/llama-3.2-1b-vanilla

Finetuned
(938)
this model