Dense-306 / README.md
Pacific-i64's picture
Upload folder using huggingface_hub
46f10fb verified
|
Raw
History Blame Contribute Delete
3.23 kB
---
license: apache-2.0
language:
- en
pipeline_tag: text-generation
library_name: vllm
tags:
- causal-lm
- pytorch
- gqa
- research
datasets:
- HuggingFaceFW/fineweb-edu
---
# Dense-306
Dense-306 is the 306.5M-parameter dense baseline from the matched run pair in
*Token Identity as a Routing Signal for Residual MLP Experts*. It is a
decoder-only base language model trained from scratch on approximately 8B
FineWeb-Edu tokens.
This is a pretrained base model, not an instruction-tuned or safety-aligned
assistant.
## Architecture
- 18 transformer layers
- hidden size 1,024
- 16 query heads and 4 key/value heads (GQA)
- QK-Norm and RoPE
- 2,048-token training context
- 32,000-token tokenizer
- dense SwiGLU width 4,096
- 306,486,528 trainable parameters
## Evaluation
Both models were evaluated zero-shot with EleutherAI LM Evaluation Harness
0.4.12, identical seeds, BF16 inference, no sample limit, and a maximum context
of 2,048 tokens.
| Model | ARC-Easy acc_norm | PIQA acc_norm | HellaSwag acc_norm | WikiText-2 word PPL |
| --- | ---: | ---: | ---: | ---: |
| Dense-306 | 49.07 | 63.49 | 33.67 | 35.79 |
| TR-MOE-306 | 48.61 | 63.49 | 34.06 | 35.20 |
The multiple-choice differences are smaller than one reported standard error.
These results support near-parity, not task-level superiority.
At the last common point on a fixed FineWeb-Edu training-split evaluation
stream, the dense model reached NLL 2.9482 and the token-routed model reached
2.9329. That stream is not an independent held-out set, and each architecture
was trained with one seed.
## Files and integrity
`model.safetensors` is the exact FP32 model export used for the reported
downstream evaluation.
- final training step: `7629`
- training checkpoint SHA-256:
`0d5bc1d8d8ead8ecde70e2705dfc8712e99a3de79bdf24f9b4e34fd1f668de7b`
- exported `model.safetensors` SHA-256:
`311c97b2233a0d182fe47fa641a427e30f59b49774872a85b59854457ed1e7ae`
The repository also contains the tokenizer, generation configuration, and
`export_manifest.json`.
## Runtime
The checkpoint uses the custom `DeepForCausalLM` format implemented by the
[Complexity-ML vLLM runtime](https://github.com/Complexity-ML/vllm-cuda_graph).
Stock Transformers and upstream vLLM do not currently register the `deep`
architecture. Use that runtime for inference, or adapt its loader before
loading these weights.
The tokenizer can be loaded independently:
```python
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Pacific-i64/Dense-306")
```
## Limitations
This is a small research base model. It may generate inaccurate, biased,
offensive, repetitive, or unsafe text. It should not be used for high-stakes
decisions. The training corpus is web-derived and may contain errors,
copyrighted material, personal information, and social biases.
## Paper and reproducibility
- [OpenReview submission](https://openreview.net/forum?id=Jd9jhTnkUy)
- [Anonymous supplementary code and results](https://github.com/Complexity-ML/tmlr-paper-pool)
The paper reports a matched single-seed observation. Multi-seed replication
remains open.
## License
The Dense-306 model weights are released under the Apache License 2.0. See
`LICENSE`.