Hummingbird-V1 / TRAINING_DATA.md
juinron's picture
Release Hummingbird-V1 500M natural continuation checkpoint
d03200e verified
|
Raw
History Blame Contribute Delete
2.69 kB

Hummingbird-V1 training-data notices

This file documents the data lineage for the released Hummingbird-V1 weights. Apache-2.0 applies to the project code and released model materials; it does not replace the licenses or notices attached to third-party source datasets.

Training lineage

The selected checkpoint has 2,000,170,752 cumulative token presentations:

  • 1,500,000,000 presentations in the original Muon causal-pretraining phase;
  • 500,170,752 presentations in the natural-corpus continuation phase.

The continuation used a 1,091,660,174-token packed training split. Its complete prepared-corpus totals across train, validation and held-out splits were:

Source Frozen revision Prepared tokens License
FineWeb-Edu (sample-100BT) 87f09149ef4734204d70ed1d046ddc9ca3f2b8f9 570,339,212 ODC-By 1.0
DCLM baseline 1.0 a3b142c183aebe5af344955ae20836eb34dcf69b 211,970,663 CC-BY 4.0
FineWeb-HQ e58199cdd52438d94405df1a4d8630cc5f13bf84 109,226,534 ODC-By 1.0
SmolLM-Corpus / Cosmopedia v2 3ba9d605774198c5868892d7a8deda78031a781f 168,663,969 ODC-By 1.0
FineMath (finemath-4plus) e92b25a616738fe95dc186b64dfb19f9c8525594 53,640,449 ODC-By 1.0

The original base phase also used FineWeb-Edu, Cosmopedia v2, TinyStories, and project-generated procedural arithmetic, tutorial and short-choice-rationale text. TinyStories revision f54c09fd23315a6f9c86f9dc80f725de7d8f9c64 is made available under CDLA-Sharing-1.0.

Filtering, deduplication and evaluation protection

The natural corpus used quality admission filters, canonical exact-document deduplication, and 32-permutation MinHash/8-band LSH near-duplicate removal confirmed by exact Jaccard similarity at a 0.80 threshold. Documents were split deterministically by SHA-256 before packing.

Rendered Open SLM and ArithMark-3 prompts plus choices were held in a protection index and excluded from training. Public benchmark results were nevertheless evaluated at the 250M, 500M, 750M and 1B continuation checkpoints and used to select the released 500M checkpoint. This checkpoint-selection bias is disclosed in the model card; benchmark records and answers were not training data.

Exact quotas, filtering, source fields, split policy and prepared token counts are included in training/corpus_contract.yaml, training/packed_metadata.json and training/provenance.json in the release package.