Safetensors
qwen2
Mitiantian commited on
Commit
1c33bd5
·
verified ·
1 Parent(s): ef93abe

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -4,7 +4,7 @@ license: apache-2.0
4
 
5
  # daVinci-origin-7B
6
 
7
- **daVinci-origin-7B** is a fully transparent, 7-billion parameter foundation model trained from scratch. It serves as a "clean-room" baseline for the research paper [*Data Darwinism -- Part1: Unlocking the Value of Scientific Data for Pre-training*](https://github.com/GAIR-NLP/Data-Darwinism).
8
 
9
  Unlike most open-source models, daVinci-origin-7B was explicitly trained on a dataset **strictly excluding scientific content** (books and research papers). This unique design allows researchers to unambiguously attribute performance gains to specific domain data injection strategies during continued pre-training.
10
 
 
4
 
5
  # daVinci-origin-7B
6
 
7
+ **daVinci-origin-7B** is a fully transparent, 7-billion parameter foundation model trained from scratch. It serves as a "clean-room" baseline for the research paper [*Data Darwinism -- Part1: Unlocking the Value of Scientific Data for Pre-training*](https://arxiv.org/pdf/2602.07824).
8
 
9
  Unlike most open-source models, daVinci-origin-7B was explicitly trained on a dataset **strictly excluding scientific content** (books and research papers). This unique design allows researchers to unambiguously attribute performance gains to specific domain data injection strategies during continued pre-training.
10