Update README.md
Browse files
README.md
CHANGED
|
@@ -4,7 +4,7 @@ license: apache-2.0
|
|
| 4 |
|
| 5 |
# daVinci-origin-7B
|
| 6 |
|
| 7 |
-
**daVinci-origin-7B** is a fully transparent, 7-billion parameter foundation model trained from scratch. It serves as a "clean-room" baseline for the research paper [*Data Darwinism -- Part1: Unlocking the Value of Scientific Data for Pre-training*](https://
|
| 8 |
|
| 9 |
Unlike most open-source models, daVinci-origin-7B was explicitly trained on a dataset **strictly excluding scientific content** (books and research papers). This unique design allows researchers to unambiguously attribute performance gains to specific domain data injection strategies during continued pre-training.
|
| 10 |
|
|
|
|
| 4 |
|
| 5 |
# daVinci-origin-7B
|
| 6 |
|
| 7 |
+
**daVinci-origin-7B** is a fully transparent, 7-billion parameter foundation model trained from scratch. It serves as a "clean-room" baseline for the research paper [*Data Darwinism -- Part1: Unlocking the Value of Scientific Data for Pre-training*](https://arxiv.org/pdf/2602.07824).
|
| 8 |
|
| 9 |
Unlike most open-source models, daVinci-origin-7B was explicitly trained on a dataset **strictly excluding scientific content** (books and research papers). This unique design allows researchers to unambiguously attribute performance gains to specific domain data injection strategies during continued pre-training.
|
| 10 |
|