GAIR
/

daVinci-origin-7B

Model card Files Files and versions

daVinci-origin-7B / README.md

Mitiantian's picture

Update README.md

1c33bd5 verified 20 days ago

|

history blame contribute delete

653 Bytes

	---
	license: apache-2.0
	---

	# daVinci-origin-7B

	daVinci-origin-7B is a fully transparent, 7-billion parameter foundation model trained from scratch. It serves as a "clean-room" baseline for the research paper [Data Darwinism -- Part1: Unlocking the Value of Scientific Data for Pre-training](https://arxiv.org/pdf/2602.07824).

	Unlike most open-source models, daVinci-origin-7B was explicitly trained on a dataset strictly excluding scientific content (books and research papers). This unique design allows researchers to unambiguously attribute performance gains to specific domain data injection strategies during continued pre-training.