Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| ada | 1 items | ||
| agda | 1 items | ||
| alloy | 1 items | ||
| antlr | 1 items | ||
| applescript | 1 items | ||
| assembly | 2 items | ||
| augeas | 1 items | ||
| awk | 1 items | ||
| batchfile | 1 items | ||
| bluespec | 1 items | ||
| c | 53 items | ||
| c-sharp | 45 items | ||
| clojure | 1 items | ||
| cmake | 1 items | ||
| coffeescript | 1 items | ||
| common-lisp | 2 items | ||
| cpp | 48 items | ||
| css | 12 items | ||
| cuda | 1 items | ||
| dart | 4 items | ||
| dockerfile | 1 items | ||
| elixir | 1 items | ||
| elm | 1 items | ||
| emacs-lisp | 1 items | ||
| erlang | 1 items | ||
| f-sharp | 1 items | ||
| fortran | 2 items | ||
| git-commits-cleaned | 55 items | ||
| github-issues-filtered-structured | 59 items | ||
| glsl | 1 items | ||
| go | 24 items | ||
| groovy | 1 items | ||
| haskell | 3 items | ||
| html | 29 items | ||
| idris | 1 items | ||
| isabelle | 1 items | ||
| java | 87 items | ||
| java-server-pages | 1 items | ||
| javascript | 65 items | ||
| json | 6 items | ||
| julia | 2 items | ||
| jupyter-scripts-dedup-filtered | 8 items | ||
| jupyter-structured-clean-dedup | 6 items | ||
| kotlin | 6 items | ||
| lean | 1 items | ||
| literate-agda | 1 items | ||
| literate-coffeescript | 1 items | ||
| literate-haskell | 1 items | ||
| lua | 3 items | ||
| makefile | 2 items | ||
| maple | 1 items | ||
| markdown | 79 items | ||
| mathematica | 2 items | ||
| matlab | 1 items | ||
| ocaml | 1 items | ||
| pascal | 2 items | ||
| perl | 3 items | ||
| php | 61 items | ||
| powershell | 2 items | ||
| prolog | 1 items | ||
| protocol-buffer | 1 items | ||
| python | 59 items | ||
| r | 1 items | ||
| racket | 1 items | ||
| restructuredtext | 4 items | ||
| rmarkdown | 1 items | ||
| ruby | 7 items | ||
| rust | 9 items | ||
| sas | 1 items | ||
| scala | 5 items | ||
| scheme | 1 items | ||
| shell | 4 items | ||
| smalltalk | 1 items | ||
| solidity | 1 items | ||
| sparql | 1 items | ||
| sql | 11 items | ||
| stan | 1 items | ||
| standard-ml | 1 items | ||
| stata | 1 items | ||
| systemverilog | 1 items | ||
| tcl | 1 items | ||
| tcsh | 1 items | ||
| tex | 6 items | ||
| thrift | 1 items | ||
| typescript | 27 items | ||
| verilog | 1 items | ||
| vhdl | 1 items | ||
| visual-basic | 2 items | ||
| xslt | 1 items | ||
| yacc | 1 items | ||
| yaml | 4 items | ||
| zig | 1 items | ||
| .gitattributes | 2.27 kB xet | 25a894cd | |
| README.md | 3.39 kB xet | 148635a8 |
StarCoder Training Dataset
Dataset description
This is the dataset used for training StarCoder and StarCoderBase. It contains 783GB of code in 86 programming languages, and includes 54GB GitHub Issues + 13GB Jupyter notebooks in scripts and text-code pairs, and 32GB of GitHub commits, which is approximately 250 Billion tokens.
Dataset creation
The creation and filtering of The Stack is explained in the original dataset, we additionally decontaminate and clean all 86 programming languages in the dataset, in addition to GitHub issues, Jupyter Notebooks and GitHub commits. We also apply near-deduplication and remove PII, all details are mentionned in our Paper: 💫 StarCoder, May The Source Be With You
How to use the dataset
from datasets import load_dataset
# to load python for example
ds = load_dataset("bigcode/starcoderdata", data_dir="python", split="train")
GitHub issues, GitHub commits and Jupyter notebooks subsets have different columns from the rest so loading the entire dataset at once may fail, we suggest loading programming languages separatly from these categories.
jupyter-scripts-dedup-filtered
jupyter-structured-clean-dedup
github-issues-filtered-structured
git-commits-cleaned
- Total size
- 311 GB
- Files
- 865
- Last updated
- Jul 1
- Pre-warmed CDN
- US EU US EU