Buckets:

69.1 GB
4,321 files
Updated 9 minutes ago
README.md

Common Crawl logo

Common Crawl Data

An open repository of web crawl data freely available to researchers, developers, and innovators worldwide.

Main Crawl Archives

The archives of the main crawl are released on a monthly basis. For details about the data format and how to access the data, see

The following monthly crawl archives have been released on 🤗 Hugging Face:

ID/Location Announcement Billion Pages Total Size Compressed (TiB)
CC-MAIN-2026-17 April 2026 2.19 106.3
CC-MAIN-2026-21 May 2026 2.16 107.1
CC-MAIN-2026-25 June 2026 2.10 102.7
CC-MAIN-2026-30 July 2026 2.14 105.2
CC-MAIN-2026-34 August 2026 2.14 105.1
CC-MAIN-2026-39 September 2026 2.17 103.4

Common Crawl is a California 501(c)(3) registered non-profit organization.

Terms of Use | Privacy

Total size
69.1 GB
Files
4,321
Last updated
Sep 25
Pre-warmed CDN
US EU US EU

Contributors