Skip to content
NakodaAI
← Research & Resources

DatasetCommon Crawl Foundation2008

Common Crawl Web Archive

By Common Crawl Foundation

A nonprofit, openly licensed archive of web crawl data, published roughly monthly since 2008 and now exceeding ten petabytes, hosted freely via AWS's Open Data Sponsorship Program.

Why it matters

It is the training-data backbone almost no end user sees named directly: filtered Common Crawl snapshots underlie Google's C4 corpus (used to train the T5 family) and were reported as the majority source of GPT-3's training tokens - understanding it is close to understanding what "trained on internet-scale text" actually means.

Key takeaways

  • Founded in 2007 as a 501(c)(3) nonprofit; has published web crawl data roughly monthly since 2008, each crawl typically containing more than two billion pages.
  • Freely downloadable via data.commoncrawl.org and hosted through AWS's Open Data Sponsorship Program at no cost to users.
  • Cited in over 12,000 research papers and used, directly or via derived corpora, as training data for most large language models released to date.
  • Raw crawl data requires substantial filtering before use - most well-known corpora built from it (C4, RefinedWeb, and others) are defined largely by their filtering methodology, not the raw crawl itself.

Part of these reading paths

Related

Also worth reading

Academic resourceGaoling School of Artificial Intelligence, Renmin University of China (with Université de Montréal)2023-03-31

A Survey of Large Language Models

A continuously updated academic survey covering the background, key findings and mainstream techniques of large language models, organised around pre-training, adaptation tuning, utilisation and capability evaluation.

  • Research Foundations

Research paperGoogle Brain / Google Research2017-06-12

Attention Is All You Need

The paper that introduced the Transformer architecture - a sequence model built entirely on attention mechanisms, with no recurrence or convolution - and became the architectural basis for essentially every large language model that followed it.

  • Research Foundations