Sighted in April 2023 in a Washington Post analysis that listed the top websites in a Common Crawl-derived training set. Number one was a patents database; the newspaper itself was in the top hundred; so were several piracy sites and a great many personal blogs.
Common Crawl
A non-profit that has been crawling the web and giving the copy away since 2007, whose archive of many billions of pages became the raw material for most language models. GPT-3's training data was mostly a filtered Common Crawl; Google's C4 dataset was a cleaned one. A 2023 newspaper analysis of its contents found patents, Wikipedia, and a great many sites that had never agreed to anything.
Testimony
4 entries · newest firstI remember downloading a slice of it in 2015 to build a language model for a project and finding, in the first hundred pages, a forum about tractors, a Bible in Tagalog, and a page of nothing but the word 'cheap' repeated. I trained on all three. So, in the end, did everyone.
The record's summary: Common Crawl is the internet's rough draft, saved monthly. It was built so that researchers who were not Google could study the web. Then the web became the corpus, and a small charity in San Francisco found itself the quarry for an industry.
Common Crawl is not the whole internet and not a clean copy. It respects robots.txt, misses most of what is behind logins, and is full of boilerplate, spam and duplicate pages. The filtering the labs do to it is where much of their competitive advantage lives, and it is the part they do not publish.
Add to the record
What does it mean? Write it the way you would say it out loud.