“As part of this initial release, the OCR-extracted text (original and post-processed) as well as the metadata (bibliographic, source, and generated) of the 983,004 volumes, or 242B tokens, identified as being in the public domain have been made available.”

Harvard released 983,004 public domain books, 242 billion tokens, as LLM training data. The scans came from the Google Books project that started in 2006. Two decades ago Google digitized libraries for search. Now the same scans are the clean training data everyone wants.