“OpenAI did not even buy the books it used. Instead, it began by torrenting books from the notorious and illegal pirate library Library Genesis.”
The datasets were called Libgen1 and Libgen2 internally. In the published papers they became Books1 and Books2. Then the files got deleted in 2022 once the lawyers noticed. The authors want summary judgment on 194 titles, and the rename is the part that kills the fair use story. You don’t relabel something you think is legal.