“It’s futile to block AI crawler bots because they lie, change their user agent, use residential IP addresses as proxies, and more.”

robots.txt told every bot to stay out. Amazon’s crawler kept hammering the git server anyway. Blocking the user agent failed because about 10% of requests did not use it. Amazon gets the training data and the small host pays for the bandwidth.