I’d link to some blog posts about this an example, but the site they’re from went down a while ago.

At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host’s resources beyond being an ⊛ to webmasters?

Edit: Let me elaborate. A lot of the answers i’m seeing are just restating the problem without explaining why LLM scrapers are apparently all either coded by idiots or assholes. I know that they’re hitting sites with unreasonable numbers of requests, wasting bandwidth, and making tools like Anubis too important. I know that’s what’s happening. I’m asking why. What they gain from not being even a little intelligent about this.

With all the money and effort (and maybe even brainpower) going into this, surely there’s some explanation beyond incompetence.

  • sanzky@lemmy.world
    link
    fedilink
    English
    arrow-up
    2
    arrow-down
    1
    ·
    10 days ago

    because you are clearly minimizing the impact of AI crawlers. sites don’t go from 1000 page requests /day to 10,000. I’ve seen sites going from 50~100 request per minute to 10.000 request per minute and they are mostly bots. no matter how efficient your site is, if the server cannot manage the connections it does not matter whether the content is cached or not.

    And caching does not solve everything because many cloud providers charge for it, so even if you could potentially deliver them, it suddenly becomes too expensive to do so.