I’d link to some blog posts about this an example, but the site they’re from went down a while ago.

At risk of giving anyone ideas, why do LLM training scrapers request pages from sites millions of times per day instead of just doing the equivalent of wget -r https://example.com/ ? If the point is just stealing things people have written, what do they accomplish by wasting a web host’s resources beyond being an ⊛ to webmasters?

Edit: Let me elaborate. A lot of the answers i’m seeing are just restating the problem without explaining why LLM scrapers are apparently all either coded by idiots or assholes. I know that they’re hitting sites with unreasonable numbers of requests, wasting bandwidth, and making tools like Anubis too important. I know that’s what’s happening. I’m asking why. What they gain from not being even a little intelligent about this.

With all the money and effort (and maybe even brainpower) going into this, surely there’s some explanation beyond incompetence.

  • Dr. Moose@lemmy.world
    link
    fedilink
    English
    arrow-up
    4
    arrow-down
    9
    ·
    edit-2
    12 days ago

    Honestly a lot of web is just extremely poorly optimized which is not a problem when you have have 1,000 page requests per day but suddenly you get 10,000 and things start breaking.

    Optimizing pages is very easy these days and a good website can serve hundreds of thousands of requests per day same as first 1,000 because caching is extremely powerful.

    However, same goes for crawlers which can be really poorly written and extremely unoptimized. This is especially the case with broad crawlers like AI crawlers that crawl all pages as they use a real web browser causing more expense. They use real browsers because websites these days use client side rendering to defer costs to the client AND anti-bot system that require a browser to bypass.

    Tl;dr: lots of very bad code

    Edit: leave it to luddites to down vote almost 30 years of web dev experience lol

    • sanzky@lemmy.world
      link
      fedilink
      English
      arrow-up
      2
      arrow-down
      1
      ·
      12 days ago

      because you are clearly minimizing the impact of AI crawlers. sites don’t go from 1000 page requests /day to 10,000. I’ve seen sites going from 50~100 request per minute to 10.000 request per minute and they are mostly bots. no matter how efficient your site is, if the server cannot manage the connections it does not matter whether the content is cached or not.

      And caching does not solve everything because many cloud providers charge for it, so even if you could potentially deliver them, it suddenly becomes too expensive to do so.