Comment by gundmc

Comment by gundmc 13 hours ago

What do you mean by "barely" respecting robots.txt? Wouldn't that be more binary? Are they respecting some directives and ignoring others?

unsnap_biceps 13 hours ago

I believe that a number of AI bots only respect robot.txt entries that explicitly define their static user agent name. They ignore wildcards in user agents.

That counts as barely imho.

I found this out after OpenAI was decimating my site and ignoring the wildcard deny all. I had to add entires specifically for their three bots to get them to stop.

Reply View 10 replies

joecool1029 12 hours ago

Even some non-profit ignore it now, Internet Archive stopped respecting it years ago: https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...

Reply View | 6 replies
- SR2Z 11 hours ago
  
  IA actually has technical and moral reasons to ignore robots.txt. Namely, they want to circumvent this stuff because their goal is to archive EVERYTHING.
  
  Reply View | 4 replies
  
  prinny_ 9 hours ago
  
  Isn’t this a weak argument? OpenAI could also say their goal is to learn everything, feed it to AI, advance humanity etc etc.
  
  Reply View | 2 replies
  
  amarcheschi 10 hours ago
  
  I also don't think they hit servers repeatedly so much
  
  Reply View | 0 replies
- AnonC 7 hours ago
  
  As I recall, this is outdated information. Internet Archive does respect robots.txt and will remove a site from its archive based on robots.txt. I have done this a few years after your linked blog post to get an inconsequential site removed from archive.org.
  
  Reply View | 0 replies
noman-land 12 hours ago

This is highly annoying and rude. Is there a complete list of all known bots and crawlers?

Reply View | 1 reply
- jsheard 12 hours ago
  
  https://darkvisitors.com/agents
  https://github.com/ai-robots-txt/ai.robots.txt
  
  Reply View | 0 replies
[removed] 12 hours ago

[deleted]

Reply View | 0 replies

LukeShu 12 hours ago

Amazonbot doesn't respect the `Crawl-Delay` directive. To be fair, Crawl-Delay is non-standard, but it is claimed to be respected by the other 3 most aggressive crawlers I see.

And how often does it check robots.txt? ClaudeBot will make hundreds of thousands of requests before it re-checks robots.txt to see that you asked it to please stop DDoSing you.

Reply View 0 replies

Animats 11 hours ago

Here's Google, complaining of problems with pages they want to index but I blocked with robots.txt.

    New reason preventing your pages from being indexed

    Search Console has identified that some pages on your site are not being indexed 
    due to the following new reason:

        Indexed, though blocked by robots.txt

    If this reason is not intentional, we recommend that you fix it in order to get
    affected pages indexed and appearing on Google.
    Open indexing report
    Message type: [WNC-20237597]

Reply View 0 replies