Hi, I’m building a personal website and I don’t want it to be used to train AI. In my robots.txt file I blocked:

  • ChatGPT-User
  • GPTBot
  • Google-Extended
  • FacebookBot

What bots should I also add? Are there any other ways to block AI bots?

IMPORTANT: I don’t want to block search engine crawlers, only bots that are used to train AI.

  • jlow (he/him)@beehaw.org
    link
    fedilink
    arrow-up
    3
    arrow-down
    2
    ·
    edit-2
    1 year ago

    I don’t really understand the reasoning behind doing any of this, they didn’t give a fuck about stealing clearly copyrighted content in the first place, why would they care about you (not OP specifically) begging them not to steal your stuff. (As long as theres no laws about this which afaik there aren’t).

    • wagoner@infosec.pub
      link
      fedilink
      arrow-up
      1
      ·
      1 year ago

      So that leaves two options then. Leave the front door wide open, don’t bother with any locks. Or shut down the web site. I’m for at least closing the door with the right robots.txt

      • jlow (he/him)@beehaw.org
        link
        fedilink
        arrow-up
        1
        ·
        1 year ago

        The analogy should be either having the door open or having the door open but putting a note on the door saying to please not steal anything. I’m not saying you shouldn’t do it, I just don’t think it’s gonna do anything, so I’m not going to bother.