Training, search indexing, and live fetching are separate
OpenAI publishes distinct user agents for distinct jobs: one that gathers content which may inform model training, one that builds the search index ChatGPT queries when it answers, and one that fetches a page live because a user's request required it. They can be allowed and disallowed independently.

