Build a robots.txt with Allow, Disallow, Crawl-delay and Sitemap lines for up to 3 crawlers. Copy it or download it.
The default policy has three settings. Allow all writes User-agent: * followed by an empty Disallow line, which means nothing is blocked. Block all writes Disallow: /, which asks every crawler to stay away from the whole site. Custom groups lets you set up to three groups, each for one crawler: All crawlers (*), Googlebot, Bingbot, GPTBot, or Other with a name you type. A group has Disallow paths, Allow paths and an optional Crawl-delay. Sitemap addresses, up to three, are typed in the box one per line and are written at the end of the file. The page opens on a sample that blocks /admin/ and /cart/ for everyone and all of the site for GPTBot.
Paths are typed separated by commas or line breaks. Each must start with / or with a wildcard *, and may not contain spaces or #; a path that fails is left out and a notice says so. A group with no Disallow and no Allow path is written as an empty Disallow, which allows everything, and the notice explains that. A Sitemap line must be a full http:// or https:// address. Only the first three are used. A crawler name may use letters, digits, dot, underscore, hyphen and *. Two groups for the same crawler get a warning, because crawlers may merge them or take only one.
Crawl-delay is a number of seconds from 0 to 600 between requests. Bing and some other crawlers follow it. Google ignores it, and the page tells you whenever a delay is in the file; to slow Googlebot you use the crawl settings in Google's own tools instead. Download .txt saves the file under the tool's own name, robots-txt-generator.txt, so rename it to robots.txt before you upload it.
The file is advisory. Well-behaved crawlers read it and follow it, and others ignore it, so it is not access control: anyone can open /robots.txt and see the paths you listed, and a path you hide there can draw attention to it. Blocking a page also does not remove it from search results. If other sites link to a blocked URL, a search engine can still list the address without a description, and it cannot read a noindex tag on a page it is not allowed to fetch. For private content use a login. To keep a page out of search results, leave it crawlable and add a noindex tag. The file must be saved as robots.txt in the root of the site.
At the root of the host, so that it opens at https://example.com/robots.txt. A file in a folder is ignored. Each subdomain needs its own file, and http and https are separate addresses.
Not reliably. It stops the crawler fetching the page, but if other pages link to it the address can still be indexed. To keep a page out of results, let it be crawled and add a noindex tag, or put it behind a login.
Set group 1 to All crawlers with an empty Disallow, and group 2 to GPTBot with Disallow /. The file then has User-agent: * with an empty Disallow, and User-agent: GPTBot with Disallow: /. This only works for crawlers that follow robots.txt.
Disallow: with nothing after it means nothing is disallowed, so the crawler may fetch everything. Disallow: / blocks the whole site. They look alike but do opposite things.
The file is built in this tab and nothing is uploaded. Your choices for the policy, paths and delay are remembered on this device so the page opens as you left it; the sitemap box is not stored. Reset clears them.
Paths must start with / or *, and may not contain spaces or #. A path like admin/ is not valid, so it is skipped and the notice says so. Write /admin/ instead.