Technical
robots.txt
robots.txt is a plain-text file at the root of a host that asks crawlers which paths not to fetch. Well-behaved bots honour it, but it is a request and not an access control.
How it is measured
Fetch `https://example.com/robots.txt` and expect 200 with `text/plain`. Rules are `User-agent:` groups with `Disallow:` and `Allow:` lines, plus an optional `Sitemap:` line. Test specific URLs with a robots tester and confirm in the log that Googlebot requests the file before crawling.
A 404 means no restriction. A 5xx makes some crawlers pause the whole site until it clears. Google reads only the first 500 KB.
Worked example
A staging WordPress site is copied to production with `Disallow: /` still in robots.txt. Within a week Googlebot requests fall from 2,400 a day to 40, and Search Console reports Blocked by robots.txt on 900 URLs.
Deleting the line returns Googlebot to about 2,000 a day within five days. Separately, a Disallow on `/private-reports/` never hid those URLs. A link from another site still got them listed, address only.
How it differs
robots.txt controls crawling, meaning fetching. A sitemap suggests what to crawl. One says where not to go and the other lists where to go. A blocked URL can still be indexed from links, and a noindex tag only works if the crawler may fetch the page to see it.
Common errors
Disallowing `/` on production. Blocking the CSS and JS a renderer needs. Using it to hide private content. Combining Disallow with noindex, so the crawler never sees the noindex. Getting path case wrong. Serving different files on www and the bare domain.
In practice
Fetch your live robots.txt now and read it line by line. Check five key URLs in a tester and add the `Sitemap:` line. Put a check in your deploy that fails if production contains `Disallow: /`.