Understanding robots.txt and RFC 9309
The robots.txt file resides at the root of a domain to direct search engine crawlers which URL paths may be downloaded. In 2022, the IETF standardized the Robots Exclusion Protocol under RFC 9309, mandating that the most specific (longest) path match wins.
Common robots.txt Mistakes
Disallow: /:Critical Warning: Blocks the entire domain from being crawled, purging search rankings within days.Disallow::A blank Disallow line explicitly allows crawling of all pages.Sitemap: URL:Declaring the full XML sitemap URL helps bots discover deep URLs automatically.