Angle

Robots.txt Tester

Paste a robots.txt, give it a path, and find out which line decides the outcome — and why one crawler gets in while another does not. Nothing is uploaded; it all runs here in your browser.

A path or a full URL. Case matters.

Overrides the dropdown.

Verdict

—

Deciding rule
—
Line
—
Matched group
—

Every crawler, same path

Blocked for one bot and allowed for another is the bug most people are actually hunting.

Crawler Result Rule

Why is my page blocked for one crawler but not another?

Because a crawler obeys exactly one group in your robots.txt, and it picks that group by matching its own name. If your file has a User-agent: Googlebot block and a User-agent: * block, Googlebot reads only the first one and ignores the wildcard block completely — including any Disallow lines you assumed applied to everyone. A crawler with no group of its own falls back to the wildcard. This is why adding a courtesy block for one bot silently changes what that bot is allowed to fetch: it stops reading your general rules the moment it finds its own name.

Does Allow override Disallow?

Not by being an Allow. Within the group that applies, the winner is whichever rule has the longest path pattern, counted in characters. So Disallow: /wp-admin/ loses to Allow: /wp-admin/admin-ajax.php because the allow pattern is longer and therefore more specific. Only when two matching patterns are the same length does Allow win as a tie-breaker. Order in the file does not matter at all — you can put the allow first or last and get the same answer.

Does robots.txt actually stop scrapers?

No. robots.txt is a request, not a fence. Well-behaved crawlers from Google, Bing, OpenAI and Anthropic read it and comply; a scraper that wants your content simply ignores it, and nothing in the protocol can prevent that. It also does not hide anything, since the file itself is public and effectively lists the paths you would rather people not visit. If a URL must stay private, it needs authentication or an IP block. Use robots.txt for crawl budget and for telling honest bots what is not worth their time.

Why is a blocked page still showing in search results?

Because Disallow blocks crawling, not indexing. If other sites link to a URL you have disallowed, a search engine can still list it — it just cannot read the page, so it shows a bare URL with no description. Worse, because it is forbidden from fetching the page, it will never see a noindex tag you placed there. To keep a page out of results, allow crawling and serve noindex, or require a login.