GeoCheckr
FeaturesFree ToolsPricingBlog
Sign inGet Started
Home/Blog/Technical SEO Check: Six Sites, Two llms.txt, One Bot Wall

Technical SEO Check: Six Sites, Two llms.txt, One Bot Wall

August 29, 2026·6 min read·GeoCheckr Team
A technical SEO check finds the failures that sit between your site and a machine that reads it. On August 29, 2026 we fetched six well-known websites the way a text-only crawler does — plain HTTP, no JavaScript, no browser — and ran the checks our technical SEO tool scores that a raw fetch can answer: security headers, crawlability, indexability, server-side content, and mobile setup. Two of the six publish a real llms.txt. Four answer the llms.txt path with a 404 page instead of a file. Etsy adds a 403 wall on top. None of these companies set out to fail the test; the failures just do not show up in a browser.

What a technical SEO check measures

Run the browser test yourself and five of these six sites will look fine. A technical SEO check tests the quiet layer underneath. Our technical SEO checker scores six dimensions, and each one answers one question:

  • Security headers — does the server send HSTS, CSP, X-Frame-Options, X-Content-Type-Options, Referrer-Policy? A bare response reads as a trust problem to crawlers.
  • Crawlability — does robots.txt exist, does it reference a sitemap, does the llms.txt path return a file?
  • Indexability — is the status 200, is there a canonical, a title, a meta description, and no noindex?
  • SSR performance — does the HTML itself carry the words, or does the page depend on JavaScript that never runs for a crawler?
  • Mobile — is there a viewport meta tag?
  • Core Web Vitals — LCP, CLS, INP, and the performance score, from real browser field data.
The check is boring on purpose. It reads the raw response rather than the rendered page, because that is exactly what search engines and AI crawlers read.

Six sites on August 29: two files, four 404s, one wall

Every site below got the identical text-only treatment on the same day. Core Web Vitals need field data from real browsers, so the table covers the five dimensions a raw fetch answers:

SiteResponserobots.txtllms.txt pathCanonicalJSON-LD on homepageSecurity headersHTML size
mozilla.org20066 B, 3 lines, no sitemap302 → 404missingnone5 of 548 KB
github.com2002.3 KB, no sitemap refpublished, 28.7 KBpresentnone5 of 5574 KB
etsy.com403 for text clients52.9 KB, 1,818 lines, no AI bots named404 — 51.5 KB pagen/a (blocked)n/a (blocked)none on block page776 B block page
stripe.com200643 B, sitemap refpublished, 65 KBpresentOrganization, WebSite, Person5 of 5662 KB
who.int20017.9 KB, sitemap ref404 — 70.8 KB pagepresentOrganization, WebPage, WebSite, SearchAction5 of 5132 KB
bbc.com2006.1 KB, 39 sitemap refs404 — 32.4 KB pagepresentNewsMediaOrganization, WebPage4 of 5 — no CSP687 KB
Two real files exist, and both belong to companies that sell to developers: GitHub's 28.7 KB and Stripe's 65 KB of structured, link-carrying markdown. The four other sites answered the same request with a 404 — a redirect chain that ends in one for Mozilla, a straight 404 for Etsy, WHO, and BBC. A crawler that asks "what should I know about you?" gets back an error page.

What broke, on each dimension

We checked five dimensions on six sites, so there are 30 individual results. The failures cluster in three places.

Crawlability fails loudly here, not softly. The llms.txt path is a text file or it is nothing, and there is no soft middle ground. Mozilla redirects /llms.txt to /en-US/llms.txt and lands on a 404. Etsy, WHO, and BBC return 404 headers with full HTML error pages sized like real content — WHO's 404 page is 70.8 KB, which is larger than Stripe's entire llms.txt file. Two organizations with plenty of engineering resources publish the file, and four organizations with plenty of engineering resources do not.

Indexability fails asymmetrically. Mozilla ships five of five security headers and a clean 200, then omits the canonical tag entirely and puts zero JSON-LD on its homepage. GitHub is the mirror image: a real llms.txt, clean headers, a canonical — and still zero JSON-LD on the homepage. Two engineering-heavy organizations, each silent about itself in exactly one place a machine reads.

The bot wall is binary. Etsy answered our plain request for its homepage with 403 before any content was exchanged. Its robots.txt is the largest of the six at 52.9 KB, but 1,818 lines of it are a legacy crawl map — 1,681 Disallow rules across three user agents, with not one AI crawler named. The block happens at the edge, not in policy. BBC is the opposite failure: it blocks nothing and publishes 39 sitemaps, then omits a Content-Security-Policy header on the one page that matters — the least-discriminating site in the test is the one missing the most security-sensitive header.

The size spread deserves its own line. BBC's homepage HTML weighs 687 KB and Mozilla's weighs 48 KB — a 14x gap between two sites that both render fine in a browser. A crawler on a slow link reads every byte of that.

Run a technical SEO check yourself, in this order

The check costs five minutes and one terminal. Do the steps in the order they clone, because each one narrows the next:

  1. Fetch the homepage with a text-only client and record the status code, the bytes, and whether the actual content words appear in the raw HTML. Plain request, no user agent spoofing.
  2. Check robots.txt — it exists, it allows the paths that matter, it references a sitemap.
  3. Request the llms.txt path — and read what comes back, not just the status code. A redirect chain that ends in 404 is still a 404.
  4. Look for the canonical tag, the title, the meta description, and the viewport in the raw head. Note which are missing.
  5. Send the same three requests to your competitors — their files, their headers, their sizes — because comparison is where priorities come from. A 48 KB homepage with a canonical and no llms.txt tells you exactly what your 687 KB homepage with a redirect chain is not doing.
  6. Read the security header block and note which of the five are absent.
Then run the free technical SEO check on your domain — it scores all six dimensions and returns fixes ordered by impact, so the gaps above become your checklist. Pair it with the AI crawler check to see whether GPTBot, ClaudeBot, and PerplexityBot can reach your robots.txt, your llms.txt, and your core pages the way our fetches just reached six other sites.

Here is the part nobody writes in the brochure. Every failure in the table is free to fix and invisible to the people who own the sites: Mozilla's canonical, GitHub's JSON-LD, BBC's CSP header are one-file changes no browser user will ever see — and every AI system answering questions about those companies will. The sites that fail a technical SEO check do not know they failed it. Yours does not have to be one of them.

Technical SEOGEOAI Search

Related Articles

GEO for Travel: Eight Booking Sites, Three llms.txt Files

August 25, 2026

GEO for Education: Six University Sites, Zero llms.txt Files

August 24, 2026

GEO for Real Estate & Proptech: AI-Powered Property Discovery

August 15, 2026

GeoCheckr

AI Search Visibility Platform. Optimize your website for ChatGPT, Claude, Perplexity, Gemini, and Google AI Overviews.

Product

  • GEO by Industry
  • Pricing
  • Blog
  • FAQ

Scoring Tools

  • Full GEO Audit
  • Citability Checker
  • LLM Visibility
  • Platform Optimization

Technical Tools

  • AI Crawler Checker
  • llms.txt Checker
  • Schema Checker
  • Technical SEO

Company

  • About
  • Topics
  • Privacy Policy
  • Terms of Service

© 2026 GeoCheckr. All rights reserved.

AI Search Visibility Platform