Common Crawl, AI Visibility, and Your Place Brand: What Actually Matters in 2026

If you run marketing for a destination, a tourism board, or any place brand (or any other brand), you've probably never heard of Common Crawl.

common-crawl-logo
the Common Crawl logo

That's about to change — because whether AI systems know anything about your destination (or brand) was substantially decided by this quiet, nonprofit web archive years before ChatGPT, Claude, or Perplexity ever existed.

Here's the plain-language version of what it is, why it matters, and — more usefully — what to actually do about it on Monday morning.

What is Common Crawl?

Common Crawl is a nonprofit that scans the internet and makes a free, public archive of web pages available to anyone. It's been running since 2008, crawls a few billion pages roughly every month, and has archived over 10 petabytes of raw web data over its lifetime.

It's not a search engine. It doesn't rank anything. It's infrastructure — the raw material that other people build on top of.

For most of its life it was funded almost entirely by its founder, Gil Elbaz, through his family foundation, with hosting donated by Amazon Web Services. In recent years, as AI companies have come to rely on it heavily, some of them have started contributing funding too — though it remains, structurally, a small nonprofit running on a surprisingly modest budget for something this foundational.

Why place brands should care

Here's the part that matters: nearly every major AI model has been trained, at least in part, on Common Crawl data.

GPT-3 drew roughly 60% of its training tokens from a filtered version of Common Crawl. Meta's Llama models used it as a primary text source. Google built its own cleaned derivative — called C4 — directly from a Common Crawl snapshot to train its T5 model. Falcon, Mistral, and most other open models followed the same pattern. Even where companies like OpenAI, Anthropic, and Google no longer publish exact data recipes for their newest models, the pattern holds: web-scale crawl data, heavily filtered, remains the backbone of the "world knowledge" baked into these systems.

In other words: if your destination's content was crawlable in the years these models were trained, it had a chance of shaping what AI systems "know" about your place. If it wasn't crawlable — blocked, gated, buried in PDFs, or simply never published in a form a crawler could parse — it likely didn't.

That's a one-way door. Blocking a crawler today doesn't undo what a model already learned from an earlier crawl. And allowing a crawler today doesn't retroactively get you into models already trained. This is slow-moving, structural stuff — which is exactly why it's easy to ignore and expensive to have ignored for the last five years.

Two separate battles: training-time inclusion vs. answer-time citation

This is the part that trips people up. There are actually two different games being played, and they have different rules.

Battle one: getting into the training data. This is Common Crawl's domain. It's a permission question — does your robots.txt allow CCBot, GPTBot, ClaudeBot, and the rest to crawl you? Notably, Common Crawl's bot (CCBot) is now the most widely blocked bot among the top 1,000 websites globally. A lot of publishers have decided the answer is no. Whether that's the right call for a place brand — which generally wants to be known and discovered — is a different question than whether it's the right call for a newspaper protecting subscription revenue.

Battle two: getting cited in AI-generated answers. This is the active battleground, and it's largely decoupled from training-time inclusion. Recent data on AI citation patterns is striking:

  • About 85% of brand mentions in AI answers come from third-party pages, not the brand's own website. Owning great content on your own domain isn't enough — you need to exist in the conversation about you, elsewhere.
  • Roughly 60% of AI Overview citations come from URLs that don't even rank in the top 20 organic search results. Classic SEO rank and AI citation are no longer the same game.
  • Comparative, editorial content — "best time to visit X," "X vs Y" — accounts for about a third of AI citations, while straightforward commercial or promotional pages account for under 5%.
  • Pages that aren't refreshed roughly every quarter are three times more likely to lose their citations over time. AI systems seem to weight freshness quite heavily.

So the old SEO playbook (rank #1, win the click) and the new AI-visibility playbook (get cited in an answer, regardless of rank) are related but not the same thing. You now need to win both.

Knowledge Assets vs. Asset Library: this distinction just got a lot more practical

If you've followed Brandkit's thinking, you'll know we've been drawing a hard line between the Asset Library — logos, templates, brand guideline PDFs, the stuff a brand manager needs for production — and the Knowledge Library: structured, AI-ready content designed to be read and cited by machines, not just downloaded by humans.

The AI citation data above explains exactly why that distinction matters in practice, not just in theory. A beautifully designed PDF brand guideline sitting in a DAM is essentially invisible to an AI system — it's not structured, it's not crawlable in a useful way, and it's not built to answer a question. A well-marked-up Knowledge Asset — one topic per page, clear headings, structured data identifying what it is and who it's about — is exactly the format modern AI search systems are built to parse and cite with confidence.

This isn't a cosmetic difference. It's the difference between content that exists and content that gets used.

The technical layer: schema and entity clarity

The single highest-leverage technical fix most place brands are missing is Organization schema with a sameAs property.

sameAs is a piece of structured data (written in JSON-LD, placed once on your homepage or about page) that tells AI systems: "this website is the same entity as this Wikidata entry, this Wikipedia page, this official social profile." It sounds small, but it solves a real problem — AI systems build confidence about who you are by cross-referencing entities across trusted sources. Without it, a model has to guess whether "Launceston Tourism" on your website is the same "Launceston" as the one on Wikipedia, or a different, unrelated thing entirely. Ambiguity makes models hedge, or skip citing you altogether.

You only need this canonical entity block once — on the homepage or about page — with other pages on your site linking back to it via a shared @id rather than repeating the whole block everywhere. Individual sub-entities (a named vineyard, a historic landmark, a signature festival) that have their own Wikidata or Wikipedia presence can carry their own sameAs too.

Getting a Wikidata entry claimed and correct, if you don't already have one, is probably the single best use of an afternoon on this whole list.

What to actually do, starting Monday

This week — audit and access

  1. Check robots.txt. Confirm CCBot, GPTBot, ClaudeBot, PerplexityBot, and Google-Extended aren't blocked — unless that's a deliberate choice, not an accident.
  2. Ask ChatGPT, Perplexity, and Google AI Overviews what your destination is known for. See if you're cited, and from which URL — your own site, or someone else's.
  3. Check Search Console crawl stats to confirm AI bots are actually reaching your pages, not just permitted to.

This week — structural fixes
4. Add Organization schema (JSON-LD) to your homepage: name, logo, and sameAs links to Wikidata, Wikipedia, and your official social profiles.
5. If you don't have a Wikidata entry, claim or create one — it's the backbone most AI systems cross-reference against.
6. Fix heading structure on your top ten pages: one H1, clear sequential hierarchy.

Next few weeks — content architecture
7. Convert your best five to ten Asset Library documents — fact sheets, guideline PDFs — into structured, single-topic Knowledge Assets with proper headings and schema.
8. Build or refresh comparative, editorial content: best time to visit, neighbourhood comparisons, data-led pieces. This is where AI citations concentrate.
9. Set a quarterly refresh cadence for your highest-value pages, and make the "last updated" date visible.

Ongoing — off-site presence
10. Prioritize placements on structurally authoritative third-party sites (Wikipedia, established travel and tourism publications) over sheer link volume.
11. Keep an eye on Reddit and YouTube conversation about your destination — you won't control it, but it shapes the co-citation picture AI systems draw on.
12. Set up a simple recurring check — weekly prompts across the major AI platforms — to track whether your citation presence is improving or slipping.

One decision to make deliberately
13. Decide, as policy rather than by accident, whether you want AI training crawlers indexing your site. Blocking has real consequences for future model training; allowing has content-control tradeoffs. Either can be the right call — but it should be a decision, not a default.

The bottom line

Common Crawl itself isn't something a brand needs to optimize for directly. It's the substrate underneath — the thing that quietly decided, years ago, which brands got baked into the base knowledge of today's AI models. T

he live battleground now is downstream of that: structured, freshly maintained, entity-linked content that AI answer engines can find, understand, and trust enough to cite.

That's not a crawler-access problem anymore. It's a content-architecture problem. Which is, unsurprisingly, exactly the problem a Knowledge Library is built to solve.


Kia ora — if you're building out your brand's Knowledge Library and want a hand thinking through structure or schema, get in touch.

Common Crawl, AI Visibility, and Your Place Brand: What Actually Matters in 2026

Common Crawl shapes how AI sees places: it’s the large, public web archive underpinning most AI training and citations. For place brands, the battle is twofold: ensure your site is crawlable for training data, and, crucially, secure ongoing AI citations by building a structured Knowledge Library with clear entity data (Organization schema with sameAs, Wikidata, etc.). Start Monday with a crawl-access check, add JSON-LD organization schema, create or claim Wikidata, and convert key assets into single-topic Knowledge Assets refreshed quarterly to win AI visibility and citations.

Asset type post
ID #887930
Word count 1608 words

Licence

Licence: Worldwide Paid and Unpaid Available to anyone for royalty free use in paid and unpaid media worldwide, provided Brandkit benefits from such use, and Brandkit is credited (optional).
Expiry: No expiry date
Release date:
Added at:
Updated at:

Tags

Loading