# How to get cited by AI answer engines (GEO)

> A step-by-step guide to getting cited by AI answer engines: what Google, ChatGPT, Claude and Perplexity actually document about how they select sources, the crawler user agents and the training-versus-retrieval distinction, whether llms.txt does anything, and why most GEO advice is folklore the vendors have contradicted in writing.

Canonical page: https://www.heybeagle.com/guides/how-to-get-cited-by-ai-answer-engines

Last reviewed: 2026-07-30

Get cited by the answer. What the vendors document, not what the blogs claim.

Generative engine optimisation is the most folklore-heavy corner of marketing right now. The good news is that the companies that run these engines have written down more than the industry admits, and most of it says the boring thing.

## The short answer

Almost everything sold as a special technique for ranking in AI answers is either unproven or something the vendors have said, in writing, that they ignore. Google's own May 2026 guidance is blunt: optimising for generative AI search is still SEO, structured data is not required, and it ignores AI-specific files including markdown and llms.txt. So the honest playbook has three real parts. First, be retrievable: these engines cite what their crawlers can fetch, and the standalone AI crawlers do not run JavaScript, so content that only renders client-side is invisible to them. Second, understand the access model: blocking a training crawler like GPTBot or ClaudeBot does not remove you from citations, because retrieval runs through different agents like OAI-SearchBot and Claude-SearchBot, and user-triggered fetches may ignore robots.txt entirely. Third, be the kind of source these systems reach for, which in practice means original, first-hand expertise and a presence in the places they lean on - and the single most-cited place, across study after study, is not your website. It is Reddit, Wikipedia, YouTube and the other communities where people who are not you vouch for things. GEO is mostly good SEO plus being genuinely talked about.

## At a glance

- **Is GEO a separate discipline?**: Not according to Google. Its May 2026 guidance says plainly that from Search's perspective, optimising for generative AI is optimising for the search experience, and thus still SEO. There is no special markup and structured data is not required.
- **Does llms.txt get me cited?**: No major AI vendor has publicly committed to reading it. Google says it ignores such files, and an analysis of ~137,000 sites found about 97% of llms.txt files received zero AI-bot requests. Treat it as unproven, not as a lever.
- **The crawler distinction that matters**: Training and retrieval are different agents. Blocking GPTBot, ClaudeBot or Google-Extended stops training use, not citation. Citation runs through OAI-SearchBot, Claude-SearchBot, PerplexityBot and, for Google, the normal Search index.
- **Do these crawlers run JavaScript?**: The standalone AI crawlers do not. Network data shows GPTBot, ClaudeBot and PerplexityBot fetch pages but do not execute JavaScript, so client-side-only content is invisible to them. Google's AI features are the exception, because they ride Googlebot's index.
- **What actually gets cited**: Community and reference sites, disproportionately. Studies disagree on the exact share, but Reddit, Wikipedia and YouTube dominate citation counts across engines. Being genuinely discussed there beats any on-page trick.
- **Can I measure it?**: Badly, for now. Google bundles AI Overview clicks into ordinary Organic Search with no separate label, and a large share of AI referrals arrive with no referrer and register as Direct traffic. Any citation-share number is a method-dependent estimate.
- **Biggest mistake**: Buying an AI-specific technique. The durable work is retrievability, originality and third-party mentions. The special files and schemas sold for this are mostly things at least one major vendor has said it ignores.

## Side by side

Six things you could spend effort on, and whether a primary source actually supports them. "Vendor-documented" is the column that separates the work from the folklore.

| Capability | Be retrievable and original | Manage crawler access | Structured data | An llms.txt file | Get discussed on Reddit etc. | Beagle |
| --- | --- | --- | --- | --- | --- | --- |
| Vendor-documented to help | Yes (Google, May 2026) | Yes (Per-agent docs) | Partly (Not required for AI) | No (Vendors ignore it) | Partly (Studies, not vendors) | No (Not a ranking tool) |
| Controls training vs citation separately | No (Not its job) | Yes (Different tokens) | No | No | No | No |
| Survives a JavaScript-only page | No (AI crawlers skip JS) | Yes | Partly (If server-rendered) | Partly | Yes (It is off your site) | Yes |
| Works without gaming anything | Yes | Yes | Yes | Yes (It just may do nothing) | Partly (Only if you are honest) | Yes |
| You control it directly | Yes | Yes | Yes | Yes | No (Others vouch, or not) | Partly (It drafts, you post) |
| Measurable today | Partly (Bundled into Organic) | Yes (Server logs) | Partly | Yes (Logs show zero hits) | Partly (Citation trackers) | Partly |

## What each one actually is

Sorted by how much a primary source backs it. The first two are documented work, the middle two are overhyped, and the last is the one nobody sells because you cannot buy it directly.

| Approach | What you get | Where it stops |
| --- | --- | --- |
| Be retrievable and genuinely original (Ongoing. It is just good SEO) | Server-render your content so the AI crawlers can read it, keep it indexable, and make it original. Google's sharpened bar is explicit: do not recycle what is already on the internet or what a model could produce itself; provide expert or first-hand takes that go beyond common knowledge. That originality is what a fan-out query rewards. | It is slow and unglamorous, and it is the same work SEO always was. There is no AI-specific shortcut layered on top, because Google says there is not one. |
| Manage crawler access deliberately (An afternoon. Then leave it) | Understand which user agents do what and set robots.txt on purpose. The key distinction: training crawlers (GPTBot, ClaudeBot, Google-Extended) are separate tokens from retrieval crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot). You can decline training while staying citable, or the reverse. | It is a permission layer, not a growth lever - it decides who may use you, not whether they choose to. And user-triggered fetchers like ChatGPT-User and Perplexity-User may bypass robots.txt because a person asked, so it is not an absolute block. |
| Structured data and clean technical SEO (Days. Worthwhile anyway) | Schema.org markup, canonical URLs, fast pages, sitemaps and IndexNow. All genuinely good practice that helps traditional search, which for Google is the same pipeline that feeds AI Overviews and AI Mode. | Do not expect an AI-specific payoff. Google states outright that structured data is not required for its generative features and there is no special schema for them. IndexNow's confirmed consumers are Bing, Yandex, Naver and DuckDuckGo, not Google and not the AI engines directly. |
| An llms.txt file (An hour. Possibly for nothing) | A markdown file listing your key pages, proposed in 2024 to give models a clean, compact view of a site. Cheap to add, and harmless. | No major AI vendor has publicly committed to consuming it. Google says it ignores such files, and roughly 97% of llms.txt files across ~137,000 sites got zero AI-bot requests. Ship one if you like - we do - but do not count it as a reason you get cited, and do not let anyone sell it to you as one. |
| Be genuinely discussed off your own site (Slow. Not directly buyable) | The uncomfortable finding behind every citation study: the most-cited sources in AI answers are community and reference sites - Reddit, Wikipedia, YouTube - not brand-owned pages. Being real, useful and talked about in those places is the closest thing to a durable edge, precisely because you cannot fake it cleanly. | You do not control it, and trying to control it dishonestly is the behaviour Reddit and the FTC both punish. The studies also disagree on the exact numbers, so treat the direction as solid and any specific percentage as an estimate. |
| An AI teammate that helps you show up honestly (Minutes. Any Slack or Teams workspace) | Not a citation trick. Beagle watches where your category is discussed, drafts genuinely useful contributions for a human to post, and helps you keep the original, first-hand content the engines reward flowing. It works on the two things you actually control: being present, and being worth quoting. | It cannot make an engine cite you, and it does not manufacture presence with fake accounts. It removes the friction from doing the real work, which is the only kind that lasts here. |

## Step by step

Start by deleting the folklore, then do the boring things that a vendor has actually put in writing. In order:

1. **Confirm an AI crawler can actually read the page** - This is upstream of everything, and it is where JavaScript-heavy sites quietly fail. The standalone AI crawlers fetch pages but do not execute JavaScript, so if your content only appears after client-side rendering, they see an empty page and you cannot be cited. Server-render or statically render anything you want quoted. View the raw HTML, not the rendered page, to check. Test: curl the URL. If the content is not in the HTML, the AI crawlers do not see it.
2. **Set robots.txt on the training-versus-retrieval distinction** - Decide deliberately, because the two are different tokens. Blocking GPTBot, ClaudeBot or Google-Extended declines training use and does not remove you from citations. Blocking the retrieval agents - OAI-SearchBot, Claude-SearchBot, PerplexityBot - is what actually affects whether you can be cited. Most brands that want visibility should allow the retrieval crawlers, whatever they decide about training. Note: User-triggered fetchers may ignore robots.txt, so it is not an absolute wall.
3. **Do the SEO, because Google says it is the same pipeline** - Indexable, server-rendered, canonicalised, fast, in the sitemap. Google's guidance is explicit that a page must be indexed and eligible to show in Search with a snippet to appear in AI features, and that there are no additional requirements or special optimisations beyond that. So the AI-features work and the ordinary search work are the same work.
4. **Raise the originality bar, on purpose** - Google's sharpened instruction is to provide unique expert or experienced takes that go beyond common knowledge, and specifically not to recycle what is already on the internet or what a model could produce itself. This is the one content lever that is both AI-relevant and vendor-documented. First-hand data, a real opinion, a thing only you know: that is what a fan-out query is reaching for. Why: AI Mode fans one question into many sub-queries and stitches the best sources
5. **Add structured data for search, not for a mythical AI bonus** - Mark up what deserves it, because it helps ordinary search and costs little. But do not let a vendor sell it to you as the AI-citation unlock, because Google states plainly that structured data is not required for generative features and there is no special schema for them. Same for markdown variants and llms.txt: fine to have, but Google says it ignores them, so do not build a strategy on them.
6. **Invest in being discussed where the engines actually look** - This is the highest-leverage and least comfortable step, because it is off your own site. The citation studies keep landing on the same place: community and reference sites carry a hugely disproportionate share of AI citations, with Reddit at or near the top. Earning a genuine presence there - honestly, the way our Reddit guide describes - does more for AI visibility than any file you can add to your domain. See: The Reddit guide, for how to do that without getting banned
7. **Measure with humility, and log what you can** - Accept that the numbers are soft. Google bundles AI Overview clicks into Organic Search with no separate label, and a large share of AI-driven visits arrive with no referrer and count as Direct. Your cleanest signal is your own server logs - which retrieval crawlers fetched what - plus citation-tracking tools read as directional, not precise. Do not report a citation-share figure as if it were a fact.
8. **Re-check the crawler docs on a schedule** - This field moves under you. OpenAI has revised its crawler documentation more than once, Cloudflare is changing its default AI-crawler posture, and the studies flip on which source leads. Put the primary docs - Google Search Central, and each vendor's bot page - on a quarterly re-read, because a guide written today, including this one, dates fast.

## What trips people up

The specific pieces of folklore that waste the most money, each one contradicted by a primary source.

- **llms.txt is not a citation lever** - It is sold as one constantly. But no major AI vendor has publicly committed to reading it, Google says it ignores such files, and when someone checked ~137,000 sites, about 97% of their llms.txt files had received zero AI-bot requests. We publish one on this very site, and we would not tell you it is why anything gets cited. Add it because it is cheap and harmless, not because it works.
- **Blocking GPTBot does not remove you from ChatGPT's answers** - The most common access mistake, in both directions. GPTBot is a training crawler; ChatGPT's search citations come through OAI-SearchBot. Block the first and you decline training while remaining citable; block the second and you affect citation. They are separate robots.txt tokens, and treating them as one is how sites either leak training data they meant to keep or vanish from citations they meant to keep.
- **Your beautiful JavaScript site may be invisible** - This is the silent one. The standalone AI crawlers do not run JavaScript, so a page whose content is injected client-side reads as blank to them, no matter how good it looks in a browser. Google's AI features are partly shielded because they use Googlebot's rendered index, but ChatGPT, Claude and Perplexity are fetching raw HTML. If it is not in the source, it does not exist to them.
- **There is no special AI schema** - Someone will try to sell you one. Google's words: structured data is not required for generative AI search, and there is no special schema.org markup you need to add. Mark up your content for ordinary rich results if it qualifies, and ignore anyone promising a GEO-specific schema, because the company running the largest AI answer surface has said in writing that it does not exist.
- **Citation is not traffic, and the numbers are soft** - Two traps in one. Being cited does not mean being visited - reported click-through on citations can be low - and the citation-share studies disagree wildly by method, from forums being 2% of citations in one to over 40% in another. Anyone quoting you a precise percentage of AI answers you could capture is selling certainty that the data does not support.
- **The access rules are being rewritten right now** - Do not treat any of this as settled. Cloudflare began blocking AI crawlers by default for new domains in 2025 and is changing the default again in 2026, publishers are starting to charge per crawl, and Reddit's licensing deals with the AI companies are themselves in renewal talks. The plumbing that decides who can cite you is a live negotiation, not a fixed backdrop.

The uncomfortable truth of getting cited by AI is that the strongest move is not on your website at all. It is being the thing other people bring up when nobody made them.

## Or skip the build

The two durable levers here - original first-hand content, and a genuine presence where the engines look - are both work, not tricks. Beagle is built to take the friction out of both without faking either.

- **It surfaces where your category is discussed** - Beagle watches the communities and threads the answer engines lean on and brings the relevant ones into your Slack or Teams, so you can contribute honestly while the conversation is live rather than discovering it in a citation months later.
- **It drafts the first-hand contribution** - The original, expert content Google says it rewards is exactly what is hard to produce at volume. Beagle drafts from what you actually know - your data, your docs, your decisions - so a real person is editing something substantive rather than staring at a blank page.
- **A human posts, and discloses** - Beagle never manufactures presence with fake accounts or auto-posts under a fabricated identity, because that is the behaviour that gets you removed from the very places you are trying to be cited from. The work stays honest, which is the only version of it that compounds.
- **It keeps the real content flowing** - Answer engines reward sites that keep publishing genuine, updated expertise. Beagle helps turn what your team already knows into that steady stream, which is the same thing that helps ordinary search - because, as Google keeps saying, they are the same thing.

## In your stack

Teams here usually connect: Slack, Notion, Drive, Linear, HubSpot, Gmail. Beagle connects to ~3,200 tools in total - see https://www.heybeagle.com/integrations.

## FAQ

**Is GEO or AEO actually different from SEO?**

Not in the words of the company that runs the largest AI answer surface. Google's May 2026 guidance states that from Search's perspective, optimising for generative AI search is optimising for the search experience, and thus still SEO. It adds that there is no special markup, that structured data is not required for generative features, and that you do not need AI-specific files. So the honest answer is that GEO is mostly good SEO plus originality plus off-site presence, not a separate discipline with its own toolkit.

**Should I add an llms.txt file to get cited?**

You can, but do not expect it to do anything, and do not pay for a strategy built on it. No major AI vendor - Google, OpenAI, Anthropic or Perplexity - has publicly committed to reading llms.txt. Google says outright that it ignores such files. And an analysis of around 137,000 sites found that roughly 97% of their llms.txt files received zero AI-bot requests. It is cheap and harmless to publish one, which is why we do, but treating it as the reason you get cited is exactly the folklore this guide is warning about.

**If I block GPTBot, do I disappear from ChatGPT?**

No, and this is the most important distinction in the whole topic. GPTBot is OpenAI's training crawler; the citations in ChatGPT's search come through a different agent, OAI-SearchBot, which OpenAI documents as not being used for training. They are separate robots.txt tokens. So you can block GPTBot to decline training while remaining fully citable through OAI-SearchBot, or vice versa. The same split exists for Anthropic, where ClaudeBot is training and Claude-SearchBot is retrieval, and for Google, where Google-Extended controls only Gemini training and explicitly does not affect Search inclusion or ranking.

**Do AI crawlers read JavaScript?**

The standalone ones do not. Network measurements show that GPTBot, ClaudeBot and PerplexityBot fetch pages but do not execute JavaScript, so any content that only appears after client-side rendering is invisible to them and cannot be cited. The important exception is Google's AI Overviews and AI Mode, which are built on Google's own index and therefore benefit from Googlebot's rendering. The safe rule is to server-render or statically render anything you want an answer engine to quote, and to check the raw HTML rather than the rendered page.

**What sources do AI answers actually cite most?**

Community and reference sites, disproportionately, with Reddit at or near the top across most studies, alongside Wikipedia and YouTube. The exact numbers vary a lot by method - one analysis put forums at around 2% of citations while another put Reddit alone above 40% - so treat the direction as reliable and any single figure as an estimate. The practical implication is uncomfortable but consistent: a genuine presence in the places people discuss your category tends to matter more for AI visibility than anything you can add to your own domain.

**How do I measure whether this is working?**

Carefully, and without over-claiming. Google bundles AI Overview and AI Mode clicks into ordinary Organic Search in Search Console with no separate label, and a large share of visits driven by AI answers arrive with no referrer and get counted as Direct traffic, so standard analytics understate it. Your most reliable first-party signal is your own server logs, which show exactly which retrieval crawlers fetched which pages. Third-party citation trackers exist and are useful as a direction of travel, but no tool gives you a clean, authoritative citation-share number today.

## Read next

- [Marketing on Reddit](https://www.heybeagle.com/guides/how-to-market-on-reddit-without-getting-banned): The single most-cited source, done right
- [ChatGPT and your knowledge](https://www.heybeagle.com/guides/how-to-give-chatgpt-your-company-knowledge): Retrieval, from the other direction
- [Our own llms.txt](https://www.heybeagle.com/llms.txt): We ship one, and we told you the truth about it

Try Beagle free: https://www.heybeagle.com/signup (1,000 credits, no credit card).
