AI search and GEO
AI search tools answer questions in their own words and link to some of their sources. Being one of those sources is becoming a second way for inbound content to be found. This page explains how it works and what you can influence — without promises, because nobody can guarantee a place in an AI answer.
What counts as AI search
“AI search” means services that answer queries with generated text drawn from web sources:
- Google AI Overviews — an AI summary above the regular results, rolled out in the US from May 2024 and shown only where Google judges it additive.
- Google AI Mode — conversational search for complex questions and follow-ups, rolled out to everyone in the US from May 2025.
- ChatGPT search — web search inside ChatGPT since October 2024.
- Perplexity — a search service built around AI answers with linked sources.
- Microsoft Copilot — Microsoft's assistant, formerly Bing Chat, drawing on Bing's index.
How an AI answer is assembled
The documented pattern is similar across providers:
- Retrieval. The prompt becomes one or more searches. Google calls this query fan-out: related searches across subtopics at once. Bing reports such grounding queries to site owners.
- Grounding. A language model writes the answer from the retrieved pages, not from its training data alone — retrieval-augmented generation (RAG). Google says its core ranking systems do the retrieving.
- Citations. The answer links to supporting pages, often from several sites.
GEO, AEO and SEO
Generative engine optimisation (GEO) means making content visible in AI-generated answers. The term was introduced in a 2023 research paper by Aggarwal and colleagues; Microsoft now uses it for the AI reports in Bing Webmaster Tools. Answer engine optimisation (AEO) means much the same, with no agreed line between the two.
Google states that, from its perspective, optimising for generative AI search is still SEO, since its AI features rest on its core ranking systems; the fundamentals are on the SEO page. What changes is the measure of success: being cited inside an answer rather than ranked in a list, which tends to favour content that is easy to quote, check and attribute.
What tends to get cited
No provider publishes a formula; these are recurring observations, not guarantees:
- Clear, direct answers near the top of a section. Microsoft notes that assistants can lift question-and-answer pairs word for word.
- Well-structured pages with descriptive headings, short paragraphs, lists and tables — though Google says content need not be chopped into tiny chunks.
- Original data and first-hand experience. Google favours “non-commodity” content with its own point of view over rehashed common knowledge. In the GEO paper's benchmark, adding statistics, quotations and sources helped most; keyword stuffing did little.
- Clear entities — consistent names and facts for your organisation, products and people across text, images and video, as Bing advises.
- Freshness — Bing recommends keeping content current and notifying search engines of changes.
- Crawlability — Google's AI features only link to indexed pages that are eligible for a snippet; Microsoft advises against hiding key answers in tabs, images or PDFs.
From keywords to prompts
AI tools invite full questions; Google describes AI Mode as built for complex, multi-part questions and follow-ups. Keyword research still shows demand, but add three layers:
- Conversational prompts — the whole question, with context such as budget, team size or industry.
- Follow-up questions — what people ask next; natural material for FAQs.
- Topics and entities — the concepts and names a subject involves. Bing groups grounding queries into topics because, in its words, AI systems reason across concepts and themes rather than isolated keywords.
Resist building a page per prompt variant: Google warns that mass-producing pages for fan-out queries can breach its scaled content abuse policy. Real questions from sales, support and community threads are a better guide.
Formats that suit AI answers
Some formats map neatly onto typical prompts:
| Format | Typical prompt | What makes it work |
|---|---|---|
| Definition | “What is lead scoring?” | A self-contained answer in the first sentence |
| Comparison | “MQL or SQL: what's the difference?” | One table, same criteria for each option |
| Step list | “How do I set up a welcome sequence?” | Numbered steps in order |
| FAQ | Follow-up questions | Real questions, short direct answers |
| Data table | “What is a typical conversion rate?” | Original figures with date and method |
The inbound angle. AI answers work mostly at the attract stage, while people research a problem. A citation builds awareness even without a click. Google's AI features can also reflect what is said about products across the web, so what others genuinely say about you matters alongside links; chasing inauthentic mentions does not help. The durable route is classic inbound: useful content, real reviews and coverage, and consistent facts about your organisation.
Crawlers, robots.txt and other controls
Several providers separate crawling for model training from crawling for search, each with its own robots.txt token:
| Token | Operator | Documented purpose |
|---|---|---|
| GPTBot | OpenAI | Content that may be used to train OpenAI's models |
| OAI-SearchBot | OpenAI | Showing sites in ChatGPT search; blocked sites are left out of its answers |
| PerplexityBot | Perplexity | Surfacing and linking sites in Perplexity; not used for model training |
| Google-Extended | A control token, not a crawler: Gemini training and grounding in Gemini apps; no effect on inclusion or ranking in Google Search |
So a site can disallow GPTBot, signalling that its content should not be used for training, yet allow OAI-SearchBot and stay eligible for ChatGPT search. User-triggered fetchers such as ChatGPT-User and Perplexity-User may ignore robots.txt, which only works if crawlers obey it anyway.
Google's AI Overviews and AI Mode use the normal Googlebot crawl, so blocking Google-Extended does not keep a site out of them. The controls there are the usual snippet and indexing rules (nosnippet, data-nosnippet, max-snippet, noindex) and, since 2026, a Search Console setting that excludes a site from Google's generative AI features without touching its regular rankings.
llms.txt is an unofficial 2024 proposal for a Markdown file that guides language models through a site; it is not a web standard. Google states that its Search, AI features included, ignores the file: it neither helps nor harms there.
Measuring AI visibility, and its limits
- Analytics. Since May 2026, Google Analytics 4 puts visits from assistants such as ChatGPT, Gemini and Copilot in a default AI Assistant channel, excluding Google's AI Overviews and AI Mode. Visits without referrer data, such as from apps or copied links, still count as direct.
- Search Console. AI Overviews and AI Mode count towards the regular Performance report. A separate Generative AI performance report, rolled out worldwide by the end of August 2026, shows their impressions by page, country and device, but no clicks.
- Bing Webmaster Tools. The AI Performance report, in preview since February 2026, counts citations in Copilot and Bing's AI summaries, with the grounding queries behind them.
- Mention and citation tracking. Elsewhere, sample: re-run a fixed set of realistic prompts and log mentions, cited pages and wording. Commercial trackers automate this; treat their figures as samples.
Many answers end the search without a visit: in a Pew Research Center analysis of 900 US adults' browsing in March 2025, people clicked a regular result on 8% of visits to Google results pages with an AI summary, against 15% without one, and a link inside the summary on 1%. Answers shift with wording, user, model updates and time, and Google itself says placement is never guaranteed. Judge AI visibility by the metrics that matter — qualified visits, leads, revenue — not by citations alone.
Frequently Asked Questions
- Is GEO different from SEO?
Mostly not. GEO and AEO describe the effort to appear in AI-generated answers, but the groundwork is the same as for SEO: crawlable, indexable pages and clear, genuinely useful content. Google says optimising for its AI features is still SEO. The differences lie in emphasis — being quoted and cited rather than ranked — and in how success is measured.
- Should I block AI crawlers in robots.txt?
It depends on what you want to allow. Training crawlers and search crawlers use separate tokens, so you can decide on each. For OpenAI, for example, blocking OAI-SearchBot keeps a site out of ChatGPT's search answers, while blocking GPTBot only signals that its content should not be used for training. Blocking Google-Extended does not keep a site out of Google's AI Overviews or AI Mode; Search Console has a separate setting for those.
- Do I need an llms.txt file?
Not for Google: it states that Google Search, including AI Overviews and AI Mode, ignores llms.txt, so publishing one neither helps nor harms there. llms.txt is an unofficial proposal from 2024 that some other tools read, mostly for software documentation. A crawlable site with clear pages matters far more.
- Does structured data help with AI answers?
It is not required. Google says no special schema.org markup is needed for its AI features, while still recommending structured data for rich results; Microsoft describes schema as helping its systems understand what a page represents. Where you use it, keep it accurate and consistent with the visible text.