Blog/Web engineering

Generative engine optimization in practice: a full GEO audit of my own site

What generative engine optimization is, what the research really supports, and the GEO audit I ran on this site: 12 checks, 7 fixes, code included.

··12 min read

  • GEO
  • AI search
  • Structured data
  • nginx
A five-step pipeline from crawl to index, retrieve, cite and measure, showing where a page can drop out of an AI-generated answer.

Key takeaways

  • Generative engine optimization (GEO) means making pages easy for AI answer engines to retrieve, quote and cite; the term comes from a 2023 paper by Aggarwal et al., presented at KDD 2024.
  • In that paper, adding quotations, statistics and cited sources raised a page's visibility in generated answers by up to 40%, while keyword stuffing did worse than no optimization at all.
  • A July 2026 survey of 45 GEO studies found those gains hold only for pages that are already retrieved; no technique showed a stable effect on being found in the first place.
  • Google says AI Overviews and AI Mode need no special files or schema: a page must be indexed and eligible for a snippet, so the SEO basics are the entry ticket.
  • My audit of this site's 105 pages ran 12 checks and led to 7 fixes, mostly discovery links, honest dates and a way to see AI crawler traffic; the code for each is below.

Generative engine optimization (GEO) is the practice of making content easy for AI answer engines, such as Google AI Overviews and AI Mode, ChatGPT search, Perplexity and Copilot, to find, retrieve, quote and cite. Where SEO optimizes for a position in a list of links, GEO optimizes for being one of the few sources an answer is built from. The term comes from a 2023 research paper presented at KDD 2024.

This post has two halves: what the evidence supports, which is less than most GEO guides claim, and a full GEO audit of this site, a static Nuxt build with 105 pages in three languages. The audit ran 12 checks and led to 7 fixes; the code for each is below, and the checklist at the end is the one I would run on any site.

What is generative engine optimization?

A generative engine does not rank ten links. It decides whether a question needs a web search at all, sends one or more queries (Google calls this query fan-out), pulls candidate pages from an index, picks passages to put into the model's context, writes the answer and cites some of its sources. A page can drop out at every one of those steps, and each step has different levers.

Where a page can drop out of an AI answerFive steps from left to right: crawl, index, retrieve, cite, measure. Under each step is what a site owner controls: robots.txt and AI user agents for crawling; being indexable and snippet-eligible for indexing; on-topic text and a Markdown copy for retrieval; facts, sources and clear claims for citation; AI citation reports and agent logs for measurement. The typical failure at each step: blocked, no snippet, not relevant, vague text, invisible. Classic SEO covers crawl to retrieve, GEO extends from retrieve to measure.From crawl to citationCrawlrobots.txtAI user agentsblockedIndexindexablemax-snippetno snippetRetrieveon-topic textMarkdown copynot relevantCitefacts, sourcesclear claimsvague textMeasureAI citationsagent logsinvisibleWHAT YOU CONTROLTYPICAL FAILURESEOGEO
Five places a page can drop out of an AI-generated answer. Under each step: what a site owner controls, and the typical failure. Classic SEO covers crawling, indexing and retrieval; GEO adds what happens once a page is retrieved: whether it is cited, and whether you can see that it was.
SEOGEO
GoalA high position in a list of linksBeing one of the few sources an answer is built from
UnitThe pageThe passage or fact that gets quoted
Success metricPosition, clicksCitations, mentions, referral visits
Measured withSearch Console, analyticsBing AI Performance, server logs, repeated prompts

GEO does not replace SEO. For Google's AI features it sits on top of it: a page that is not indexed cannot be cited.

What does the research actually show?

The founding paper is GEO: Generative Engine Optimization by Pranjal Aggarwal and colleagues from Princeton and other institutions, first published in November 2023 and accepted at KDD 2024. They built GEO-bench, a benchmark of 10,000 queries, rewrote source pages with nine different methods and measured how visible each source became in the generated answer.

  • Quotations, statistics and cited sources work. The best methods improved on the unoptimized baseline by 41% on position-adjusted word count and by 28% on a subjective impression score; the headline figure is "up to 40%".
  • Keyword stuffing does not. Adding more query keywords, the classic SEO move, scored below the baseline.
  • Lower-ranked pages gain most. Citing sources raised the visibility of pages ranked fifth in the search results by 115.1%, while top-ranked pages lost 30.3% on average.
  • It carried over to a live engine. On Perplexity.ai, the same methods raised visibility by up to 37%.

Then the caveats arrived. Olivier Martinez's critical survey of 45 GEO studies (July 2026) argues that GEO is not a single ranking task but a noisy pipeline, and that the founding paper's gains are valid in its setting but conditional on a source already sitting in a fixed context. In the reviewed work, topical relevance and position in the context were the most reproducible levers. Generic rewriting heuristics transferred poorly, citation-oriented rewrites could even hurt retrieval, and no technique showed a stable, long-term, cross-platform effect on being discovered in the first place.

A March 2026 paper by Tian and colleagues points the same way from the other side. Instead of applying one rewrite to every page, their AgentGEO system diagnoses why a specific document is not cited and repairs that. It raised citation rates by over 40% relative while changing about 5% of the content, and the authors found that generic optimization can harm long-tail content.

What do Google, OpenAI and Microsoft say?

The platforms' own documentation is short and consistent.

  • Google says there are no additional requirements for AI Overviews or AI Mode: a page must be indexed and eligible to be shown with a snippet. You don't need new machine-readable files, AI text files or special schema.org markup; structured data must match the visible text; and nosnippet, data-nosnippet, max-snippet and noindex control what is shown. Both features may use query fan-out, and their traffic is counted in Search Console under the Web search type.
  • OpenAI uses OAI-SearchBot for ChatGPT search and GPTBot for training, and the two settings are independent. Sites that block OAI-SearchBot are not shown in ChatGPT search answers, apart from navigational links, and a robots.txt change takes about 24 hours to apply.
  • Microsoft added AI Performance to Bing Webmaster Tools as a public preview on 10 February 2026. It reports total citations in Copilot and Bing's AI answers, the average number of cited pages, the grounding queries the AI used to retrieve content, and citations per URL.

Note what Google leaves out: llms.txt and Markdown copies are not needed to appear in its AI features. They serve agents that fetch pages directly, which is a different audience; my post on llms.txt versus Markdown content negotiation has the log data. They cost little on a static site, so I keep them, but I don't count them as ranking levers.

How I ran the audit

I ran the audit on the production build, not the source code: nuxt generate writes 105 static HTML pages (35 pages in English, German and Hungarian), and a post-build step writes a Markdown copy of each one plus /llms.txt. Two small Node scripts then went through the output.

  • Technical audit over every generated HTML file: title and description, canonical and hreflang, robots directives, Markdown and llms.txt links, the JSON-LD graph (node types and dates), one h1 per page, and whether the Markdown copy exists.
  • Content audit over the 24 English posts in the blog database: does the first sentence define the topic, how many sources and inline citations, how many numbers, whether there are key takeaways and an FAQ, and how many h2 headings are questions.
// GEO audit over the generated site (excerpt): one pass over every HTML file
for (const file of htmlFiles) {
  const html = readFileSync(file, 'utf8')
  if (!/type="text\/markdown"/.test(html)) add('no rel=alternate text/markdown', page)
  if (!/rel="describedby"/.test(html)) add('no rel=describedby llms.txt', page)
  const robots = html.match(/<meta name="robots" content="([^"]*)"/)?.[1] ?? ''
  if (!robots.includes('max-snippet')) add('robots without max-snippet', page)
  for (const node of jsonLdGraph(html))
    if (/WebPage|CollectionPage|ProfilePage/.test(node['@type']) && !node.dateModified) add('page node without a date', page)
}

The checks follow the pipeline in the diagram: access, discovery, understanding, content and measurement. The findings below are in the same order.

What did the audit find?

CheckBeforeAfter
AI crawlers allowedPass: robots.txt allows everyone and names 16 AI user agents; TDMRep allows text and data miningUnchanged
Snippet eligibilityIndexable, but 33 of 105 pages set no max-snippet directivemax-snippet:-1 on every indexable page
Markdown copy per pagePass: 105 of 105Unchanged
Markdown discovery (rel="alternate")21 of 105 pages; no blog post or expertise page had itEvery page, plus an HTTP Link header
llms.txt discovery (rel="describedby")0 of 105; llms.txt was linked as rel="alternate" type="text/plain"Every page, plus the Link header
Entity graph (JSON-LD)Pass: Person, WebSite, WebPage and BreadcrumbList on every page; BlogPosting, FAQPage, citations and speakable on postsUnchanged
Dates on page nodes0 of 105 WebPage nodes had a datePosts and the blog index carry datePublished and dateModified
Sitemap lastmodEvery static page stamped with the build datelastmod only where there is a real date
Quotable summaryKey takeaways marked as speakable, but not in the schemaabstract built from the takeaways
Answer-first intros21 of 24 posts open with a definition; 3 are first-person build logsKept as they are
SourcesMedian of 6 sources per post; 3 posts cite fewer than 4Noted for their next revision
AI crawler visibilityNone: analytics loads only after consent, and crawlers don't run JavaScriptA separate nginx log for AI agents, with the Accept header

The audit also flagged about 30 page titles and descriptions that are longer than search results display. That is snippet hygiene rather than GEO, and I left it for a separate pass. The content side held up because the posts were written to a template with takeaways, an FAQ and a source list from the start; the gaps were almost all in the plumbing.

The fixes, step by step

Every page links its Markdown copy and llms.txt

Version 2 of llms.txt answers the question of how an agent finds a page's Markdown version with two standard link relations: rel="alternate" type="text/markdown" for the Markdown copy, and rel="describedby" for the llms.txt that covers the page, either as HTML link elements or as an HTTP Link header. The site had a plugin for the first one, limited to seven top-level pages. It now covers every page that has a copy, error pages excluded:

// app/plugins/agent-links.ts (excerpt)
const PAGE = /^(\/(about|references|game|accessibility|privacy|imprint|blog)|\/(blog|expertise)\/[a-z0-9-]+)?$/

if (!m || error.value || !PAGE.test(page)) return {}
return {
  link: [
    { key: 'markdown', rel: 'alternate', type: 'text/markdown', href: page ? `${prefix}${page}.md` : `${prefix}/index.md` },
    { key: 'llms-txt', rel: 'describedby', href: '/llms.txt', title: 'llms.txt' },
  ],
}

The same pair goes out as an HTTP header, so a client that only sends a HEAD request sees it too. One nginx detail needed a second look: try_files changes $uri to /about/index.html before the headers are written, so a map on $uri produces nothing for most pages. Keying the map on $request_uri avoids that:

# Keyed on $request_uri: try_files changes $uri to /page/index.html before the headers go out
map $request_uri $bc_agent_links {
  default                                  "";
  "~^/(\?.*)?$"                            '</index.md>; rel="alternate"; type="text/markdown", </llms.txt>; rel="describedby"';
  "~^(?<p>/[a-z0-9/-]*[a-z0-9])(\?.*)?$"   '<$p.md>; rel="alternate"; type="text/markdown", </llms.txt>; rel="describedby"';
}

location / {
  add_header Link $bc_agent_links;   # an empty value sends no header
  # ...
}

Snippet directives and honest dates

Google only uses a page in AI Overviews if it may show a snippet. Snippets are allowed by default, so max-snippet:-1 changes nothing in principle, but it states the intent on every page and matches what the blog posts already sent.

Dates were the more interesting gap. The sitemap gave every static page the date of the build, which is a false freshness signal: Google says it uses lastmod only when the value is consistently and verifiably accurate. Static pages now have no lastmod at all, and only the posts and the blog index, which have real dates, carry datePublished and dateModified in their structured data. No date is better than a wrong one.

A quotable summary in the structured data

Every post starts with five key takeaways, written as sentences that stand on their own. They were already marked as speakable, and the post's sources were already in the schema as citation. The takeaways now also go into the BlogPosting's abstract, so a system that reads the JSON-LD gets the summary without parsing the page:

{
  "@type": "BlogPosting",
  "headline": "Generative engine optimization in practice: …",
  "abstract": "Generative engine optimization (GEO) means … (the five key takeaways)",
  "datePublished": "2026-09-28",
  "dateModified": "2026-09-28",
  "citation": [{ "@type": "CreativeWork", "name": "GEO: Generative Engine Optimization", "url": "https://arxiv.org/abs/2311.09735" }],
  "speakable": { "@type": "SpeakableSpecification", "cssSelector": ["#takeaways"] }
}

Measuring what AI crawlers fetch

You can't improve what you can't see, and this site could not see AI crawlers at all: Google Analytics only loads after cookie consent, and crawlers don't run JavaScript anyway. The server log is the only honest source. nginx now writes requests from 16 AI user-agent patterns to a log of their own, with the Accept header, so it shows both which pages they fetch and who asks for Markdown:

# http context: which requests come from AI crawlers and agents
map $http_user_agent $bc_ai_agent {
  default 0;
  "~*(GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|…)" 1;
}
log_format bc_agents '$time_iso8601 $status $request_method $host$request_uri -> $uri $body_bytes_sent "$http_accept" "$http_user_agent"';

# server block: an access_log here replaces the inherited one, so the default is repeated
access_log /var/log/nginx/access.log;
access_log /var/log/nginx/balazscsorba-agents.log bc_agents if=$bc_ai_agent;

Two commands are enough to start with:

# The pages AI agents fetch most (after negotiation, so /about.md means "asked for Markdown")
awk '{print $6}' /var/log/nginx/balazscsorba-agents.log | sort | uniq -c | sort -rn | head -20
# Requests per agent
grep -oE '(GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User)' \
  /var/log/nginx/balazscsorba-agents.log | sort | uniq -c | sort -rn

Server logs show crawling, not citing. For citations, the AI Performance report in Bing Webmaster Tools is the one first-party number available; Search Console includes AI Overviews and AI Mode in its Web totals; and referrals from chatgpt.com, perplexity.ai and copilot.microsoft.com show up in analytics like any other referrer. Beyond that, the survey's advice applies: repeat the same prompts over time, with paraphrases, because single answers vary from run to run.

What I deliberately did not do

  • No keyword stuffing. It scored below doing nothing in the original GEO paper.
  • No blanket rewrite of all 24 posts. Generic rewrite rules transfer poorly and can hurt retrieval, according to the survey. The three posts with few sources get more when they are next revised, for the readers' sake.
  • No hidden text for language models. Text meant only for AI systems is cloaking by another name and uses the same trick as prompt injection. Everything a model can read on this site, a person can read too.
  • No structured data that isn't on the page. Google asks for markup that matches the visible text; the FAQ in the schema is the FAQ you see.
  • No invented entity data. The Person node lists only profiles that exist. More sameAs links are on my list, but only for accounts I actually use.

GEO audit checklist

  1. Let the answer engines in. robots.txt allows OAI-SearchBot, Claude-SearchBot, PerplexityBot and the rest, and no CDN bot filter overrides it.
  2. Be indexed and snippet-eligible. No stray noindex or nosnippet; check in Search Console and Bing Webmaster Tools.
  3. Serve the content as HTML. Crawlers that don't run JavaScript must see the text; static or server rendering does that.
  4. Open with the answer. The first sentence of a page defines its topic in the words someone would search for.
  5. Make claims specific and sourced. Numbers, dates, named sources and links are the part of GEO the research supports best.
  6. Keep structured data true. One connected entity graph, markup that matches the visible text, and dates only where they are real.
  7. Offer a clean copy. A Markdown version of each page, linked with rel="alternate", and an llms.txt, linked with rel="describedby".
  8. Log AI crawlers separately. A server log with the user agent and the Accept header, not client-side analytics.
  9. Track citations, not just rankings. Bing AI Performance, Search Console and referral traffic, checked over weeks.
  10. Re-run the audit after every build. The checks are scripts, so a regression shows up the same day.

For the agent side of the same work, see llms.txt versus Markdown content negotiation and the guide to WebMCP on a real site. If you want this audit run on your own site, get in touch.

Sources

  1. Aggarwal et al.: GEO: Generative Engine Optimization (KDD 2024)
  2. Martinez: Optimizing Visibility in Generative Engines, a critical survey of GEO 2023–2026 (July 2026)
  3. Tian et al.: Diagnosing and Repairing Citation Failures in Generative Engine Optimization (March 2026)
  4. Google Search Central: AI features and your website
  5. Google Search Central: Build and submit a sitemap
  6. OpenAI: Overview of OpenAI crawlers
  7. Bing Webmaster Blog: Introducing AI Performance in Bing Webmaster Tools (10 February 2026)
  8. llmstxt.org: Changes from v1 to v2

Frequently asked questions

What is generative engine optimization (GEO)?

Generative engine optimization is the practice of making content easy for AI answer engines such as Google AI Overviews and AI Mode, ChatGPT search, Perplexity and Copilot to retrieve, quote and cite. The term comes from a paper by Aggarwal et al., first published in November 2023 and presented at KDD 2024, which showed that adding quotations, statistics and cited sources can raise a page's visibility in generated answers by up to 40%.

Is GEO different from SEO?

It builds on SEO rather than replacing it. SEO aims for a high position in a list of links; GEO aims to be one of the few sources an answer is built from, so the unit is the quotable passage rather than the page. For Google's AI features the entry requirements are the same as for search: the page must be indexed and eligible for a snippet.

Does llms.txt help with generative engine optimization?

Not as a ranking signal. Google says no AI text files or special markup are needed to appear in AI Overviews or AI Mode, and log studies show most llms.txt files are never requested. llms.txt and Markdown copies help agents that fetch pages directly, which is a separate audience; on a static site they are cheap to keep.

What content changes improve AI citations?

The best-supported ones are specific, verifiable claims: statistics, quotations and links to credible sources, in text that clearly matches the question being asked. Keyword stuffing performed worse than no optimization in the original study. A 2026 survey of 45 studies warns that these gains apply to pages that are already retrieved and that generic rewrite rules transfer poorly, so measure on your own pages.

How do I measure GEO results?

Combine three sources. Bing Webmaster Tools' AI Performance report shows citations in Copilot and Bing's AI answers per page, with the grounding queries behind them. Server logs show which pages AI crawlers fetch. Referral traffic from chatgpt.com, perplexity.ai and similar sites shows up in analytics. Generated answers vary between runs, so repeat the same prompts over time instead of trusting a single answer.

Sounds like what you need?

Tell me about your project or role – I’d love to hear from you.