<!--
ROBOTS ON PAYROLL / THE VAULT
What this is: a data-backed checklist of what actually makes AI engines cite your pages, as of July 2026.
How to use it: download this file and give it to your AI (Claude, ChatGPT,
Cursor, Lovable chat) as a reference. Say: "Use this as my playbook for
making my pages citable by AI engines."
Latest version + full guide: https://robotsonpayroll.com/resources/seo-geo-engine
-->

# The GEO Checklist

What actually makes AI engines cite your pages, as of July 2026. Every item is backed by published data (Princeton GEO study, Ahrefs 1.4M-prompt study, Semrush, Peec AI); the "skip" section is backed by data too. Work top to bottom: structure and evidence move the needle most.

Score yourself: check every box on a money page before you call it done.

---

## 1. Content structure (how AI engines read your page)

LLMs retrieve at passage level, not page level. Every section must survive being ripped out of context.

- [ ] **Answer-first blocks.** Every H2 opens with a direct 40-80 word answer that makes sense with zero surrounding context. Depth comes after the answer, never before.
- [ ] **Question-phrased or claim-phrased H2s.** "How much does X cost in 2026" beats "Pricing considerations." A skimmer reading only H2s gets the full argument.
- [ ] **Self-contained chunks.** No section depends on the previous one to be understood. No "as mentioned above" logic.
- [ ] **Tables for any comparison** of 2+ things on 2+ attributes. Bullets for any list of 3+. These are the most extractable formats.
- [ ] **3+ quotable one-liners per post.** Self-contained, opinionated, one sentence, placed in the sections they support. Write them so a human would forward them and an AI would quote them verbatim.
- [ ] **Clean heading hierarchy.** One H1, logical H2/H3 nesting, no skipped levels, no headings used for styling.
- [ ] **Cover the query fan-out.** Google AI Overviews split one query into sub-queries and cite pages that appear across the most sub-query result sets. List the 5-10 sub-questions around your topic (cost, timeline, risks, alternatives, "vs", "for beginners") and answer each one on the page or in a tightly linked cluster. Only 38% of AIO-cited pages rank top-10 for the query (Ahrefs, 2026), so covering sub-questions can earn citations before rankings.
- [ ] **FAQ section** with question-phrased H3s and 40-80 word answers for the fan-out questions that didn't earn a full H2.

## 2. Evidence density (the strongest proven levers)

From the Princeton/Georgia Tech/IIT Delhi GEO study (Aggarwal et al., KDD 2024), still the benchmark in 2026.

- [ ] **Statistics with named source and year.** Adding sourced statistics lifted AI-answer visibility up to ~40%, the strongest single tactic tested. Format that wins: specific number + named source + year, inline. "(Ahrefs, 2026)".
- [ ] **Attributed expert quotes.** ~30-40% visibility lift. Real people, real attribution, not invented "experts say".
- [ ] **Cite your own sources.** Content that cites credible sources gets cited significantly more itself. Link out to the studies you reference.
- [ ] **No unsourced numbers.** Every figure in the post traces to a source or to your own test. Anything else gets cut.
- [ ] **Original data at least annually.** First-party data (your survey, your benchmark, your pricing teardown) is the only content AI engines can't get elsewhere. Publish one original-research piece per year and become the stat other pages cite.
- [ ] **No keyword stuffing.** The Princeton study found it actively reduced GEO visibility. Write for extraction, not density.

## 3. Freshness (a real ranking-of-citations lever)

- [ ] **Update money pages every 30 days or less.** ~76% of ChatGPT citations come from content updated within the last 30 days. ChatGPT-cited URLs averaged 458 days newer than Google's organic results (Ahrefs, 2026).
- [ ] **Visible, machine-readable dateModified.** On-page "Updated [date]" plus `dateModified` in Article schema.
- [ ] **Honest updates only.** Change actual content (a stat, a price, a section), not just the date. Fake freshness is a trust asset you only get to burn once.
- [ ] **Kill stale numbers.** Anything dated "202X" that's no longer true gets updated or removed on the monthly pass.

## 4. Technical prerequisites (table stakes, do once)

- [ ] **Allow AI crawlers in robots.txt:** GPTBot and OAI-SearchBot (ChatGPT), PerplexityBot, Google-Extended (Gemini grounding; AI Overviews use normal Googlebot), ClaudeBot.
- [ ] **Indexed in Bing.** ChatGPT's live search runs on Bing; no Bing index means no ChatGPT citations, full stop. Set up Bing Webmaster Tools and IndexNow.
- [ ] **Indexed in Google** (Search Console verified, sitemap submitted).
- [ ] **Server-side render key content.** Many AI crawlers execute little or no JavaScript. If your answer only exists after hydration, it doesn't exist.
- [ ] **Fast TTFB.** Crawlers with budgets skip slow origins.
- [ ] **Schema for SERP features, not for GEO.** Keep Article, Organization, FAQPage (where eligible), Product. But know the data: Ahrefs' May 2026 analysis of 1,885 pages found schema showed -4.6% differential in AI Overviews and +2.2% in ChatGPT, neither statistically significant. It earns rich results, not AI citations.

## 5. Entity clarity (help engines know who you are)

- [ ] **One consistent name** for your product/brand everywhere: site, socials, directories. No alternating between "Acme AI" and "AcmeApp".
- [ ] **Organization schema** on the homepage with `name`, `url`, `logo`, `sameAs` pointing to your real profiles.
- [ ] **An about page that states plainly** what you are, who runs it, and why you're credible. Engines resolve entities from plain statements, not vibes.
- [ ] **Author bylines with credentials** on posts, marked up with Person schema where it's true.

## 6. Off-site citation surfaces (the part you don't own)

Across ChatGPT, Gemini, Perplexity, and AI Overviews, the most-cited domains are Reddit, YouTube, LinkedIn, Wikipedia, Forbes (Peec AI, Semrush, Profound studies, 2026). Reddit alone is up to ~1 in 5 Perplexity citations.

- [ ] **Authentic Reddit presence** in your niche's subreddits: answer questions with real substance, disclose affiliation, link only when it genuinely answers the question. Astroturfing gets accounts and domains nuked.
- [ ] **YouTube versions of your best content.** YouTube is the single most-cited domain in AI Overviews (~18% of outside-top-100 citations).
- [ ] **LinkedIn distribution** of your original data and strongest claims.
- [ ] **Per-platform check:** only ~11% of domains are cited by both ChatGPT and Perplexity. Winning one engine does not transfer. Audit each engine separately (see the 10-prompt audit in the main resource).

## 7. Skip these (overrated, per 2026 data)

- [ ] **Confirmed you are NOT relying on schema markup to earn AI citations** (see section 4; keep it for SERP features only).
- [ ] **Confirmed you are NOT prioritizing llms.txt.** No major AI company (OpenAI, Google, Anthropic, Meta, Mistral) reads it in production as of Q1 2026. Of ~38,000 domains with a valid llms.txt, 97% got zero requests for it in May 2026. Google confirmed it will not support it. Only real use case: developer-docs retrieval in coding tools. Ship one in 20 minutes if you have docs; spend zero hours beyond that.
- [ ] **Confirmed you are NOT doing fluency-only AI rewrites** of existing pages and calling it optimization. The Princeton study found rewrites without added evidence do little; keyword stuffing made things worse.

---

## The monthly pass (30 minutes per money page)

1. Update at least one stat, price, or section. Bump dateModified.
2. Re-check the direct answer block still matches reality.
3. Run your 10-prompt citation audit (main resource, section 5) and log results.
4. Add one new sourced statistic or quote if the page is under-evidenced.

Source basis: Princeton GEO study (KDD 2024), Ahrefs 1.4M-prompt ChatGPT study and 4M-URL AIO study (2026), Ahrefs schema analysis (May 2026), Semrush/Peec AI/Profound citation studies (2026). Figures current as of July 2026.
