The GEO Checklist: 14 Things That Get You Cited by AI
This GEO checklist covers fourteen changes that get a page cited by ChatGPT, Perplexity and Google AI Overviews, ordered by effort against impact rather than by topic.
The short version: items 1 to 5 take an afternoon and account for most of the movement: answer the question in the first two sentences, put a real number in that answer, write headings as questions, add a 40 to 60 word summary under the H1, and make your entity details identical everywhere. Items 6 to 14 are the compounding work: honest schema, a named author, real tables and lists, self-contained passages, original data, topical depth, published pricing, and brand mentions. llms.txt is not on the list, and the reason is explained below.
Told when this changes?
The engines keep moving and this list gets revised when they do. Leave your email and I’ll let you know what changed.
Nothing lands in your inbox today. The checklist is right here on the page, free and complete.
Most GEO advice is a list of tactics with no ranking of effort against impact, which leaves you doing the easy things that don’t matter and skipping the boring things that do.
This is the order I work in on client sites. Items 1 through 5 take an afternoon and account for most of the movement. Items 6 through 14 are the compounding work.
The three engines do not pick sources the same way
Optimising for “AI search” as though it is one destination is the most common mistake I see.
Google AI Overviews
Pulls overwhelmingly from what already ranks in the top 20 organic results. If you rank, you are in the candidate pool. If you don’t, no amount of GEO work will put you there.
Lever: existing SEO
Perplexity
Heavily favours freshness. Recently published and recently updated content gets cited disproportionately, which makes publishing cadence a direct visibility input.
Lever: cadence
ChatGPT
Leans on training data and high-authority publishers, which is why a large share of its most-cited pages are Wikipedia, government, education and major news, places you cannot pitch your way into.
Lever: authority and entity clarity
Profound ran 100,000 prompts through both engines and found ChatGPT and Perplexity landed on the same domain just 11% of the time. That is domain-level agreement, so page-level agreement is lower still. Same query, different worlds, and the checklist below covers all three.
Do these five first
Highest impact for the least effort. Run them on your top 20 pages by impressions before you touch anything else.
Answer the question in the first two sentences
Open every page and every H2 section with the direct answer. Not context, not a preamble, not a story about why the topic matters.
Models extract passages. A passage that contains the complete answer can be lifted and cited. A passage that sets up an answer arriving three paragraphs later cannot.
The test: delete everything except your opening two sentences. Does what remains answer the question a reader typed?
Put a real number in that answer
Specific figures get cited far more than qualitative claims. “Significantly faster” is unquotable. “Cut load time from 4.1s to 1.3s” is quotable.
If you have no number, that is a content problem, not a formatting problem. Go get one. Run the test, audit the accounts, count the thing.
Make every H2 a question someone would actually type
Not “Implementation Considerations.” Try “How long does it take to implement?”
Your headings become the retrieval anchors. Match them to real query phrasing and the extraction gets easier for the model.
Add a 40 to 60 word summary block under the H1
One tight paragraph that answers the page’s core question completely, before any subheading. This is the single highest-leverage change on most pages and it takes ten minutes.
Fix your entity consistency
Your company name, what you do, your location and your key claims must be identical across your website, your Google Business Profile, your LinkedIn, your directory listings and any store listing you have.
Models cross-reference. Contradictions between sources create uncertainty, and uncertainty is a reason to cite someone else instead.
Quick audit: search your brand name in ChatGPT and ask it to describe your business. Whatever it gets wrong is what is inconsistent somewhere.
Make the page machine-readable
Slower to ship than Part 1, but this is what makes a passage safe to extract and attribute.
Ship real schema, not decorative schema
Organization, Article, FAQPage and Product where applicable. Fill the properties honestly.
Two rules matter more than the markup itself: never mark up content that is not visible on the page, and never publish review schema for reviews you do not have. Both are the kind of thing that costs you trust for years to save you a week.
Name your author, with credentials that can be verified
Anonymous or generic bylines underperform. A named author with a linked profile, real credentials and a body of published work outperforms, and a large majority of top-ranking pages now carry detailed author bios.
An author box with a stock photo and a two-line generic bio is not the point. Verifiability is. Link to something external that confirms the person exists and knows the subject.
Add a “last updated” date and mean it
Perplexity in particular rewards freshness. But changing a date without changing the content is a bad idea, because models and readers both compare.
Update genuinely, then update the date.
Structure lists and tables properly
Use real <ul>, <ol> and <table> markup rather than styled divs or paragraph text with dashes. Comparison tables get extracted well when they are actual tables.
Keep passages self-contained
Each section should make sense if lifted out of the page alone. Avoid opening a section with “As mentioned above” or “Building on this”, because that passage becomes useless the moment it is extracted.
What actually separates cited sites from ignored ones
These take months, not afternoons. They are also the reason some sites get cited for years while their competitors churn out posts nobody quotes.
Publish something the internet does not already have
This is the one that actually moves things, and it is the one most people skip because it is work.
Original data from a survey you ran. A case study with real numbers from real client work. An audit of a sample of accounts. A framing nobody else used. Write the “when this is the wrong answer” post instead of another “what is” post.
One original data point outperforms twenty rewrites. If your page could be assembled from the top five results, models have no reason to prefer it.
Build depth in one subject area rather than breadth across many
Narrow, expert destinations are outperforming broad shallow sites. Cover every angle and sub-topic of one thing before you start a second thing.
Publish your pricing and your case studies
Counterintuitive, but these outperform top-of-funnel guides for AI-referred traffic. Bottom-of-funnel pages get cited when someone asks a buying question, and buying questions are where the revenue is.
Most agencies hide pricing. That is a citation you are handing to a competitor.
Track brand mentions, not just links
AI systems weight how often and how consistently your brand is discussed across the web. Unlinked mentions count. Podcast appearances, community answers, guest pieces, being quoted in someone else’s article. All of it feeds the same signal.
Running all 14 by hand does not scale past about 20 pages
Items 1 through 10 are mechanical. They are also the ones that break down at volume. A 400-page site means 400 opening paragraphs to rewrite, 400 headings to rephrase and 400 schema blocks to check. That is the point where most GEO projects quietly stop.
So I built wptaskify, an AI SEO and GEO tool for WordPress, WooCommerce and Shopify. It connects your site to Claude or ChatGPT and runs this checklist as actual operations on your pages rather than as a report telling you what to fix.
It is bring-your-own-AI, so you are not paying a markup on tokens, and it works on live sites without a rebuild.
Full disclosure: wptaskify is our product. The checklist above works whether you use it or not. Every item is doable by hand, and the manual method is described in enough detail to follow without buying anything.
Why llms.txt is not on this list
Ship it if you want, because it does real work helping coding agents like Cursor, Copilot and Claude Code fetch your documentation efficiently. That is a genuine use case.
It is not an AI search visibility lever. Gary Illyes of Google said plainly in July 2025 that Google does not support llms.txt and is not planning to.
The usage data says the same thing. Ahrefs checked 137,210 domains and found 97% of llms.txt files received no traffic at all in May 2026. Dries Buytaert measured his own logs and counted about 5,000 llms.txt requests out of 400 million, or 0.001% of traffic. Nearly two years after Jeremy Howard proposed it in September 2024, the file is widely published and almost never fetched.
| What llms.txt is for | What it is not for |
|---|---|
| Helping coding agents fetch docs efficiently | Getting cited in ChatGPT or Perplexity answers |
| A curated map of your documentation | A ranking or visibility signal in Google AI Overviews |
| Reducing token waste for agents reading your site | A substitute for items 1 to 14 above |
Right file. Wrong job.
Figures in this section verified against their primary sources on 16 August 2026. This space moves fast, so if you are reading much later, check them again.
One warning if you have already implemented it
The common implementation of generating individual markdown copies of every page will create duplicate content at scale if those files are indexable. If you have done that, check your index coverage.
How to actually run this
- Open Google Search Console and sort by impressions.
- Take your top 20 pages, the ones that already have demand.
- Run items 1 through 5 on all 20. That is your afternoon.
- Work down items 6 to 14 over the following weeks, highest-traffic pages first.
- Measure before you scale. Do not do this to your whole site at once. Do it to the pages that already have demand and watch what happens.
Want to know when this list changes?
Bookmark the page, or leave your email and I’ll tell you when an item gets added, dropped, or reordered.
No welcome email, no sequence. Your address sits in a list until there is genuinely something to say.
Common questions
What is GEO and how is it different from SEO?
GEO, or generative engine optimization, is the work of getting your content cited inside AI-generated answers rather than ranked in a list of blue links. It overlaps heavily with SEO, because Google AI Overviews draw almost entirely from pages already ranking in the top 20, but it adds requirements SEO does not have, mainly that a passage must answer the question completely on its own so it can be lifted and attributed.
Do I need llms.txt to get cited by AI?
No. Gary Illyes of Google said in July 2025 that Google does not support llms.txt and is not planning to, and Ahrefs found that 97% of llms.txt files across 137,210 domains received no traffic at all in May 2026. The file is genuinely useful for helping coding agents fetch documentation efficiently, but it is not an AI search visibility lever.
How long does GEO take to show results?
Items 1 to 5 are structural changes to pages that already rank, so movement in AI Overviews can appear within a few weeks of recrawl. Items 11 to 14, which cover original data, topical depth and brand mentions, compound over months. Anyone promising citation results in days is describing something else.
Does GEO work if my pages do not rank on Google yet?
Partly. Perplexity and ChatGPT can cite pages that do not rank well, so freshness, entity clarity and original data still help. But Google AI Overviews pull from the top 20 organic results, so for that engine specifically you need conventional ranking first. Fix the ranking problem before expecting AI Overview citations.
Can I run this checklist on a large site without doing it page by page?
Items 1 to 10 are mechanical and can be automated. That is what we built wptaskify for. It connects a WordPress, WooCommerce or Shopify site to Claude or ChatGPT and applies answer-first rewrites, schema and GEO audits in bulk. Items 11 to 14 cannot be automated, because original data and brand mentions are earned rather than generated.
Something not applying to your setup?
Send me a message and tell me your stack. I’ll tell you what to do instead.
I build MCP servers and run technical SEO on WordPress, WooCommerce and Shopify sites. If you would rather someone just did this for you, that is the conversation.
