Request Quote
Lume Digital

Blog

Checklist for auditing whether AI assistants cite your business

The AI Citation Audit: 27 Checks That Decide Whether ChatGPT Recommends Your Business

If ChatGPT recommends your competitors instead of you, it is almost never because your business is worse. It is because their website is easier for an AI system to crawl, identify and quote. The 27 checks below cover the four things that decide this — whether AI crawlers can reach you, whether they can tell who you are, whether your content can be extracted, and whether they have any reason to trust you.

You can work through the whole list in an afternoon.

Why AI citations now decide who gets the enquiry

For twenty years, being found meant ranking. A buyer typed a question, scanned ten blue links, and clicked one. That flow is breaking down, and the numbers are not subtle.

When an AI Overview appears on a Google result, click-through rate drops by 46.7% in relative terms — from 15% down to 8% — according to Pew Research data covering 68,000 queries. Ahrefs measured a 34.5% click reduction for position-one informational keywords. Similarweb tracked zero-click searches rising from 56% to 69% in a single year.

Ranking first is no longer the same thing as being found.

But one number moves the opposite way, and it is the one worth building a strategy around. Amsive found that branded queries with an AI Overview saw an 18% increase in click-through rate. Generic informational traffic is being absorbed. Demand for businesses people already know by name is growing.

That is the whole game now. When a buyer asks an AI assistant “who should I hire for X,” you either come out of that answer as a named recommendation or you do not exist in that conversation. There is no page two to fall back on.

What an AI citation audit actually checks

Most “AI SEO” advice stops at “write helpful content.” That is necessary and nowhere near sufficient. An AI assistant has to clear four hurdles before it can recommend you, and failing any one of them makes the others irrelevant:

  1. Access — can its crawler reach and render your pages?
  2. Identity — can it work out that you are a specific, real business?
  3. Extraction — can it pull a clean, quotable claim out of your content?
  4. Trust — does anything outside your own website corroborate you?

The checks are grouped by those four, plus a fifth group on measurement, because a change you cannot measure is a change you will abandon.

Group 1: Can AI crawlers reach you? (Checks 1–5)

The most common cause of invisibility in AI answers is the least interesting one. The crawler was blocked.

  1. Your robots.txt allows AI crawlers. Check for GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended and Bingbot. Visit yourdomain.com/robots.txt and read it. A single Disallow line aimed at one of these ends your citations from that platform completely.
  2. Your firewall is not blocking them silently. This catches people out far more often than robots.txt. Cloudflare and most managed WordPress hosts ship bot-fighting rules that block AI crawlers by default, and they do it without touching robots.txt. Check your firewall’s bot management settings directly.
  3. Your key pages render without JavaScript. Disable JavaScript in your browser and load your homepage and top service page. If the content disappears, most AI crawlers cannot read it either. React and heavily plugin-driven sites are the usual offenders.
  4. Your sitemap is generated, not static. Open yourdomain.com/sitemap.xml and look at the lastmod dates. If they are all identical and months old, your sitemap is a fossil, and new pages may never be discovered.
  5. You are not relying on Bing alone. This is where most 2025-era advice is now wrong. ChatGPT no longer simply passes queries to Bing — OpenAI has been building its own index, and pulls from multiple providers alongside it. Being in Bing still helps. It is no longer sufficient.

Group 2: Can they tell who you are? (Checks 6–11)

AI assistants recommend entities — identifiable organizations — not URLs. If a model cannot resolve your website into a specific business, it will not put your name in an answer.

  1. Organization schema is on every page, with legal name, logo, founding date, address and phone.
  2. Your sameAs array lists every profile you control — LinkedIn, Facebook, Instagram, X, YouTube, Crunchbase, Clutch. This is how a model connects scattered mentions into one entity.
  3. Your name, address and phone number are byte-identical everywhere. “Suite 200” on your site and “Ste 200” on Google Business Profile is enough to fragment your identity.
  4. Your Google Business Profile is claimed, complete and correctly categorized. For anything local, this is the single heaviest input into AI recommendations.
  5. You appear in the directories your industry’s AI answers actually cite. Ask ChatGPT for the best providers in your category and watch which sources it cites. Those are the directories to get listed on. For most B2B services that means Clutch and G2; for local services, Yelp and industry associations.
  6. Your entity is disambiguated. If another business shares your name, your schema, bio copy and directory listings need enough distinguishing detail — city, industry, founding year — that a model does not merge you with them.

Group 3: Is your content extractable? (Checks 12–18)

An AI assistant does not “read” your page. It lifts a claim out of it. Your job is to make that easy.

  1. Every important page answers its own question in the first 60 words. No throat-clearing, no “in today’s fast-paced digital landscape.” Models draw heavily from the opening.
  2. Your H2s are phrased as questions people actually type. “How much does commercial cleaning cost in Sacramento?” gets extracted. “Pricing Considerations” does not.
  3. Your key sentences survive being lifted out of context. Read any important sentence on its own. If it needs the paragraph above it to make sense, rewrite it. “Rates run $0.08–$0.15 per square foot” works alone; “as mentioned, this varies” does not.
  4. Every key claim contains a number, date or named entity. Specific claims get quoted. Vague ones get skipped.
  5. Comparative information is in a table or a numbered list. Structured blocks are extracted disproportionately often.
  6. The date appears in your body copy, not just the metadata. Write “as of September 2026” in the text. Models weight recency and read it from the visible content.
  7. FAQPage and Article schema are implemented and valid. Run every page through Google’s Rich Results Test. Broken schema is worse than none, because it signals carelessness on a page you are asking to be trusted.

Group 4: Is there any reason to trust you? (Checks 19–24)

This is the group nearly everyone skips, and it is the one that separates businesses that get named from businesses that get crawled.

  1. Your content has named human authors. Real names, real photos, real credentials, a LinkedIn link. A byline reading “admin” or your company name tells a model there is no identifiable expertise behind the page.
  2. Person schema marks up those authors, linked to the Organization.
  3. You cite authoritative sources and link out to them. Outbound citation to named studies and primary documentation is one of the clearer correlates of being cited yourself.
  4. You have review volume and recency. Three testimonials from 2023 will not carry an AI recommendation. Twenty-five current Google reviews will.
  5. Other websites mention you by name. Models weight what third parties say about you far above what you say about yourself. Press coverage, podcast appearances, expert quotes, association listings, guest posts. This is the highest-effort item on the list and the highest-yield.
  6. You publish something nobody else has. Original data, a survey, a benchmark, results from your own client work. Unique information is the most citable content that exists, because there is no alternative source to cite instead.

Group 5: Are you measuring any of this? (Checks 25–27)

  1. GA4 has a custom channel group for AI referrals, capturing chatgpt.com, perplexity.ai, gemini.google.com and Copilot. Without it, this traffic hides inside Direct and Referral, and the work looks like it did nothing.
  2. You keep a prompt log. Twenty questions a real buyer would ask, run monthly across ChatGPT, Gemini, Perplexity and Copilot, recording whether you appear and who appears instead. This is the cheap manual version of AI visibility tracking and it is enough to start.
  3. You track share of voice against named competitors. Absolute citations matter less than whether you are gaining on the three businesses you lose deals to.

How to run this audit in an afternoon

Work in this order. Each stage takes roughly an hour.

  1. Access (checks 1–5). Open your robots.txt, your firewall’s bot settings and your sitemap. Load your homepage with JavaScript disabled. These are binary — they pass or they do not.
  2. Baseline (check 26). Write down twenty questions a real customer would ask before hiring you. Run all twenty through ChatGPT, Gemini and Perplexity. Record whether you appear and who appears instead. This takes about forty minutes and it is the most useful forty minutes of the whole exercise, because it tells you which competitors the models currently prefer and what sources they lean on.
  3. Identity and structure (checks 6–18). Work down the list against your homepage and top three service pages.
  4. Trust (checks 19–24). Be honest here. Most businesses fail at least four of these six.

What to fix first

If you cannot do everything, the order matters more than the total.

Priority

Checks

Effort

Why first

1

1–5 — access

Hours

Nothing else can work while a crawler is blocked

2

19–20 — named authors

One day

Raises the ceiling on every page you have ever published

3

12–17 — extractability

Days

Rewriting openings and headings is cheap and compounds

4

6–11 — entity signals

Days

Schema and directory work, mostly one-time

5

22–24 — reviews, mentions, original data

Months

Slowest to build, hardest for competitors to copy

The pattern is worth noticing. The fast fixes are technical and anyone can do them, which means they are table stakes rather than an advantage. The durable advantage is in group 4 — reviews, third-party mentions and original data — precisely because those take months and cannot be bought quickly.

The honest part

None of this makes an AI assistant recommend a business that does not deserve recommending. These 27 checks remove the reasons a model cannot cite you. They do not manufacture a reason it should.

If you do good work, have real customers who say so, and know something your competitors do not, this audit is how you stop that being invisible. If you do not, fix that first — no amount of schema markup substitutes for it.

Find out where you actually stand

Most businesses have never checked whether AI assistants mention them at all. It is usually a short and uncomfortable conversation.

We will run the prompt test for you — twenty buyer-intent questions across ChatGPT, Gemini, Perplexity and Copilot — and send you a report showing whether your business appears, which competitors come up instead, and which of these 27 checks your site currently fails.

No charge, no obligation, and you keep the report either way.

Frequently Asked Questions

How do I get ChatGPT to recommend my business?

Make your site crawlable by GPTBot and OAI-SearchBot, add Organization schema so ChatGPT can identify you as a specific business, write content that answers questions directly in the first 60 words, and build third-party mentions and reviews. ChatGPT recommends businesses it can verify through sources other than their own website.

Why does ChatGPT recommend my competitors instead of me?

Usually one of three reasons: your site blocks AI crawlers, your business is not clearly identifiable as an entity through schema and consistent directory listings, or your competitors have more third-party mentions and reviews. It is rarely about which business is actually better.

Is GEO different from SEO?

They overlap but optimize for different outcomes. SEO aims for a ranking position that earns a click. Generative engine optimization aims to be quoted inside an AI-generated answer, which often produces no click at all. GEO weights entity clarity, extractable claims and third-party corroboration more heavily than traditional ranking factors.

How long does it take to appear in AI search results?

Technical fixes such as unblocking crawlers can show results within two to four weeks. Entity and schema work typically takes one to three months. Building the third-party mentions and reviews that drive consistent recommendations takes three to six months or longer.

Does llms.txt help me get cited by AI?

Evidence is currently thin. No major AI platform has confirmed it uses llms.txt as a ranking or citation input. It takes under an hour to add and does no harm, so it is worth doing — but it should come after crawler access, schema and content structure, not before.

How do I know if AI assistants are sending me traffic?

Create a custom channel group in GA4 capturing referrals from chatgpt.com, perplexity.ai, gemini.google.com and Copilot. Without it, this traffic is misattributed to Direct or Referral. Bear in mind that many AI citations produce a mention without a click, so referral traffic understates your true visibility.