For practitioners

GEO audit template: what a client-ready AI visibility audit contains

Clients are paying $500 to $2,000 for a one-off AI visibility audit in 2026. Here is the complete structure of one that earns that fee: every section, what goes in it, and the honesty rules that keep the client at the second invoice.

A client-ready GEO audit contains seven parts: a plain-language summary of how AI engines currently see the site, an executive summary, evidenced live visibility tests on real engines, an eight-dimension diagnostic with a transparent score, copy-paste fixes the owner can apply, a developer task sheet for the technical items, and a prioritised 90-day plan that ends in a re-test.

That is the whole answer. The rest of this post walks through each part, so you can build the template yourself or sanity-check the one you already use. The market band above is the verified 2026 rate for exactly this shape of work (LLM Pulse, GEO Agency Guide).

Part 1: How AI sees this website

Open the report with one page in plain language. Not a score, not a checklist. Two or three paragraphs describing what an AI engine can and cannot work out from the client's site today: what the business does, where it operates, who it serves, and what evidence exists for any of it. The client will read this page even if they read nothing else, and it frames every finding that follows.

Part 2: Executive summary

The score, the three biggest gaps, and the one-sentence plan. Write it last. If a finding does not survive compression into this page, it probably does not belong in the report at all.

Part 3: Live visibility tests, evidenced

This section is what separates a practitioner audit from a $99 automated scan. Run real queries on the engines the client's customers actually use, and record what came back.

  • A 3 by 3 grid. Three query types: one local ("best emergency plumber in Stockport"), one commercial ("boiler replacement cost Stockport"), one national stretch. Each run on three engines, for example Perplexity, Google AI Mode and ChatGPT.
  • One run per engine, date-stamped, screenshotted. AI answers are volatile, so a run you cannot evidence is a run you cannot defend. The date-stamp is not bureaucracy; it is what makes the 90-day re-test meaningful.
  • Log the cited domains, not just the named businesses. While each answer is on screen, record which domains the engine cited. Semrush's study of over 100 million citations across 230,000+ prompts found AI answers are assembled from a recurring set of third-party sources, and each engine forms its own source hierarchy (Semrush, 2025). Every cited domain the client is absent from is a finding, and usually a sellable action.

Disclose the volatility to the client in this section. An independent study of 530,875 AI citations found most cited sources change day to day, with daily churn between 44 and 88% depending on the engine (GetMentions.ai, June 2026). Telling the client what nobody can honestly promise is exactly what separates you from the vendors who promise it anyway.

How to build the three queries

The grid only works if the queries are ones a real customer would type, so construct them from the client's intake call, not from a keyword tool. The local query pairs the service with the place, phrased as a person asks it. The commercial query carries buying intent: a cost, a comparison, a "which should I choose". The national stretch query is the one the client wishes they ranked for and probably does not; it earns its place because the answer shows who the engines treat as the authorities in the niche, which feeds directly into the citation-source analysis below. Write the queries down before you run anything, and never rerun a query until you get an answer you prefer. One run, whatever it says, is the rule that keeps the audit honest.

Part 4: The eight-dimension diagnostic

The live tests show the symptom. The diagnostic explains it. Eight dimensions cover the ground without padding:

  1. Clarity. Can an AI engine work out what the business does from the homepage alone? One H1, a first paragraph that states the what, the where and the who in plain words.
  2. Quotability. Are there passages an engine can lift whole? Question-shaped headings with direct answers underneath. A page of badges and slogans gives the model nothing to quote.
  3. Entity coverage. Does structured data (JSON-LD, the machine-readable block that names the business, its type and its offers) match what the pages say?
  4. Trust. Reviews, named people, a verifiable address, third-party mentions. The cross-referencing an engine does before it names anyone.
  5. Extraction structure. Headings in a logical hierarchy, lists where lists belong, tables for anything tabular.
  6. Semantic density. How much verifiable, specific information each page carries per hundred words. Filler copy dilutes the signal.
  7. Information gain. Does the site say anything the model cannot already get from ten other sites? First-hand knowledge, real numbers, real photographs.
  8. Technical basics. Crawlability for AI user agents, render speed, canonical hygiene, a working sitemap.

Mark each check pass, fail or not applicable, and note the evidence. On llms.txt specifically: record it, do not score it. Ahrefs analysed 137,210 domains and found 97% of published llms.txt files were never requested by any crawler (Ahrefs, May 2026). Scoring it inflates the checklist without evidence, and clients eventually find out.

Part 5: The score, with the weighting shown

A single AI visibility score is useful shorthand, but only if the client can see how it is computed. Publish the weights next to the dimensions in the report. A score with hidden weights is marketing; a score with visible weights is a measurement the re-test can repeat.

Part 6: Fixes and developer tasks, separated

Split the actions by who does them. Copy-paste fixes the owner can apply the same afternoon: a rewritten H1, a meta description, an FAQ block. Developer tasks for the rest: schema installation, heading restructure, feed work. Every fix traces back to a failed check, and every failed check feeds a fix. Nothing floats.

Part 7: The 90-day plan and the re-test

Prioritise by impact against effort: this week, this month, this quarter. Then book the re-test. Same queries, same engines, same weights, 90 days later. Because the first run was date-stamped and screenshotted, the re-test produces movement you can point at, and the re-audit is the natural recurring engagement.

The built version of this template

Everything above ships pre-built in the AI Visibility Audit Kit, £39: the white-label client report template (Markdown and styled HTML), the 8-tab scoring workbook with the 40-check runsheet and disclosed weights, the live-test grid, the citation-source tracker, the prompt pack and the proposal template with verified 2026 pricing benchmarks. Unlimited white-label client use.

No visibility outcomes are guaranteed, AI engine results are volatile; this is an evidence-based diagnostic methodology.

The four mistakes that sink first audits

  • Scoring without evidence. Every pass or fail needs a line of evidence next to it: the URL checked, the passage quoted, the screenshot reference. A checklist without evidence collapses under one hard client question.
  • Auditing only the homepage. Run the diagnostic on the homepage and at least one money page. The homepage tells the engine who the business is; the money page is where a commercial query should land. They fail differently.
  • Reporting findings without owners. A finding that is not assigned to either the owner's fix list or the developer's task sheet will not get done, and the re-test will show no movement you can take credit for.
  • Promising what the volatility data forbids. If any sentence in the report guarantees a citation or a ranking, the 530,875-citation study above is the evidence a sharp client will quote back at you.

The rule that holds the whole template together

Never promise rankings, citations or placements, anywhere in the report. The volatility numbers above are the reason. Your product is a rigorous diagnostic of the durable inputs, honestly measured, with a re-test to prove movement. In a market full of "guaranteed AI citations" sellers, the evidence-based audit is the one clients keep paying for.