How it works

What actually happens when you run an audit

When someone asks an AI search engine a question, the engine answers using pages it can read and quote. If your page does not answer that question plainly, it quotes somebody else's. An audit tells you which questions you are losing, and why.

1

We fetch the page

You paste a URL and we download that one page — the same way a crawler would. We strip out navigation, headers, footers, cookie banners and scripts, and keep the main content, because that is the part an engine would actually quote. Nothing is installed and no account is needed.

We read one page at a time, not your whole site. Audit the page you care about ranking.

2

We work out what the page is about

Two questions, asked separately: what kind of site is this, and what is this particular page about. Then they are combined — “HVAC installation company”, “heat pump technology on an online encyclopedia”.

Separately, because asking once gets it wrong. Audit the Wikipedia article on heat pumps with a single “what is the business type” question and you get “online encyclopedia” — followed by 25 questions about editing Wikipedia. The site and the page are two different facts, so we detect them as two.

Everything downstream depends on this, so it is worth a glance when your report loads: if the topic is wrong, the questions will be too.
3

We generate 25 questions people ask AI about that topic

This is the question we are asked most, so here is the whole mechanism.

We ask for the 25 questions people most commonly put to AI assistants about your topic, spread deliberately across intents: what it is and how it works, cost, choosing between options, how-to, troubleshooting, and trust or safety. They are written the way someone types before they have chosen a supplier.

The generator never sees your page. Its inputs are the topic, your hostname and the language — nothing about what you happen to have written. That is the point: these are the questions your market asks, not a summary of your page handed back to you. A test written from your own content would be one you always pass.

The set is generated once per URL and then frozen. Re-audit the same page and you are measured against the same 25 questions, so a change in your score is a change in your page.

We learned that the hard way. Before freezing, two audits of the same unchanged page three hours apart scored 54 and 28 on coverage — every technical signal byte-identical. That is 30% of the composite moving on its own, and it made the history graph worse than useless: you could change nothing and read it as progress.

We detect the language your page is written in and write the questions in that language, phrased the way a native speaker would ask them — not translated from English.

4

We check which ones your page answers

Each question is graded against your actual page content — not by keyword matching, but by reading it. Every question lands in one of three buckets:
Covered
The page states an answer the reader can act on without going anywhere else.
Partial
The page raises the topic but stops short — the reader would still have to infer it, search further, or contact you.
Missing
The page does not address it, or mentions it only in passing.

Covered counts once, Partial counts half, Missing counts nothing. That ratio is your Question Score, and it is the single largest part of your overall score at 30%.

Partial is usually the cheapest thing to fix: the subject matter is already on the page and it needs a heading and a direct answer. If a verdict comes back unusable for any reason it counts as Missing — we round against you, not in your favour.

5

We run 6 technical checks on the page

The questions measure whether your content answers what people ask. These six measure whether an engine can find, parse and trust the page in the first place. Unlike the question score, every one of these is a verifiable fact about the page — we read your robots.txt, parse your JSON-LD and look at your markup.
AI crawler access20%

Whether your robots.txt actually lets GPTBot, ClaudeBot, PerplexityBot and Google-Extended read the page. Block them and nothing else on this list matters — the page cannot be cited if it cannot be read.

Structured data15%

JSON-LD markup that tells an engine what the page is: an article, a product, an organisation, a set of questions and answers. Without it the engine has to infer everything from prose.

Answer structure15%

Whether the page is laid out as questions and direct answers. A heading that asks a question, followed by a paragraph that answers it, is far easier to quote than the same fact buried mid-paragraph.

Citations and statistics10%

Concrete numbers and links to outside sources. Pages that back claims with data get cited more than pages that assert them.

Freshness5%

A machine-readable publish or update date. Static pages like an About or Pricing page rarely carry one and that is fine — it matters most for guides, blog posts and anything where recency changes the answer.

E-E-A-T signals5%

Evidence of who wrote this and who stands behind it: an author byline, a link to an author or about page, an Organization schema.

If the crawler-access check cannot complete — your robots.txt returns 403, or the request times out — you get unable to verifyand the reason, and that one signal is left out of the weighted average rather than filled in with a guess. It is the only signal that can end up unverified, because it is the only one we fetch separately from your page. A tool that scores an unreachable robots.txt as “no blocks found” is handing you a green tick that means “we failed to look”.

Five of these six are pure functions over the HTML we fetched: the same page produces the same score every time, and we test that. Question Coverage is a model judgement, and the same page can score a point or two differently on two consecutive runs — treat a small move on an unchanged page as noise rather than a result. The methodology has the measurements. (Crawler access can move too, but only because your server's answer moved.)

If GPTBot, ClaudeBot or PerplexityBot are outright blocked, your overall score is capped at 40 however good the rest of the page is. A page the engines cannot fetch cannot be cited, and a 78 next to a blocked crawler would be a comfortable lie.

llms.txt is checked too, but deliberately kept out of the score and shown as a bonus. No AI provider has confirmed it affects citation ranking, so scoring you on it would be inventing a standard that does not exist yet.

6

We ask the engines directly

Everything above reads your page. This step asks the engines. We take five of your 25 questions, put them to Perplexity — and to OpenAI as well on paid plans — and record which answers used your domain as a source. The result lands on the same table as the coverage verdict, a chip per engine on each question, so “my page answers this” and “an engine cited me for it” can finally be read on one line.

Covered questions are checked first. A question your page answers that an engine still did not cite is the one result that changes what you should do — every other gap says write the answer, and that one says the answer is already there. Most pages have few covered questions, so the sample usually continues into Partial and Missing.

One check, five questions, one run each. It is a snapshot. Ask the same engine the same question tomorrow and you may get a different answer. Anyone reporting this as a stable percentage is reporting noise as a metric, and 0 out of 5 does not mean “never”.

Being named in an answer is recorded separately and does not count as a citation. Getting mentioned is not the same as being the source the engine relied on, and merging the two would flatter everyone.

When an engine answers without using you, we show which domains it used instead — and across the sample, the domains cited most often, with a link to put your page beside any of them. Five questions is a pointer, not a ranking of your market, and the report labels it that way.

Where our verdict and the engine disagree, the question gets a badge. Answered, not cited means your page answers it and the engine still chose someone else — a trust or format problem rather than a content one, so writing more is the wrong fix.

The other twenty questions show “not checked yet” — not a bad result, just an absent one — and paid plans can check any of them on demand, against both engines, from the row itself. That comes out of a monthly allowance shown next to the button, because each click costs real money and we would rather ration that visibly than bury it in the price. What you get back is another snapshot, not a rate: engines vary run to run, and one more observation is one more observation.

7

You get fixes, not just a number

Every gap comes back with something you can paste: the answer text for a missing question, an FAQ block, or the JSON-LD snippet for a schema you are lacking — written in your page's own language. Each fix shows the effort involved and what it is worth, so you can start with the cheap ones — and the labels are not all “high impact”.

Generated content is checked against your page before you see it. Any date, price or percentage your page does not state is replaced with a placeholder like [INSERT PRICE], plus a line explaining why. Claims without a figure in them cannot be checked automatically, and we say so at the point you are about to publish rather than pretending the check is total.

Projected scores after fixes are deliberately conservative — 40% of the theoretical maximum — because implementation quality varies, and a projection you beat is better than one you resent.

What the score means

The overall score is the seven signals — question coverage plus the six technical checks — weighted and combined into one number out of 100.

71–100

Well optimised

An engine can read the page and it answers most of what people ask.

41–70

Needs GEO optimisation

The page is readable but leaves real questions unanswered — usually the quickest wins.

0–40

Needs significant work

Either engines are blocked outright, or the content answers very little of what is asked.

One thing worth being clear about: this measures how well your page is set up to be cited. It is not a measurement of whether an engine cited you today — that changes constantly and no tool can promise it. What we can tell you is whether the page gives an engine a reason to.

Want the technical detail?

The methodology page has the exact formula behind every signal, the thresholds, and what we deliberately do not measure.

Read the methodologySee an example reportAudit a page