OuterciteOutercite

How verification works

What happens between an AI engine's answer and a counted citation: the prefilter, the judge, the verifier that audits it, and the rules that decide a tie.

Outercite does not ask an engine whether it mentioned you and write down the reply. Each answer is read, judged, and where the judgement is not clear cut, audited by a second model before anything is counted.

What you'll learn

  • The four things that happen to every answer, in order
  • What the confidence score on a citation actually measures
  • The rules that decide the cases people argue about

Stage 1: The prompt is classified

Before an engine is asked anything, the prompt is sorted into one of five intent types: local, buying, informational, comparison or branded. A citation on a buying-intent prompt is a different commercial event from a passing mention in an explainer, and the classification is what lets the platform report them separately. See intent classification.

Stage 2: The answer is prefiltered

If an answer contains no concrete signal of your business at all, no name variant, no website, no phone number, there is nothing for a model to weigh. That check is recorded as not cited without spending a model call on it. It is cheap, and it is honest: the absence is the measurement.

Stage 3: A judge reads the answer

Every remaining answer goes to a judge model, which reads the whole response rather than searching for a string. It returns a structured verdict:

  • whether your business is explicitly cited, and how confident it is on a 0 to 100 scale
  • which kinds of citation it found: brand, product, service, URL or phone number
  • the exact sentence that cites you, quoted verbatim
  • where you sat in a ranked list, if the answer made one
  • the sentiment of the mention
  • which competitors were named, and where
  • which sources the answer leaned on

Two judging rules matter more than people expect:

Being named is not enough. If the answer names you in order to warn readers off, or steers them to a competitor instead of you, it is not a citation. It is recorded as not cited with negative sentiment, and it shows up under negative mentions rather than quietly inflating your rate.

A generic term that matches your name is not you. If the business is "GRP Tubing Pty Ltd" and the answer discusses GRP tubing the material, that is not a citation of the company. Matches that survive only because normalisation stripped the legal suffix are rejected.

Stage 4: A verifier audits the uncertain cases

Where the judge is less than certain, its findings go to a second, independent model, which re-reads the original answer and reaches its own verdict. What happens next is the part worth understanding, because it is what your confidence score means:

What happenedHow it is recordedCounted as cited?
Both models agreetwo_model_agree, confidence averaged across the twoYes, if both said cited
The models disagreetwo_model_dispute, confidence is the lower of the twoNo
The judge was certain, so no verifier ranjudge_only, the judge's own confidence, unchangedYes, if the judge said cited
No signal to judgesignal_prefilterNo

A disagreement fails closed. When the two models do not agree, the citation is not counted, because a platform whose numbers you take to a board meeting should be wrong in the direction of understating you.

A check that the judge decided alone is labelled as exactly that. It does not borrow the confidence of an agreement that never happened, which is the single most important honesty rule in the pipeline: one model's opinion is one model's opinion.

What the confidence score is, and is not

The confidence score is how sure the models are that the citation is real and substantive. It is not a quality score for the citation, not a ranking, and not a measure of how prominent you were in the answer.

  • High. Clear, substantive, unambiguous.
  • Middle. Genuine, with something ambiguous about the context or the prominence.
  • Low. Present but weak. Worth reading the sentence yourself, which you can do in Proof.

Not every citation carries every field. Position only exists when the answer ranked things. Sentiment and descriptors only exist when the answer expressed them. An empty field means the answer did not contain one, not that the number is zero.

AI Overviews is scored differently

Six of the seven surfaces are asked a question. Google's AI Overviews is read off the results page, and Google does not draw a panel for every query. When there is no panel, the check is recorded as an absent overview and left out of your citation rate entirely, rather than counted as a failure to be cited. Nobody failed to cite you in an answer that was never written.

A raw search counts appearances. Verification counts recommendations. The difference is deliberate, and it is the whole reason the number is usable. See why Outercite shows fewer citations.

Try this in Outercite

Open the assistant and ask "Show me the exact sentences AI wrote about us." The Proof panel gives you the verbatim sentences, the engine, and the date, so you can check the judgement yourself.

Was this helpful?

Last updated on