OuterciteOutercite
Core Concepts

Confidence scores explained

The confidence score tells you how certain Outercite is that a citation is genuine, based on agreement between two independent AI models.

Every citation Outercite records comes with a confidence score. That score tells you how certain the system is that your business was genuinely mentioned in an AI answer. The higher the score, the more clearly both verification models agreed on the result.

What you'll learn

  • How the two-model consensus process produces the confidence score
  • What high, medium, and low scores mean in practice
  • How near-zero false positives are achieved and why that matters

Where the confidence score comes from

Outercite does not count a citation just because a keyword appeared somewhere in an AI response. It runs every result through a two-model verification pipeline.

Here is how that pipeline works:

  1. Intent classification. The prompt is classified by intent type so the system knows what kind of answer to expect. See intent classification for detail.
  2. Deep analysis. A high-performance model reads the full AI response and identifies every business it believes is mentioned, named, or linked.
  3. Cross-verification. A second model independently audits what the first model found. It does not see the first model's output. It reads the same response and forms its own conclusion.
  4. Consensus result. The system compares both models' findings. Only citations both models independently agree on are counted and assigned a confidence score.

The confidence score is the output of that consensus step. A perfect score means both models saw the citation clearly and agreed without ambiguity. A lower score means one model was less certain, or the wording in the original AI answer was ambiguous.

Outercite verifies every citation with two independent models. The two-model approach keeps false positives near zero across all six tracked engines.

Reading the score: high, medium, and low

The confidence score runs from 0 to 100. A higher number means the two models agreed more strongly that the citation is genuine. The bands below are a practical rule of thumb for how to treat each citation, not fixed settings in the product.

High confidence (roughly 85 and above). Both models identified the citation clearly. The AI response named or linked your business in an unambiguous way. These citations are reliable signals you can act on. They are safe to include in reports and trend analysis.

Medium confidence (roughly 60 to 84). Both models agreed a citation was present, but one model was less certain. This can happen when the AI answer used indirect phrasing ("businesses like yours in that area") or when your brand name is similar to another. Medium-confidence citations are real but worth reviewing if you are doing detailed competitor analysis or building a content argument.

Low confidence (below 60). The models reached some agreement but with notable uncertainty. This often reflects genuinely ambiguous language in the AI response. Outercite records these but flags them clearly. Treat low-confidence citations as signals to investigate rather than confirmed data points.

Avoid building strategy on low-confidence citations alone. They are early indicators, not confirmed facts. Use them to identify areas worth monitoring more closely.

Why two models instead of one

A single model can be confidently wrong. Language models sometimes hallucinate, misread context, or overfit to certain phrasing patterns. By using two independent models with different architectures as reviewers, Outercite introduces a check that neither model can override on its own.

Think of it like a two-person audit. One auditor works through the accounts independently. The second auditor does the same. You only sign off on findings both auditors reached. Disagreements go back for review rather than being counted.

This matters because the cost of a false positive is real. If your citation count includes phantom mentions, you might invest in a content strategy based on a position you never actually held. Clean data leads to better decisions.

Why false positives near zero matters

The phrase "near zero" is specific. It does not mean zero. Ambiguous AI responses exist. Brand names that overlap with common words exist. The system handles edge cases by flagging uncertainty rather than hiding it.

What "near zero" means in practice: the citations you see in your dashboard are ones the system is confident are real. The confidence score tells you exactly how confident. You are not working with a black box. You can see the score, decide your own threshold for action, and filter your data accordingly.

How to use confidence scores in your workflow

  • For weekly reporting: as a rough guide, focus on citations with a high confidence score (roughly in the upper band). These are your cleanest, most defensible numbers.
  • For trend analysis: include medium-confidence citations. The trend line is still valid even if individual data points have some uncertainty.
  • For competitor benchmarking: focus on high-confidence citations for both your brand and competitors. This keeps the comparison fair.
  • For alerts and monitoring: Outercite surfaces low-confidence results in surge alerts so you can investigate spikes before they reach your main reports. See set up alerts for configuration options.

Try this in Outercite

Open your citations list and sort by confidence score. Identify any medium or low-confidence entries for your most important keywords. Click into each one to see the original AI response and decide whether the citation reflects how you want your brand to appear.

Was this helpful?