How the Pipeline Works
A plain-English walkthrough of the four-stage verification pipeline that turns raw AI responses into trustworthy citation data.
Outercite does not simply ask an AI engine whether it mentioned your brand and take the answer at face value. Every response goes through a four-stage verification pipeline designed to catch errors, filter noise, and produce citation data you can trust.
What you'll learn
- The four stages of the verification pipeline in plain English
- Why two independent models produce more reliable data than one
- What a confidence score represents and how it is calculated
The problem with naive citation tracking
If you asked an AI engine "Did you mention my business?" it might say yes when the mention was vague, tangential, or even hallucinated. A simpler tracking tool might count every response that contains your brand name, regardless of context. That produces inflated, misleading numbers.
Outercite solves this with a multi-model pipeline. Each AI response is analysed twice, by two different models, and only citations that both models agree on are counted. This keeps false positives near zero.
Outercite checks every prompt across all six major AI engines and verifies each citation with two independent models. Across the prompts we track, brands are cited in about half of all checks (a 48% citation rate). Every counted citation passed through all four stages below.
Stage 1: Intent classification
Before any AI engine is queried, each prompt is classified into one of five intent categories.
| Intent type | Example prompt |
|---|---|
| Local | "Best physiotherapist in Geelong" |
| Buying intent | "Which project management tool should I buy?" |
| Informational | "How does cloud accounting work?" |
| Comparison | "ChatGPT vs Perplexity for research" |
| Branded | "What do people say about Outercite?" |
Intent classification matters because citation value varies by intent. A citation in a buying-intent answer is commercially very different from a passing mention in an informational one. Classification is how Outercite groups your results meaningfully and powers features like share of voice and competitor analysis.
The five categories above are illustrative of how intent is classified. The exact taxonomy may evolve as the pipeline is refined.
You can explore intent classification in more depth in the Intent Classification article, which also explains how to filter your reports by intent type.
Stage 2: Deep analysis
Once the AI engine returns a response, that full response is passed to a high-performance language model for deep analysis.
This model does more than check for a keyword match. It reads the complete response and asks:
- Is any business named in this response?
- Is the mention substantive (a recommendation, a description, a comparison)?
- Is the brand being cited in a positive, neutral, or negative context?
- What position does the citation hold in the answer?
The output of this stage is a structured list of every business mentioned in the response, along with metadata about each mention.
A simple keyword search would catch any mention of your brand name, including irrelevant context. The deep analysis stage filters for meaningful citations. This is why your Outercite citation rate may be lower than a simple text-search would suggest, and why it is more useful.
Stage 3: Cross-verification
The structured output from the deep-analysis model is passed to a second, independent model for cross-verification.
The cross-verification model re-reads the original AI response and audits the first model's findings. It is not told what the first model concluded. It makes its own independent judgement about which citations are present and meaningful. Two models with different architectures reaching the same conclusion is a strong signal that the conclusion is correct.
Stage 4: Consensus result and confidence score
The pipeline compares the outputs from both models. Only citations that both models identify are counted as verified citations.
Where the models agree completely, the confidence score is high. Where there is partial agreement (for example, both see a mention but differ on the intent context), the confidence score reflects that uncertainty.
The confidence score you see on each citation in your Outercite dashboard is a direct output of this consensus step. A high confidence score means both models saw a clear, substantive citation. A lower score means the citation exists but had some ambiguity in context or prominence.
Prompt → AI engine response → Intent classification
→ Deep-analysis model
→ Cross-verification model
→ Consensus check: both agree = verified citation + confidence score
disagreement = not countedWhy two models matter
A single model can be wrong for reasons that are hard to predict: a quirk of its training data, a formatting pattern it over-weights, or a simple error. Two independent models sharing the same error is far less likely.
This is why Outercite can claim near-zero false positives. It is also why the verified citation rate sits at roughly 48%: citations require agreement from both models, so only genuine mentions clear the bar. Checks that do not produce a citation are equally valuable. They are accurate data points showing where your brand was absent.
When Outercite says you were cited, you can trust it. A dashboard full of uncertain mentions is not useful for decisions. Near-zero false positives is the design goal, not a side effect.
How this connects to your dashboard
Every metric in your dashboard flows from this pipeline. Citation rate, visibility score, share of voice, and competitor threat scores are all calculated from verified, consensus-confirmed citations. Predictions and forecasts are built on historical trends from this clean data.
If you want to understand what any of these metrics mean, the confidence scores explained article is a good next step.
Try this in Outercite
See the pipeline in action on your own brand. Start the onboarding walkthrough and Outercite will run your first keyword through all six engines, classify the intent, run both verification models, and show you the confidence score on every citation.
Related
Confidence Scores Explained
A deeper look at what confidence scores mean, how to interpret them, and when a lower score still indicates a genuine citation.
Intent Classification
How the pipeline classifies each prompt by intent and why it affects how you interpret your citation data.
What is Outercite?
The big-picture intro: six engines, who Outercite is for, and the full list of metrics in the dashboard.
