How AI decides who to cite (and how to become the answer)
The field guide to AI search visibility for Australian businesses and the agencies that serve them: how the engines choose who to cite, the moves that actually work, and the plumbing nobody checks.

Your best customer just asked ChatGPT for a recommendation in your category. It named three businesses. You were not one of them.
They did not pay to be there. They did not rank first on Google. The model simply chose them as the answer, and the customer never saw a list of ten blue links to scroll. This is the new front door, and most businesses are still polishing the old one.
I run Outercite, where we track how AI search engines cite Australian businesses across ChatGPT, Claude, Gemini, Perplexity, Grok and DeepSeek. I have spent a lot of hours watching who gets named and who gets skipped, and the pattern is not random. It is learnable. This is the field guide I wish existed when I started.
The one shift that reorganises everything
Stop trying to rank a page. Start trying to be an extractable answer. Traditional SEO gets you ranked on a results page. AI search gets you cited inside an answer. They share a foundation, but they reward different things: SEO optimises for the click, AI search optimises for the reuse.
The unit is no longer the page. It is the passage. A page sitting on result two can be cited over the page on result one if its answer is cleaner and easier to lift. Three words people use interchangeably are not the same thing:
- Ranked. You sit high on a Google results page. A separate game.
- Mentioned. Your name appears somewhere in an AI answer. Nice, and passive.
- Cited. You are the source the answer was actually built on.
Key takeaway
Citation is the one that moves customers, and the one most AI visibility tools quietly fail to measure. Counting mentions is easy. Counting real citations is hard.
The secret that should reorder your strategy
You get cited far more often from other people’s websites than from your own. In one analysis of more than a million AI prompts, around 85% of citations came from earned, third-party sources, not the brands’ own websites. That is roughly five to six times more likely to be cited off your own site than on it, and we see the same pattern across the Australian businesses we track.
So if your whole AI search plan is to write more content on your own site, you are fighting for the smallest slice of the answer. The engines trust corroboration over self-description, so most of the real work happens off your domain, in what I call the citation supply chain:
- Reference and community sources. Wikipedia and Reddit turn up constantly.
- Review and comparison sites. The trusted directory for your category.
- Best-of and roundup pages. The third-party lists buyers compare against.
- YouTube. Frequently cited by Google’s AI Overviews for how-to and product questions.
Find where the engines that matter to you already pull their answers for your category, then earn an accurate, current presence in those exact places. Most businesses have never looked. That is the gap.
How engines actually choose who to cite
They read a page top to bottom, lean heavily on headings to understand intent, and prefer sources that answer directly, clearly, completely and credibly. Underneath a cited answer is a short pipeline: the engine reads the intent, retrieves candidate sources by searching live, selects the few it trusts most, and cites the ones it leaned on. The answers that carry citations are almost always retrieval answers, and retrieval is the part you can actually influence.
What actually works: three levers
1. Structure: be extractable
Make every section a self-contained answer block that works if it is lifted out with no surrounding context.
- Open each section with the answer, in about 40 to 60 words, then expand.
- One idea per section, one idea per paragraph. Keep paragraphs to two to four sentences.
- Write your headings the way a person phrases the question.
- Use lists for options and tables for comparisons.
- Add an FAQ with self-contained answers.
2. Authority: be citable
This is measured, not a matter of taste. The Princeton GEO study found the right moves can lift a source’s visibility in AI answers by up to 40%, with the biggest gains from citing your sources, adding specific statistics, and adding expert quotations. Keyword stuffing, the old SEO reflex, was one of the weakest tactics it tested. Freshness is weighted heavily, so show a visible last-updated date, and name your authors with real credentials.
3. Presence: be where AI looks
This is the third-party supply chain from earlier, and it is the lever with the most headroom precisely because it is the most work. Map the sources, earn the placements, keep them current. If you do only one new thing this quarter, do this.
The plumbing nobody checks
Four technical moves decide whether the engines can read you at all. They take an afternoon and almost nobody does them.
- AI bot access. If GPTBot, PerplexityBot, ClaudeBot or Google-Extended are blocked in your robots.txt, that engine literally cannot cite you. Check this first.
- Schema markup. Structured data tells engines what your content is without guessing.
- An llms.txt file. A clean map of what you do for AI systems.
- A pricing.md file. AI agents are becoming the ones comparing products for a buyer. If your pricing is behind JavaScript or a contact-sales wall, they skip you.
How do you know if it is working
You measure at the passage level, repeatedly, across engines, because a single check is noise. The trap almost everyone falls into is to run one prompt, screenshot the answer, and treat it as proof. But the same prompt returns different answers run to run and engine to engine.
Key takeaway
Proof is a pattern measured over time, not a screenshot. And you have to separate a real citation, where the engine used you as a source, from a hallucinated mention. At Outercite we verify every citation with two independent models before it counts.
The metrics that actually reflect AI visibility are citation frequency, share of voice against your competitors, and query coverage across the related questions buyers ask.
Turn it into a system: See, Understand, Act, Prove
AI search visibility is not a project you finish. It is a loop you run.
- See. Where you stand today: coverage, citation rate, share of voice.
- Understand. Why a competitor is cited and you are not: entities, authority, third-party trust, freshness.
- Act. The highest-leverage moves first, sequenced by impact.
- Prove. Track each change, tie it to citation movement, and repeat.
Run it every quarter and it compounds. That is the difference between a flurry of activity that moves nothing, and a system that keeps making you the answer.
If you are an agency, this is a service
There are no established AI search agencies yet, and no one has a ten-year head start. The agency that walks into a client meeting with a clear method and a plain-English brief wins the retainer. Package the loop into a productised service: an audit that shows the client where they stand, a one-page brief that turns the gaps into a scoped, priced ask, and a monthly report that proves it worked.
The honest part
This takes about 90 days to compound, especially the off-site work. Leading signals like citation rate move first, with traffic and revenue following. Anyone promising instant AI visibility is selling you something. The good news is that the field is young, the moves are knowable, and most of your competitors are not doing them yet.
You do not need any tool to start. Open ChatGPT, ask it the question your best customer would ask, and see whether it names you. That answer is your starting line.
Sources
- The 40% visibility lift and the top tactics: Aggarwal et al., GEO: Generative Engine Optimization, KDD 2024. arXiv.
- The earned-media citation share: Muck Rack, analysis of over one million AI prompts. Study.
- AI crawler controls: official docs for OpenAI GPTBot, PerplexityBot, Anthropic ClaudeBot and Google-Extended.

