How the leaderboard is built.
No marketing fluff. Here is exactly what we measure, how we score it, and where our estimates come from.
A fortnightly basket of real buying questions.
Each category has a hand-curated basket of 12 buying-intent prompts, the kind a real customer would type or ask. Every prompt is put to six AI engines: ChatGPT, Claude, Gemini, Perplexity, Grok and DeepSeek. Each engine gets one answer per prompt, once per round, so a category is scored on up to 72 answers a round. Every category on the board is measured in the same round, and rounds run a fortnight apart.
Judged, then verified, then weighted.
A judge model reports which tracked brands each raw answer mentions. A mention only counts when the brand's name, or one of the aliases we list for it, appears word for word in the answer text. There is no fuzzy matching, so a near miss such as “Australians” never counts as AustralianSuper. Short acronyms, brand names that are also everyday words (ART, ANZ, Rest, Boost, and longer names that start with one, such as Boost Mobile) and descriptive aliases such as Sydney Uni must appear in the brand's own capitalisation. Other names match in any capitalisation and with a possessive, so Commbank and PayPal's Pay in 4 both count. Phrases that only look like a brand, such as Western Sydney University or “the rest of”, are ignored. Position matters: a 1st-place recommendation gets full weight, each position after that decays by 0.15, down to a floor of 0.4. An unranked recommendation (mentioned favourably but not placed) scores 0.7. A bare mention with no recommendation scores 0.3. A brand's per-engine score is its share of the full basket, so a missing answer counts as zero rather than shrinking the basket. The blended score shown on the ladder is the mean across the engines that were scored that round.
Coverage rule: questions that get no usable answer are asked again once. An engine that still answers fewer than 10 of the 12 questions is left out of that round's blended score and is not shown as its own ranking, and the category page says how many questions it answered. If fewer than three engines clear that bar, no result is recorded for the round. Ties: brands on the same score share the same rank (shown as =1), so neither is placed ahead of the other, and brands tied for first both count as leading that round.
Movement: rank changes, streaks and “The twist” compare a round with the last published round. The twist names the brand that climbed the most, and only when that brand was on the previous board and climbed at least 3 points. A smaller change is within ordinary round-to-round variation, so no brand is named and a general note is shown instead. Rounds published before the coverage rule are corrected the same way: an engine under the bar is taken out of that round's blended score, which is recalculated from the per-engine results we stored, and the movement shown on the following round is recalculated to match.
Sentiment themes: the judge also notes, in a few words, what each answer praises or criticises a brand for. We tidy those notes before they are shown. A theme that is too long, or that quotes a price, fee or year, is dropped rather than cut short. A theme that several brands share in the same round is dropped, because it describes the whole list rather than one brand, and so is a theme that appears as both praise and criticism for the same brand. Near-identical themes are merged, and names and spelling are shown in Australian English. The quote under “The receipts” is a short extract from one real answer, cut at a sentence or word boundary. We skip answers that date their own knowledge to an earlier year (“as of 2023”), because they can name services that no longer operate.
An estimate, always shown as a range.
Where we have the inputs, we estimate what an absent brand may be losing: monthly demand multiplied by an AI adoption share, multiplied by the recommendation gap share, multiplied by average customer value, multiplied by a 10% capture assumption. The adoption share is 0.18, an assumption derived from published consumer research on AI search adoption, not a figure we measure ourselves, and we review it as the market shifts. We always show the result as a range (0.5x to 1.5x the central estimate), and we omit the panel entirely when the underlying inputs are missing. These are estimates, clearly marked as such. They are never presented as measured revenue.
What this ladder is, and is not.
This leaderboard measures AI visibility, not business quality. A low score does not mean a bad business; it means AI engines are not recommending it right now. A mention only counts toward a score when the AI judge reports it and the brand name or a listed alias is then found word for word in the actual text of the answer. When an engine fails to return usable data for a category in a given round, or answers too few of the questions to be scored, we disclose it on that category's page rather than silently filling the gap. If your brand wants a correction, or wants its data reviewed, contact us.