How the ladder is built.
No marketing fluff. Here is exactly what we measure, how we score it, and where our estimates come from.
A weekly basket of real buying questions.
Each category has a hand-curated basket of dozens of buying-intent prompts, the kind a real customer would type or ask. Every prompt is put to six AI engines: ChatGPT, Claude, Gemini, Perplexity, Grok and DeepSeek. Each engine gets one answer per prompt, once per week.
Judged, then verified, then weighted.
A judge model extracts every brand mention from each raw answer. A second pass deterministically verifies each mention against the actual answer text, so nothing is scored on a hallucinated citation. Position matters: a 1st-place recommendation gets full weight, each rank after that decays by 0.15, down to a floor of 0.4. An unranked recommendation (mentioned favourably but not placed) scores 0.7. A bare mention with no recommendation scores 0.3. A brand's per-engine score is its share of the basket. The blended score shown on the ladder is the mean across every engine that returned usable data that week.
An estimate, always shown as a range.
Where we have the inputs, we estimate what an absent brand may be losing: monthly query volume multiplied by an AI adoption share, multiplied by the recommendation gap share, multiplied by average customer value, multiplied by a 10% capture assumption. The adoption share is 0.18, an assumption derived from published consumer research on AI search adoption, not a figure we measure ourselves, and we review it as the market shifts. We always show the result as a range (0.5x to 1.5x the central estimate), and we omit the panel entirely when the underlying inputs are missing. These are estimates, clearly marked as such. They are never presented as measured revenue.
What this ladder is, and is not.
This ladder measures AI visibility, not business quality. A low score does not mean a bad business, it means AI engines are not recommending it right now. Every mention that counts toward a score is verified with a two-source check before it is scored. When an engine fails to return usable data for a category in a given week, we disclose it on that category's page rather than silently filling the gap. If your brand wants a correction, or wants its data reviewed, contact us.