How the AI Index is built.
The exact method behind every ranking: the questions we ask, the engines we ask them of, how the score is computed, and the rules that keep it honest.
Methodology v1.3 · current as of September 2026
The question we ask
For every category, each AI model is asked the question a real buyer asks: “What are the best {category}? Recommend the top brands or products that people actually use.” Product categories, whose boards rank individual models, ask for specific models instead: “Recommend the specific models people actually buy”. No brand names are supplied either way. Regional editions ask the same question the way a local buyer would, naming the country.
Every monthly refresh asks each AI model the same buyer question once, and the exact run count behind every edition is published in its JSON record. Every figure is therefore reproducible. One engine’s quirky answer cannot create a ranking on its own: the consensus gate below requires at least two AI models to agree before a brand is ranked at all.
The engines
All 9 leading AI models, every refresh: no engine is skipped because it is inconvenient to reach. The exact model version behind every captured answer is recorded with it.
The score: the full formula
Brands are ranked by their AI Recommendation Score (0 to 100). The complete calculation is published here so any position can be recomputed from the answers. There is no editorial weighting and no hidden step: every brand recommendation is extracted from every captured answer, a brand counts at most once per answer however many of its products that answer names, a consensus gate requires at least two AI models, and the score combines share of voice, mention rate and how early the AI models name the brand.
Ties: when two brands land on the same score, the tie breaks on average position (earlier wins), then on total answers naming the brand. Ranks are strictly sequential, so three brands tied on 67.5 rank 2, 3 and 4, never joint 2nd.
Why there is no citation component: our July 2026 audit showed citations reward the wrong thing. Review sites that answers cite as sources outranked products that six of eight AI models actually recommended, so citations were removed from the score.
The gates
A brand must be recommended by at least two different AI models to be ranked at all. One engine’s quirk is not a market position.
Every ranking keeps its receipts: the verbatim answers the AI models gave are stored and shown, so any position can be checked against the answers that produced it.
Each monthly refresh is captured as an immutable snapshot. Movement arrows compare against the previous snapshot; past editions are never edited.
What never ranks: entity exclusions
A ranking must contain entities of the category, so two published exclusion rules run before scoring (added 7 August 2026, dated in the changelog below). First, in service-provider categories, the directories and marketplaces that rank providers are never themselves ranked: an entity is excluded when its domain sits on our curated source table as a directory or marketplace. Second, the per-category exclusion list below removes brands AI models name in passing that are not members of the category, such as an ad platform named inside an agency recommendation. Both rules require the entity’s name to identify it as the domain’s owner, so a real provider an answer happened to attribute to a directory’s domain always stays ranked, and a per-category exception list keeps legitimate incumbents in place (Amazon remains a third-party-logistics provider).
Excluded mentions leave the share-of-voice pool, so remaining scores reflect only category members. Nothing is removed from the record itself: every excluded mention stays verbatim in the published answer corpus and its extractedBrands, so any board remains recomputable from its own record. The list changes only with a dated changelog entry.
Tamper-evident records
Every refresh freezes its receipts twice. At capture time, the complete verbatim answer corpus is hashed with SHA-256 and the hash is stored on the immutable snapshot and published in the record’s JSON download, next to the answers themselves. Anyone can recompute the hash from the published corpus, using the spec shipped in the same file, and confirm the answers were not edited after publication.
Once a monthly sweep is verified complete, the per-category hashes are combined into a single run-level root hash, published here and inside that month’s Index report. One changed character in any archived answer changes its category hash, which changes the root.
Show earlier editions (4)Hide earlier editions
Non-determinism, sampling and variance
The same AI model, asked the same question twice, does not reliably give the same answer. This is a property of the models, not a flaw in any measurement, and any methodology that does not address it is marking its own homework. We measured it on our own corpus: of 535 engine-prompt pairs asked at least twice over 60 days, only 16.3% produced the identical set of search queries every run (published August 2026, with the full sample). A single point-in-time check of any AI answer is one roll of the dice.
The Index is designed around that fact rather than around pretending it away. Rankings never rest on a single roll: the consensus gate requires at least two independent AI models before a brand ranks at all, so one run’s quirk cannot create a position. Movement is read edition over edition against immutable snapshots, never inside a single run. The exact run count behind every edition is published in its JSON record, the exact model version is recorded on every answer, and model changes are dated in the changelog below with their expected variance noted, so month-over-month movement is always attributable to either the market or the method, never silently to both.
Sampling cadence is deliberate and disclosed: the public Index refreshes monthly in a single verified sweep, and CiteHawk workspaces collect weekly. More runs per refresh would smooth variance further, and the complete verbatim corpus is published precisely so anyone can quantify the remaining variance themselves rather than taking our word for it. That is the trade we choose: fewer, fully published, tamper-evident runs over many unpublishable ones.
Regional rankings
Some categories (banks, insurers, agencies, professional services) get genuinely different answers in different countries, so those categories are also asked as a local buyer would ask, for the United States, the United Kingdom, Australia and Canada. A regional edition is published only when its ranking meaningfully differs from the global one: a different #1, or fewer than three-quarters of the top 10 in common.
Rankings are computed from AI responses only. Claiming a brand cannot change its position, and positions are not for sale.
Brand owners can claim their brand to verify identity details and follow their movement. Identity corrections are reviewed and never affect scores. No payment, partnership, or relationship with CiteHawk influences any ranking.
The changelog
Dated version history of the methodology.
Show earlier changes (15)Hide earlier changes
What the Index is not
The Index reports what AI models recommend. It is not an endorsement by CiteHawk, and it is not a review site. If an AI model is wrong about a category, the Index will faithfully show you that it is wrong. That is the point.

Want this measurement for your own brand?
CiteHawk tracks the same signals for your brand across all 9 AI models, every week, receipts included.
Free AI visibility report · No credit card · 50 prompts · 10 engines
