Feature lists make AI search tracking tools look alike. A trial on the brand's own questions makes the differences visible within two days, and it needs only a fixed question set and a second run.
This guide covers that trial. The full list of evaluation criteria sits in the main evaluation guide, and the manual baseline covers the first afternoon of work before any tool is involved.
Key takeaways
- A fair trial puts the same fixed set of questions through every candidate tool, on the same engines and in the same market. Different question sets cannot be compared.
- A second run the next day shows how much movement is normal. In a Zumi run of 304 beauty-category questions, 7.3% of day-to-day repeats mentioned exactly the same brands, and the average change on a 0 to 1 scale was 0.60.
- A tool showing almost no change between the two days deserves a question about whether it asked the engines again. A tool showing large change should show how many answers sit behind each figure.
- The underlying answers settle disagreements. A figure that cannot be traced to a real answer cannot be checked.
How should the trial be set up?
Four decisions come before any tool is switched on:
- The questions. Twenty to thirty questions a real buyer would ask, written by the team rather than generated from keywords. Some should include the brand name and some should describe the category without it.
- The engines. The engines the brand's buyers use, which every candidate must cover. A tool that cannot track one of them is out before the trial starts.
- The market. One country and one language per run.
- The record. A sheet with one row per tool and one column per check below, filled in as the trial runs.
What should be recorded on each day?
| Check | What to record | Worth a question to the vendor |
|---|---|---|
| Day-to-day change | How far the brand list moves for the same question between day one and day two | Almost no movement, or no way to see the repeats |
| Answer counts | How many answers sit behind each percentage, per engine and period | A percentage with no count beside it |
| Traceability | Whether any figure opens to the real answer text | A score that cannot be traced to answers |
| Agreement between tools | Whether two tools differ on the same question and the same day | Differences nobody can explain by opening the answers |
| Engine and plan limits | Which engines and how many questions the plan includes, and what the next tier adds | Coverage that depends on an upgrade |
| Output | Whether the tool recommends specific fixes or stops at a score | A dashboard with no next step |
What counts as normal day-to-day change?
Answers from the same engine to the same question vary from one day to the next, so two runs are never expected to match. Zumi measured this on 304 beauty and personal care questions in India, asked daily on ChatGPT, Perplexity, and Google AI Overviews from 23 July to 2 August 2026. The 7,280 pairs of back-to-back answers broke down like this:
| Same question, same engine, one day apart | Share of pairs |
|---|---|
| Exactly the same brands mentioned | 7.3% |
| More than half the brand list changed | 58.8% |
| No brand in common at all | 15.3% |
Other categories may move more or less, so the figures are a reference point, not a target. They do show what the trial should look like: a tool whose two days differ is behaving like the engines. A tool whose two days match exactly warrants a question about whether the engines were asked twice.
The engines also differ from each other, so a trial should be read engine by engine:
| Engine | Average day-to-day change (0 = same, 1 = nothing shared) |
|---|---|
| Perplexity | 0.64 |
| ChatGPT | 0.63 |
| Google AI Overviews | 0.54 |
A tool's Perplexity figures should move more between the two days than its Google AI Overviews figures. A tool that reports the same amount of change for every engine deserves a question about how its per-engine figures are built. The record sheet should therefore split each check by engine.
Two tools can also disagree on the same question on the same day, and both can be right for the moment they ran. Opening the answers is what settles it.
What decides the winner?
Not the dashboard. The tool worth paying for shows how often each question runs, how many answers back each figure, and the answers themselves, and it covers the engines the brand's buyers use. Price and agency fit come after those checks.
The guide to citation views covers what the source data in a tool should show.
Zumi, an AI Search Intelligence Platform, refreshes tracked questions daily and reports mention rate, share of voice, average position, and citation share for each tracked engine. Book a demo to run the trial on a brand's own questions.
Sources
- Zumi, India beauty and personal care question run, 304 questions asked daily on ChatGPT, Perplexity, and Google AI Overviews, 23 July to 2 August 2026.