A brand can appear in one ChatGPT answer and disappear from the next. A different prompt, account state, location, or answer product can change the result again. A manual baseline must freeze the measurement first.
The engine run comes after.
That baseline does not need to prove that a platform is necessary. Its purpose is to replace accidental screenshots with a bounded observation, then support a decision to repeat manually, automate the work, or stop.
Key takeaways
- A defensible baseline fixes the prompt version, engine or answer product, geography, language, account state, date, run count, and denominator.
- The worksheet retains raw answers and cited URLs before it calculates mention rate, share of voice, average position, or citation share.
- Same-period repetitions show answer variation. A later run with the same measurement contract shows movement over time.
- Prompt count and engine count are scope choices. No universal total fits every market, buyer group, or decision.
- Continuous monitoring is warranted when the decision recurs, the scope is costly to run manually, and several teams need the evidence.
What decision should the baseline support?
A baseline begins with a decision, not a prompt list. The decision might be whether an important buyer-question cluster deserves monitoring, whether one market has a discovery gap, or whether a recurring executive report can be maintained manually.
The scope note contains:
- Decision owner
- Market, geography, language, and audience
- Product, category, or segment in scope
- Candidate engines or answer products
- Declared competitor set
- Baseline collection window
- Planned review and next decision
The note also names what is outside the exercise. Traffic, conversion, pipeline, and revenue are not AI Visibility signals. A separate analytics system may examine those outcomes, but the manual baseline should not infer them from answer observations.
Which questions belong in the prompt panel?
The prompt panel represents the questions that can expose, compare, or exclude brands during discovery. Sales conversations, on-site search, support logs, product-category pages, and traditional search queries can inform the panel.
The first version usually contains several intent classes:
- Category discovery: Observes which brands enter an unprompted category answer. A pattern is “Which platforms help a mid-market team manage [job]?”
- Use case: Observes which brands fit a task or audience. A pattern is “What is a suitable [category] for [segment] that needs [requirement]?”
- Comparison: Observes how the engine separates known options. A pattern is “How do [option] and [option] differ for [use case]?”
- Problem framing: Observes which categories and vendors appear before a shortlist exists. A pattern is “How can a team solve [problem] under [constraint]?”
- Branded validation: Observes whether known brand facts are described correctly. A pattern is “What does [brand] support for [requirement]?”
Unbranded questions carry most of the discovery read. Branded questions answer a different question because the prompt has already supplied the entity.
Every prompt receives a stable ID, exact text, intent label, topic label, and version. A material wording change creates a new version. It does not silently replace the question in the original baseline.
There is no universal prompt total. A focused category in one market may need a smaller panel than a multi-product brand operating across several regions. Coverage of the decision matters more than a round number.
Which engines and test conditions must be fixed?
Engine selection follows the audience and the decision. The record names the exact interface or answer product rather than grouping every product from one company into a single result.
Google AI Overviews and Google AI Mode, for example, are different answer products. A result from one should not be labeled as evidence from the other. An API model with attached web search also should not be described as a pixel-for-pixel consumer interface.
The measurement contract records:
- Engine or answer product: Exact product name and visible mode or model where available.
- Geography and language: Market context used for the collection.
- Account state: Signed in or signed out, with personalization state where known.
- Session rule: Fresh session for each independent answer.
- Prompt panel: Version and stable prompt IDs.
- Sampling: Predetermined repetitions and collection cadence.
- Date and time: Timestamp for each answer.
- Failures: Empty, blocked, or incomplete answers retained and classified.
- Counting: Mention, competitor, position, and citation rules.
Consistency matters more than an artificial claim of neutrality. An incognito window does not erase every contextual signal. A documented, repeatable state gives the result a clear boundary.
Which fields belong in the worksheet?
The workbook keeps definitions, observations, citations, calculations, and decisions on separate tabs. That separation prevents an analyst's interpretation from overwriting the original answer.
Prompt registry
The prompt registry stores:
- Prompt ID and exact text
- Intent and topic labels
- Prompt version and active status
- Inclusion reason
- Canonical brand and competitor names
Raw answer log
Each row represents one prompt, one engine or answer product, and one repetition:
- Identity: Run ID, prompt ID, prompt version, and engine.
- Context: Date, time, geography, language, account state, and session state.
- Evidence: Full unedited answer, screenshot or archive reference, and cited URLs.
- Brand observation: Brand mentioned, mention position, and owned page cited.
- Competitive observation: Competitors mentioned and their positions.
- Quality control: Parse status, failure reason, reviewer, and notes.
The raw answer remains unchanged. A correction or coding review adds a new field or revision record instead of editing the evidence.
Citation log
The citation log records the source URL, domain, source owner, cited claim where visible, and whether the page belongs to the tracked brand. A third-party article that mentions the brand remains a third-party source.
Decision register
The decision register separates four fields:
- Observation: What appeared in the answer set.
- Interpretation: The plausible meaning of that pattern.
- Action: The owned- or earned-media work selected.
- Remeasurement: The unchanged panel and date used to review movement.
The distinction makes uncertainty visible. A cited competitor page can explain part of an answer, but its presence does not prove that the page caused the brand's absence.
How should the four signals be calculated and read?
The baseline uses the four registered AI Visibility signals without combining them.
- Mention rate: How often the brand appears in eligible tracked answers. The record includes the eligible-answer denominator and absence handling.
- Share of voice: How often the brand appears relative to declared competitors. The record includes the competitor set and share denominator.
- Average position: Where the brand lands among mentioned brands. The record includes the position rule and treatment of unordered prose.
- Citation share: How often the brand's own pages are cited. The record includes the total citation denominator and owned-domain rule.
The AI Search Intelligence metric dictionary defines each signal and its common misread. The baseline should retain both numerator and denominator wherever a percentage is reported.
The four signals can move in different directions. A brand may gain mentions while appearing later in the answer. It may retain the same presence while its own pages begin to support more answers.
Those patterns are the reason the signals stay separate. One composite would hide which source, prompt cluster, or engine needs attention.
What is the difference between repetition and monitoring?
Same-period repetition asks the same question several times under broadly consistent conditions. It shows how much the answer varies within the baseline window.
Monitoring repeats the measurement contract later. It shows whether the observed pattern changed as engines, source indexes, owned pages, earned coverage, and competitors changed.
Neither is a controlled causal test. A before-and-after observation can record that a content change preceded movement in the affected prompt cluster. It cannot rule out every other change in the information environment.
The number of repetitions should be chosen before the answers are reviewed. A single answer remains a spot check. More repetitions can describe variation better, but no universal floor applies to every decision.
When is continuous monitoring worth buying?
The baseline should end with a buy, repeat, or pause decision.
- Repeat manually: The panel is narrow. The review cadence is low, one team owns the work, and raw evidence remains manageable.
- Automate or buy monitoring: The panel spans several engines, markets, products, or clients. The report recurs, answer retention matters, or several teams need access.
- Pause: The selected questions have little strategic value. The brand is consistently present where it matters, or no team has authority to act.
A platform earns its place by preserving the measurement contract, retaining evidence, reducing collection work, and helping the team route action. More decimal places do not make a weak prompt panel useful.
Once the baseline shows where a brand is missing, how to write content that AI engines cite covers the page-level fixes.
Zumi is an AI Search Intelligence Platform. It monitors a scoped set of up to nine AI engines, refreshes AI Visibility daily, and produces prioritized Owned Media and Earned Media recommendations.
A platform workflow review shows how the manual record becomes a continuous operating system once the spreadsheet no longer fits the job.