TL;DR. Tracking brand mentions in one model is a tutorial. Tracking across three is a workflow. This guide shows a repeatable manual process for ChatGPT, Claude, and Perplexity, the spreadsheet that keeps it consistent, and the signals to capture so you can diagnose what to fix next. This is Step 2 of the 4-step AI visibility audit, pair it with the free 15-prompt starter set for the full workflow.

▶ Watch the full walkthrough on YouTube →
Different assistants can give different answers to the same product question. If you track only one, you see one platform’s output rather than a cross-platform view of how your brand is represented.
This post is the operational guide for Step 2: running the same prompt set across ChatGPT, Claude, and Perplexity, capturing comparable outputs, and logging the details that make the results actionable.
If you need the operating model behind this workflow, start with our AI brand monitoring guide. It covers the scorecard, review cadence, and when a manual spreadsheet stops being practical.
Why track brand mentions across multiple AI platforms?
The practical reason is variance. Different systems can produce different brand sets, different phrasing, different citations, and different confidence even when the prompt is identical. The more you rely on a single surface, the more likely you are to mistake one model’s behavior for your market reality.
A multi platform pass helps you seperate distinct outcomes that often get mixed together:
- Mention: your brand name appears at all.
- Citation: your brand or a claim about your brand is supported by a linked source (where the interface provides citations).
- Recommendation: the assistant positions your brand as a suggested option for the user’s goal.
- Sentiment: the language is positive, neutral, or negative.
- Factual accuracy: the description matches what is true about your product.
Treat these as separate fields. A brand can be mentioned without being recommended, recommended without being cited, or cited with inaccurate details.
Also, do not treat list order as a stable rank. Many assistants vary ordering across runs, and interfaces evolve. Your job in tracking is to capture what happened in a specific run, then look for consistant patterns over time.
What you need before you start tracking
You can do a first round with free accounts and a spreadsheet.
-
Accounts on each platform
Create accounts for ChatGPT, Claude, and Perplexity. Tier, default model, and controls change over time, so treat your account setup as a variable. -
A locked prompt set
Use the 15 starter prompts from the 4-step audit pillar, or build a custom set of 10 to 15 prompts from sales calls, support tickets, and community threads. Freeze the set so you can compare runs. -
A spreadsheet template
Keep it simple but explicit. Suggested columns: - Prompt - Platform - Model or mode (as labeled in the UI) - Run date and time - Run number (1, 2, 3) - Mentioned (yes/no) - Recommendation (yes/no) - Citation present (yes/no, and source domains if applicable) - Position in answer (capture as observed, but do not treat as rank) - Accuracy notes (wrong features, stale pricing, wrong category, etc.) - Sentiment notes (optional) -
Focused time and consistent conditions
Manual tracking time varies by prompt complexity, response length, and interface speed. Avoid promising a fixed time budget. Instead, aim for consistency: same prompt set, same run window, and the same logging detail each round. -
A simple repeatability rule
For each prompt, run multiple times and log the variance. If you can't run multiple times, record that explicitly and treat the output as directional.
How to build a tracking worthy prompt set
Stage 1 is choosing prompts that match how buyers actually ask. For a repeatable audit, you want prompts that are:
- Natural language: how a person would type it
- Category anchored: the category term appears, not only brand names
- Decision oriented: implies comparison, constraints, or a goal
If you want a default starting point, use the 15 starter prompts linked from the 4-step audit pillar.
If you want to customize, pull candidates from: - Sales call questions and objection handling notes - Support ticket phrasing - Community thread titles in your category - Competitor comparison pages that rank for your category terms
Whatever set you choose, freeze it. Replace prompts only when category language or your ICP changes enough that the existing prompts stop reflecting real buyer queries.
How to track brand mentions in ChatGPT
Stage 2 begins with ChatGPT.
What to log for every ChatGPT run
For each prompt, capture:
- Mention: yes/no
- Recommendation: yes/no, and the phrasing used
- Accuracy: correct category, correct capabilities, correct positioning
- Any citations shown in the interface: copy the linked domains or URLs when present
- Run context: model or mode label shown in the UI, and the date
ChatGPT interfaces change, and the same applies to model defaults and search behavior.
Do not rely on a stable toggle name or placement. The safest approach is to record the mode and date for each run.
ChatGPT Search notes (what to do, and what not to assume)
OpenAI describes ChatGPT Search as a feature that can search the web and provide inline citations, and that it can use multiple third party search providers. Reference: https://help.openai.com/en/articles/9237897-chatgpt-search
Practical implications for tracking:
- Do not treat ChatGPT as training data only. Depending on mode and prompt, it may search automatically and may show citations.
- Do not attribute results to a single provider. OpenAI indicates multiple third party providers can be used.
- Make runs comparable. If you are comparing quarter over quarter, record whether Search appeared to be in play and whether citations were shown.
Run hygiene for ChatGPT
- Start a fresh chat for each prompt to reduce cross prompt contamination.
- Keep the prompt text identical across platforms.
- If you do multiple runs, spread them over a short window so your run set reflects the same content environment.
If you want the deeper platform specific setup, see how to track ChatGPT brand visibility.
How to track brand mentions in Claude
Stage 3 is the Claude pass.
What to log for Claude
Use the same columns as ChatGPT so your spreadsheet aggregates cleanly:
- Mention and recommendation
- Accuracy issues
- Any consistent hedging language that affects whether your brand is framed as a safe choice or an edge case
- Run context: model label, workspace context if relevant, date
Claude run hygiene
- Start a fresh conversation for each prompt.
- Avoid running inside a context heavy workspace if it injects additional documents or memory that makes the run non comparable to baseline usage.
- Capture variance across runs rather than chasing a single best answer.
A common failure in Claude tracking is to over interpret tone. Claude can be conservative about naming brands or may ask clarifying questions. That can be a style or safety posture, not a visibility problem by itself. The tracking question is whether your brand is present, whether it is recommended, and whether the description is correct.
How to track brand mentions in Perplexity
Stage 4 is the Perplexity pass.
What Perplexity is useful for in a tracking workflow
Perplexity describes its approach as searching the web and providing answers with citations. It is designed to help users inspect sources. Reference: https://www.perplexity.ai/help-center/en/articles/10352895-how-does-perplexity-work
For brand tracking, this typically makes Perplexity your most source inspectable surface:
- Capture whether your brand is mentioned and recommended.
- Capture the linked citations and the domains that appear repeatedly.
- When your brand is missing, capture which sources are cited for competitors.
Avoid universal claims about how Perplexity ranks or weights sources. Use it as an evidence trail for the specific prompt and run you executed.
Perplexity run hygiene
- Keep the prompt identical to your ChatGPT and Claude prompt.
- Record the mode label shown in the UI and the date.
- Save citations as URLs in your sheet so you can audit them later, especially when the assistant describes your product incorrectly.
For deeper Perplexity tactics, see tracking brand visibility in Perplexity.
See also: AI visibility tools compared
If you are reaching the point where you want to graduate from manual tracking, these companion posts cover the tools landscape and evaluation angles:
- Best AI visibility tools compared (2026), a broader landscape with feature and pricing comparisons
- Profound vs Otterly: AI visibility tools compared, a public-plan comparison of the two platforms
- AI visibility platform comparison (2026), Profound vs AthenaHQ and the mid-market alternatives
If you are using these pages to evaluate vendors, check the pricing sections for freshness. Last verified: July 26, 2026.
How to score and aggregate results without pretending there is a stable rank
After you finish runs across platforms, reduce the raw logs into a few summary metrics that remain meaningful even when ordering shifts.
Recommended summary metrics
Per platform:
- Mention coverage: number of prompts where you are mentioned at least once within your run set.
- Recommendation coverage: number of prompts where you are recommended (not just listed).
- Citation coverage (where applicable): number of prompts where your brand is linked or supported by citations.
- Accuracy issues count: prompts with at least one factual error about your product.
Cross platform:
- Cross platform overlap: prompts where all tracked platforms mention your brand at least once in the run set.
- Cross platform recommendation overlap: prompts where all tracked platforms recommend your brand at least once.
If you still want to record order, treat it as an observation:
- Log position as seen for the run.
- Compute an average only if you are also logging variance and run context.
- Do not interpret a single run list order as a durable leaderboard.
Illustrative summary table
Use an output format like this to make your next step obvious. The numbers below are illustrative, not benchmarks.
Table 1: Illustrative example output format.
| Platform | Prompts mentioned (of 15) | Prompts recommended (of 15) | Prompts with citations (of 15) | Accuracy issues (count) | Notes |
|---|---|---|---|---|---|
| ChatGPT | Record model or mode and date; capture citations if shown | ||||
| Claude | Record model or mode and date; note hedging patterns | ||||
| Perplexity | Save cited URLs and domains |
How to read patterns and prioritize what to fix
Treat the tracking round as diagnosis input, not a score to celebrate.
1) Start with recognition, then recommendation, then validation
If you are missing from brand anchored prompts, that is a recognition problem. If you are present but rarely recommended, that is a positioning and comparative authority problem. If you are recommended but described inaccurately, that is an accuracy and narrative control problem.
These are different fixes. Tracking works when your sheet clearly separates them.
2) Use Perplexity citations as your evidence trail
Where Perplexity provides citations, it can tell you which URLs are shaping the answer for that prompt. This is often the fastest way to identify:
- Which third party pages repeatedly define your category
- Which reviews, listicles, or forum threads keep appearing for competitors
- Which sources are outdated or incorrect about your product
Then you can prioritize: update your owned content, improve third party coverage, or correct misinformation on the sources that are actually being used.
3) Fix prompt level gaps before platform level narratives
It is tempting to say you are weak in one platform. The higher leverage view is usually prompt level:
- If you are absent across all platforms for the same category only prompt, that suggests a category association gap.
- If you are present in some platforms but absent in one, look at the cited sources and the phrasing used. It may be a retrieval difference, or it may be that one platform is drawing from different surfaces for that query.
For more on common patterns, see 11 AI visibility failure modes guide.
When to graduate from manual to automated tracking
Manual tracking is useful when you need clarity on what is happening and why. It gets harder when you need cadence, scale, and reliable historical comparisons.
Automation becomes more compelling when:
- You need frequent runs and consistent timestamps without manual effort
- You track multiple brands or multiple product lines
- You need clean trend reporting over time, including changes in wording and citations
If you are exploring automation, use your manual sheet as your requirements document. It tells you exactly which fields you need a tool to capture: mention, reccomendation, citations, accuracy notes, run context, and timestamps.
This is the problem friction AI is designed to address. If you evaluate tools, compare them against the criteria your workflow actually needs, and validate outputs by spot checking prompts inside the platforms.
Frequently Asked Questions
How do I know if ChatGPT mentions my brand?
Run a small set of buyer realistic prompts in ChatGPT and log whether your brand is mentioned, recommended, and described accurately. For each run, record the model or mode label and the date. ChatGPT may use web search and may show inline citations depending on the experience you are in, so capture citations when they appear. Reference: https://help.openai.com/en/articles/9237897-chatgpt-search
Should I run each prompt in a new chat?
If you want comparable results across prompts, use a fresh chat or conversation per prompt. Prior context can bias responses, and that bias is hard to detect later when you are aggregating.
How do I track brand mentions in Claude?
Run the same locked prompt set in Claude and log mention, recommendation, and accuracy. Record the model label and date. Focus on whether Claude names your brand and how it frames tradeoffs, not on whether the wording is more cautious than another platform.
How do I see sources in Perplexity?
Perplexity searches the web and links citations so you can inspect sources. For every prompt, save the cited URLs or at least the source domains, especially when your brand is missing or described incorrectly. Reference: https://www.perplexity.ai/help-center/en/articles/10352895-how-does-perplexity-work
What is the best AI brand monitoring tool in 2026?
It depends on your requirements: platforms covered, fields captured (mentions vs citations vs recommendations), cadence, reporting, and whether you need collaboration. Use the tool comparisons here and validate by spot checking outputs against live runs in the assistants:
- best AI visibility tools compared (2026)
Pricing sections: Last verified: July 26, 2026
How often should I run brand visibility checks?
Pick a cadence you can maintain with consistent prompts and logging. Quarterly is a common baseline for manual tracking. If you need more frequent reporting, record run context carefully because platform interfaces and defaults change, and you want your time series to remain interpretable.
How long does a multi platform tracking round take?
It varies by prompt length, platform speed, and how many runs you do. Instead of targeting a universal time budget, standardize your process: identical prompts, consistent run counts, and explicit recording of model or mode and date.
Can I use paid tiers instead of free tiers?
You can, but treat paid tier runs as a separate dataset. Model defaults, modes, and limits vary over time and across tiers. The most important step is to record what you used in each run so comparisons remain valid later.
What is the minimum useful prompt count?
Ten prompts is a practical floor for a directional read. Fifteen prompts is a common size because it lets you cover brand anchored prompts, category only prompts, and comparison or validation prompts. If you go larger, the value comes from better coverage of your category vocabulary, not from sheer volume.
How do I track a prompt where my brand is not mentioned?
Log it as not mentioned, then capture which brands were mentioned, whether any were recommended, and any citations or sources shown. Missingness is part of the competitive map, and the cited sources often point to what is shaping the answer.
Run this audit on your own brand
Want Step 2 running across ChatGPT, Claude, and Perplexity on a continuous schedule, with run context and deltas captured automatically?
▶ Start your free trial of friction AI →
Or grab the free 15-prompt starter pack → and run the manual workflow.