# How to Set Up AI Visibility Tracking: A Step-by-Step Guide

**Author:** John Morabito (Founder, /winston)
**Published:** September 17, 2026
**Reading time:** 9 minutes
**Canonical:** https://www.winstondigitalmarketing.com/playbooks/how-to-set-up-ai-visibility-tracking/

Plenty of articles will tell you why AI visibility matters. Far fewer tell you how to actually stand the tracking up, step by step, so you go from wondering whether the AI recommends you to having a real, running measurement program. That is what this is: a plain, tool-agnostic setup checklist you can follow this week. The steps are the same whether you run it by hand to start or automate it with a tool, so I will walk the whole thing and then show where a tool takes over the grind. No theory, just the setup.

If you want the concepts underneath first, why this metric matters and how it is defined, that is [how to measure AI share of voice](https://www.winstondigitalmarketing.com/playbooks/how-to-measure-ai-share-of-voice/). This piece assumes you are sold on the why and want the how.

## Step 1: Build the prompt set

The prompt set is the foundation, and everything else depends on getting it right. It is the list of questions you will run over and over, so it needs to reflect what your buyers actually ask an assistant, phrased the way they would phrase it, not a list of keywords.

Cover the buying-intent questions where being named actually matters: the best option for a job, a recommendation for a specific situation, alternatives to a competitor, and the local or category variants that fit your business. Deliberately include prompts you suspect a competitor wins, because confirming that is useful. Then tier the set: a small core of the highest-stakes prompts you care most about, and a wider set you monitor more loosely. Two rules keep this sound. Write real buyer language, and once you have built the set, keep it fixed, because a stable prompt set is what makes the trend meaningful; if you change the questions each month you have nothing to compare. Building this well is its own craft, covered in [GEO prompt research](https://www.winstondigitalmarketing.com/playbooks/geo-prompt-research/), but for setup you need a fixed, tiered list of the real questions that matter.

## Step 2: Pick the engines

Track the five engines where cited, real-time answers are actually read and acted on: ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, and Gemini. The reason to track all five rather than just the biggest is that they genuinely disagree with each other, drawing on different sources and giving different answers to the same question, so a single-engine check gives you a false picture of where you stand. If some engine truly does not matter for your audience you can drop it, but understand what you are giving up. Note that Claude is not in the set, because it does not surface the same kind of cited, live-web recommendations these five do, so tracking it would not reflect the answers your buyers are getting.

## Step 3: Set the cadence per tier

Do not track everything at one frequency. Match cadence to stakes: run your core high-stakes prompts often, ideally daily, so you catch a drop or a competitor gain while you can still act on it, and run the wider set weekly or monthly. The reason tiered cadence matters is that AI answers are volatile, so a monthly-only check can miss a change that happened and reversed inside the month, while running your entire long tail daily is usually more than you need and more than you will read. Decide each tier by a simple test: how much would it cost you to find out about a change a month late? Where the answer is a lot, track it often. The full reasoning is in [daily vs monthly AI visibility tracking](https://www.winstondigitalmarketing.com/playbooks/daily-vs-monthly-ai-visibility-tracking/).

## Step 4: Decide what to record

For every prompt on every engine, record the same four things, so your data is consistent and comparable over time:

- **Named or not.** The base presence signal: does the answer mention you at all.
- **How you are described.** The sentiment and the specific framing, because tone matters as much as presence.
- **Which competitors appear.** Named instead of or alongside you, which gives you the competitive picture.
- **Which sources are cited.** The most actionable field and the one people most often skip. The cited sources are the map of where the engine built its answer, so they tell you exactly where to earn presence to change it.

Record all four the same way every time. Consistency is what lets you compare periods and spot real movement. If you capture only whether you are named, you have a scoreboard with no instructions; the cited sources are what turn tracking into a to-do list, which is why they matter so much.

## Step 5: Capture a baseline

Before you change anything, run the full set once and record it. That first reading is your baseline, and it is what everything afterward is measured against. Without it you will never be able to say whether your GEO work moved the needle, because you will have no before. Run it, save it, and resist the urge to start fixing things until you have that starting point captured. This baseline is also exactly what a free AI visibility audit produces, so if you want a fast, done-for-you starting point rather than a manual first pass, that is the shortcut, and [what a free AI visibility audit reveals](https://www.winstondigitalmarketing.com/playbooks/what-a-free-ai-visibility-audit-reveals/) walks through what that baseline shows.

## Step 6: Run the review loop

Tracking that nobody reads is wasted, so the last setup step is a habit, not a tool. On a set schedule, review the deltas rather than the raw levels: what moved, which prompts you gained or lost, which competitors are rising, and which cited sources are driving the answers you care about. Turn the gaps into action, a source you are absent from is a place to earn presence, a worsening description is a reputation issue to address, and then check next period whether the work moved the number. That diagnose, act, re-measure loop is the entire point of setting this up, and running it as an ongoing discipline rather than a one-time audit is the subject of [AI visibility monitoring](https://www.winstondigitalmarketing.com/playbooks/ai-visibility-monitoring/).

## Manual to start, tool to operate

You can do all six steps by hand, and doing a manual pass once is genuinely worth it: run a handful of prompts across the engines yourself, and you will see for yourself whether you are named, and have something concrete to show stakeholders. It costs nothing but an afternoon.

What manual tracking cannot do is survive a real schedule. Running dozens of prompts across five engines every day or week, parsing each answer, logging competitors and sources, deduping the run-to-run noise, and holding the trend is a grind nobody sustains by hand, and a one-time manual check tells you nothing about the trend, which is the whole value. So the honest path is manual to validate, tool to operate. The setup steps above do not change; a tool just automates the running, recording, and trending so your time goes to acting on the data instead of collecting it. If you would rather understand the cross-engine capture in depth, that is [tracking AI citations across engines](https://www.winstondigitalmarketing.com/playbooks/tracking-ai-citations-across-engines/).

That is what the [Winston GEO Tracker](https://www.winstondigitalmarketing.com/geo-tracker/) is: the fast path through this whole checklist. You give it your prompt set, it runs across the five engines that matter (ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, and Gemini; it does not track Claude, which does not surface the same cited answers) at whatever cadence you set, records all four fields per answer including the cited sources, holds the baseline and the trend, and gives you the review surface so the loop actually happens. At $0.75 per prompt it is affordable to run the whole set at the cadence each tier deserves rather than rationing. And the simplest way to complete steps 1 through 5 in one move is the free AI visibility audit, which builds a starter prompt set, runs it across the engines, and hands you the baseline, no call required. We run this setup and the GEO work that acts on it as part of our [generative engine optimization](https://www.winstondigitalmarketing.com/services/generative-engine-optimization/) practice.

The whole setup, one more time: build a fixed, tiered prompt set of real buyer questions; track the five engines; set cadence by stakes; record named, described, competitors, and cited sources; capture a baseline; and run the review loop. Do that and you have turned an anxious guess about whether the AI recommends you into a measured program you can actually improve.

## Frequently asked questions

### What do you need to set up AI visibility tracking?

Five things, and none of them require a big lift to start. A prompt set, which is the list of questions your buyers actually ask an assistant. A choice of engines, realistically the five that matter. A cadence, meaning how often you run each prompt. A definition of what to record for each answer, which is whether you are named, how you are described, which competitors appear, and which sources are cited. And a baseline plus a review loop, so you have a starting point and a habit of acting on what the tracking shows. That is the whole setup. You can do a rough version by hand to prove the value, then move to a tool to run it reliably on a schedule. The steps are the same whether you do it manually or automate it; the tool just removes the grind.

### How do you build a prompt set for AI visibility tracking?

Start from real buyer questions, not keywords. Write the prompts the way a customer would actually phrase them to an assistant, covering the buying-intent questions where being named matters (the best option for a job, a recommendation for a situation, alternatives to a competitor, local or category variants that fit your business). Include the prompts you expect competitors to win, because confirming a suspicion is useful. Then tier them: a small core of high-stakes prompts you care most about, and a wider set you monitor more loosely. Keep the set fixed once you build it, because a stable prompt set is what makes the trend over time meaningful; if you change the questions every month you cannot compare periods. Proper prompt research is its own craft, but for setup you need a fixed, tiered list of the real questions that matter.

### Which engines and cadence should you track?

Track the five engines where cited answers are actually read and acted on: ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, and Gemini. They disagree with each other, so tracking only one gives a misleading picture. For cadence, match frequency to stakes rather than tracking everything at one rate: run your core high-stakes prompts often, ideally daily, so you catch changes while you can act, and run the wider set weekly or monthly. The reason to tier cadence is that AI answers are volatile and a monthly-only check can miss a change that happened and reversed, while tracking a large set daily is usually more than you need for the long tail. Decide cadence per tier by asking how much it would cost you to find out about a change a month late.

### What should you record for each AI answer?

Four things per answer, because together they turn a vague impression into an actionable record. Whether you are named at all, which is the base presence signal. How you are described, the sentiment and the specific framing, because tone matters as much as presence. Which competitors are named instead of or alongside you, which gives you the competitive picture. And crucially, which sources the answer cites, because those cited sources are the map of where the engine formed its answer and therefore where you need to earn presence to change it. Record all four consistently for every prompt on every engine, and do it the same way each time so the data is comparable over time. The cited sources are the most actionable and the most commonly skipped; do not skip them, because they are what turn tracking into a to-do list.

### Should you set up AI visibility tracking manually or use a tool?

Do a manual pass once to prove the value, then use a tool to run it for real. A by-hand run of a handful of prompts across the engines is a good way to see for yourself whether you are named and to convince stakeholders it matters, and it costs nothing but time. But manual tracking does not survive a real schedule: running dozens of prompts across five engines every day or week, parsing each answer, logging competitors and sources, and holding the trend is a grind nobody sustains, and a one-time check tells you nothing about the trend, which is the whole point. So the honest answer is manual to start and validate, tool to operate. The setup steps are identical either way; a tool just automates the running, recording, and trending so you spend your time acting on the data instead of collecting it.
