← Back to blog

How to Track Website Visibility in ChatGPT and AI Search

By Joe Della MoraFounder, GroundScore

monitoringbuyers-guide
Three methods for tracking website visibility in ChatGPT and AI search engines

There are three ways to track website visibility in ChatGPT and AI search results, and you can start the first one in the next five minutes. Ask the engines your buyers' questions yourself and log what comes back. Detect referral traffic arriving from AI tools in your analytics. Or run dedicated monitoring that asks a fixed question set on a schedule and charts the trend. Each measures something genuinely different, and none of them measures all of it. What makes this harder than it looks is that AI answers vary from run to run, so one check proves close to nothing. This guide covers what you are actually measuring, the three approaches weighed as a buying decision, and how to choose between them.

What you are actually measuring

Direct answer: Three different things get called visibility: whether an engine mentions your brand name, whether it cites your URL as a source, and whether it recommends you as the answer. They are not the same signal, they fail independently, and any tracking method you choose measures some subset of them.

Before picking a method, know what you are counting. The three states look similar in a chat window and mean very different things.

Mentioned. Your brand name appears in the text of the answer. No link, no attribution, just the name. This is worth something for recall and nothing for clicks.

Cited. Your URL appears as a source, usually in a footnote, a sidebar, or an inline link. This is the state most people mean when they say visibility, and it is the one that can actually send traffic.

Recommended. The engine does not just list you among sources; it puts you forward as the answer. For a local service business, this is the whole game.

A site can be cited constantly as background material and never recommended; another recommended by name and never linked. If your tracking collapses all three into one yes-or-no, you will misread your position.

Then there is the harder problem: variance. Ask the same engine the same question twice and you can get two different sets of sources. Model updates, live search results, session context, and plain sampling randomness all move the answer. Building GroundScore, the pattern I keep seeing is owners running one check, getting a bad result, and concluding they are invisible, or getting a good result and concluding they are fine. Neither conclusion is supported by a single sample. Whatever you use to track this, the unit of truth is a set of repeated observations, not one screenshot.

Approach one: ask the engines yourself

Direct answer: Open ChatGPT, Perplexity, and Claude, ask the questions your buyers ask, and write down whether your site appears. It costs nothing, takes about twenty minutes, and it is the fastest honest read available. Its weakness is sample size: a handful of manual runs cannot separate a real trend from noise.

This is a legitimate starting point and nobody should feel talked out of it. If you have never checked, it is the highest-value twenty minutes available to you, and it needs no budget, tool decision, or account.

The method is simple enough to do properly:

  1. Write down five to ten real buyer questions. Not your brand name. Real questions, phrased the way a customer would type them: "best emergency plumber in Tacoma", "what software handles recurring invoices for freelancers", "who does commercial roof inspections near me".

  2. Ask each question in each engine. ChatGPT, Perplexity, and Claude behave differently enough that checking one tells you little about the others.

  3. Record three columns per run. Mentioned, cited, recommended. Plus which competitors showed up, because that is often the more useful data.

  4. Repeat the whole set on the same day each month. Same questions, same order. Changing the questions between rounds destroys comparability, which is the most common way a manual program quietly stops being useful.

What you get is real evidence from the actual engines, with no vendor in between. What you do not get is consistency. You will skip a month, reword a question, run three checks instead of ten because it is Friday. And because answers vary run to run, a small irregular sample cannot tell you whether a change is real or just the dice.

There is also a subtler trap: asking from an account that has discussed your business before is not a clean read. Use a fresh session, and resist nudging the engine toward yourself.

We compare the two methods in more depth in manual spot checks versus automated monitoring. The short version: manual checking is excellent for discovery and poor for measurement.

Side by side comparison of three ways to track AI visibility over time

Approach two: detect AI referral traffic in your analytics

Direct answer: Filter your analytics for visits arriving from chatgpt.com, perplexity.ai, and similar hosts. This measures outcomes rather than presence: real people who read an answer and clicked through. It is free and it is proof of value, but it goes dark whenever an engine answers the question without sending anyone.

Every analytics package can do a version of this — the referring host is in the data. Segment for AI sources and you have a traffic line made of people who were mid-decision when they clicked.

The appeal is obvious: this is the only approach that produces a business number rather than a marketing one.

The limits are equally real, and they are structural rather than fixable:

  • No click, no data. AI answers frequently satisfy the question outright. You can be cited in a thousand answers and see a handful of sessions. Referral traffic measures the clicked subset of your visibility, which is a small and unrepresentative slice.
  • You learn nothing about absence. Zero AI referrals is ambiguous. It could mean you are never cited, or cited constantly in answers nobody clicks through from. Those two situations call for opposite responses.
  • Attribution is lossy. Referrers get stripped, and traffic from a desktop app or a copied link often lands in direct. The count you see is a floor, not a total.
  • It is backward-looking. By the time referral traffic drops, the citation loss happened weeks earlier.

None of that makes it optional. First-party AI-referral tracking is a real feature in GroundScore for exactly this reason: presence data tells you whether you are in the answer, referral data tells you whether that mattered, and you want both on the same screen. Just do not mistake the second for the first.

Approach three: dedicated monitoring on a schedule

Direct answer: Monitoring tools ask a fixed set of questions across several engines on a repeating schedule and store every result. That fixed question set plus repetition is the whole point: it converts scattered anecdotes into a trend line you can act on. You pay for that consistency, in money rather than in your own time.

Dedicated monitoring is the manual method with the human reliability problem engineered out. Same questions, same engines, same cadence, stored and comparable. GroundScore checks ChatGPT, Claude, and Perplexity, scores the result 0 to 100 across Authority and Trust, Site Readiness, and AI Presence, and re-runs it weekly on paid plans. Anything above 70 we band as strong, 40 to 69 moderate, below 40 weak.

What you are buying is not access to the engines. You already have that, free. You are buying four things:

  • Consistency. The same questions run whether or not you remember, which is what makes month-over-month comparison meaningful.
  • History. A stored record you can point at when a score moves, so you can tie a change to something you did.
  • Diagnosis alongside measurement. Presence is downstream of crawler access and page structure. A tool that checks whether GPTBot, ClaudeBot, and PerplexityBot can actually reach your pages tells you why you are absent, not just that you are.
  • Someone else's problem. Engines change their behavior and surfaces; keeping up is work you can pay to avoid.

Honest about the cost side: this is the only approach with a bill. GroundScore is $49 per site per month on Pro, and $299 a month on Agency for five sites with $15 per additional site. If you run one site, have no clients, and are not yet sure AI search matters in your category, that is a real expense against uncertain value, and the manual baseline is the better first move. Vendor selection is its own decision with its own criteria, which we cover in how to choose an AI visibility monitoring tool.

How the three approaches compare

Direct answer: Manual checking wins on cost and speed and loses on consistency. Referral tracking wins on proving business value and loses on coverage, since most AI answers never produce a click. Dedicated monitoring wins on trend and coverage and loses on price. Pick by which weakness you can least afford.

Laid out against the criteria that actually decide the purchase:

Criterion Ask the engines yourself Analytics and referrer detection Dedicated monitoring
What it measures Presence in a live answer Clicks that arrived from an AI tool Presence across a fixed question set
Cost Free Free Paid, from $49 per site per month
Effort 20 minutes per round, every round One-time setup Setup once, then scheduled
Coverage Whatever you remember to ask Only answers that produced a click Same questions, engines, and cadence
Shows a trend Not reliably Yes, for traffic Yes, for citations
Diagnoses causes No No Yes, if it checks crawler access
Cannot tell you Whether today was typical Whether you were cited at all Why one specific answer changed

Two things stand out when you read the table across rather than down.

First, the free options are complementary rather than redundant. Manual checking sees presence without outcomes. Referral detection sees outcomes without presence. Run both and you have covered the two ends of the funnel for nothing but your time, which is why the honest recommendation starts there.

Second, the paid option is not a different measurement. It is the same measurement performed reliably, plus the diagnostic layer that turns a result into a task. Whether that is worth paying for comes down to one question: is the answer going to change what you do next?

Which approach fits your situation

Direct answer: Start with a manual baseline and a free automated check, because both cost nothing and together they tell you whether you have a problem at all. Add referral tracking next. Move to paid monitoring only when the question changes from am I visible to is this getting better.

The sequence matters more than the choice, because each step tells you whether the next one is worth taking.

If you have never checked at all, do the manual round today. Ten questions, three engines, one spreadsheet. Then run a free AI visibility check to get the structural half: whether AI crawlers can reach your pages, whether your content is parseable, and where you land on the 0 to 100 scale. No account needed. Between the two you will know within an hour whether you have a visibility problem, a technical problem, or neither.

If you are already getting cited sometimes, turn on referral detection so you can see whether those citations produce anything. This is the cheapest way to find out whether AI visibility is a revenue channel in your category or a vanity metric.

If you are actively working on it, you have crossed the line where monitoring pays for itself. Once you are publishing, fixing schema, or unblocking crawlers, you need to know whether the work moved anything, and manual sampling is too noisy to answer that.

If you manage sites for clients, the calculus changes early. Manual checking across ten sites is not a twenty-minute task, and clients want a chart rather than your recollection.

One caution regardless of path: do not track everything. Five to ten questions that map to real buying intent beat fifty covering your whole category. A question set you actually read every week is worth more than a comprehensive one you ignore.

Four step sequence for tracking website visibility in AI search results

Frequently asked questions

Can I see my website's visibility in ChatGPT for free?

Yes. Open ChatGPT in a fresh session, ask five to ten questions your customers would ask, and record whether your site is mentioned, cited, or recommended. That costs nothing and takes about twenty minutes. Repeat it monthly with the same questions so the results stay comparable across rounds.

Why do I get different answers every time I check?

Because AI answers are generated, not retrieved from a fixed ranking. Model updates, live search results, session context, and sampling randomness all shift which sources appear. This is why one check proves very little, and why any serious measurement depends on repeated observations of the same question set over time.

Does ChatGPT show up in my analytics as a traffic source?

Sometimes. Clicks from chatgpt.com and similar hosts usually arrive with a referrer you can segment. But desktop apps, copied links, and stripped referrers push a share of that traffic into direct, so treat your AI referral count as a floor rather than a complete total.

Is analytics referral data enough on its own?

No, because it only captures visibility that produced a click. AI answers often satisfy the question outright, so you can be cited frequently and see very little referral traffic. Zero AI referrals is ambiguous on its own: it could mean no citations, or citations nobody clicked.

When is paid monitoring actually worth it?

When you are actively changing your site and need to know whether the changes worked. Manual sampling is too irregular to show a trend, so once the question shifts from whether you are visible to whether you are improving, scheduled monitoring against a fixed question set is what answers it.

How many questions should I track?

Five to ten is the right range for most single-site businesses. Pick questions that map to real buying intent rather than broad category terms. A small tracked set you review every week is more useful than an exhaustive list nobody reads, and it keeps the comparison across rounds clean.

The bottom line

Tracking your visibility in ChatGPT and AI search is not one method, it is three, and they answer different questions. Asking the engines yourself tells you whether you are in the answer. Referral detection tells you whether that produced anything. Scheduled monitoring tells you whether the line is moving. Start with the free two, because they cost nothing and will tell you whether there is anything here worth paying to watch.

If you want the structural half of the picture in the next few minutes, run a free AI visibility check and see where your site stands.

How visible is your site in AI search?

Check your AI visibility score in seconds — free, no account needed.

Check your score