If you ask ChatGPT "who's the best plumber in Denver?" you get one list. If you ask "who's an affordable plumber in Denver?" you get a different list. Ask "who do I call for an emergency plumber in Denver?" and you get a third. Same city, same category, same buyer. Different rosters.
This isn't a quirk. It's the whole game. And it means something uncomfortable for anyone trying to measure AI visibility: your measurement is only as good as the prompts behind it. Change the wording and you change the answer, which means a sloppy prompt set gives you a sloppy number and no way to know it's wrong.
A quick test
I ran a small version of this myself across a few local-services categories. I kept the location and the category fixed and only swapped the qualifier — the adjective or the intent word. Three variants:
- "best [category] in [city]"
- "affordable [category] in [city]"
- "emergency [category] in [city]"
The named businesses moved every time. "Best" tended to surface the established, review-heavy names — the ones with long track records and lots of citations. "Affordable" pulled in different players, sometimes ones that explicitly market on price, sometimes ones I hadn't seen in the "best" list at all. "Emergency" reshuffled again, favoring businesses that advertise 24/7 availability and fast response.
The overlap between the three lists was partial at best. In a couple of categories, the top name in one variant didn't appear at all in another. If I'd only run "best," I would have walked away thinking a business had strong AI visibility. Run "affordable," and the same business was nowhere.
That's the problem in one sentence: which prompt you happen to pick determines what you conclude.
This is a survey-instrument problem
Anyone who's run a customer survey knows that question wording moves the results. "How satisfied are you?" and "What frustrated you today?" produce different data from the same person. Good researchers don't wing the wording — they standardize it, document it, and keep it stable so that this quarter's numbers can be compared to last quarter's.
AI visibility measurement is the same discipline, and most people aren't treating it that way yet. They type a question into ChatGPT, see whether their business shows up, and call it a reading. But a single ad-hoc prompt is one question from an unstandardized survey. It tells you something, but you can't trust it over time and you can't compare it to anything.
If you want a measurement that means something, the prompt set has to be treated as the instrument. That means:
-
Define the intents you care about. For a local-services business, "best," "affordable," and "emergency" are three genuinely different buyer situations. They may map to different jobs, different margins, different urgency. Decide which ones matter to the business and measure those deliberately — not whichever phrasing came to mind first.
-
Write the prompts down and freeze them. The exact wording, the location format, the category term. If you change the prompt, you've changed the instrument, and the before-and-after comparison breaks. Version your prompts the way you'd version anything else you rely on.
-
Run the full set, not one query. A single "best" prompt is a single data point. The picture only holds up when you run the whole intent set and look at the roster across all of them.
-
Repeat on a schedule. The value is in the trend, and a trend requires the same instrument every time.
Why this matters more than it sounds
Here's the practical stakes. Say you're an agency reporting AI visibility to a plumbing client. You run "best plumber in [city]," your client shows up second, and you report strong positioning. The client is happy. But their actual paid-search business is built on emergency calls — that's where the margin is. In the "emergency" roster, they don't appear at all. You've measured the wrong thing and reported good news about it.
Or the reverse: a business looks invisible under "best" because the category is dominated by legacy names, but shows up consistently under "affordable" and "emergency," which is exactly where its customers are. If you only ran one prompt, you'd either panic or celebrate for no good reason.
The named roster is real data. AI assistants are giving specific answers to specific questions, and those answers are shaping who gets considered. But the data is conditional on the question. A visibility number without a documented prompt set is like a survey result without the survey — you can't interpret it and you can't defend it.
What good measurement actually requires
Standardizing prompts isn't glamorous work, but it's the difference between a number you can act on and a number you made up. The businesses and agencies that get this right will treat their prompt set the way a good analyst treats a tracking survey: fixed wording, documented intents, consistent cadence, and comparison over time.
This is the part of AI visibility that's underappreciated right now. Everyone wants to talk about influencing what AI says. But you can't influence what you haven't measured correctly, and you can't measure correctly with a prompt you typed off the top of your head. Measurement comes first, and measurement done right starts with the instrument.
At LLMClarity this is exactly the problem we work on — tracking what ChatGPT, Claude, Gemini, Perplexity, and Google AI say about a business across a standardized, documented prompt set, over time, so the number means the same thing next quarter as it does today. But you don't need us to take the first step. Pick your real buyer intents. Write the prompts down. Run the whole set, not one query. Then run it again next month.
The word you put in the prompt decides who AI names. Once you've internalized that, you stop trusting any single reading — and you start building a measurement you can actually stand behind.
