ABOUT
Cited measures how often AI answer engines recommend a specific business when buyers ask the kinds of questions that lead to a purchase — and which competitors get named in its place. This case study covers a baseline audit run for a Denver-area residential HVAC company, tracked under the name Northline HVAC. Twenty unbranded buyer questions were put to two AI engines repeatedly, and the resulting forty answers were scored for whether the subject appeared at all.

THE PROBLEM
Buyers increasingly ask an AI engine rather than a search engine which company to hire. When they do, they get a short list of recommended names — and a business either appears on that list or it does not.
Businesses have no visibility into this. Traditional rank tracking measures position in a results page; it says nothing about whether a model names you in a paragraph of prose. The informal version — asking a chatbot about yourself once — is worse than useless, because these models do not give the same answer twice, and a question containing your own name cannot measure whether you would have been mentioned unprompted.
What was needed was a measurement: repeatable, comparable over time, and honest about what it does and does not cover.

THE SOLUTION
We built the audit as a measurement method rather than a one-off report, with the design choices that make a figure meaningful built into the process.
A frozen question set. Twenty unbranded buyer questions were drafted from the subject’s service description and market, spanning five intent families — category, service, constraint, urgency and locality — then frozen, so this audit and any future rerun use an identical set. That is what makes a before-and-after comparison a real change rather than an artefact of asking differently. None of the questions mention the subject’s name, because a mention is exactly what is being measured.
Repetition, because models vary. Each question was sent multiple times per engine rather than once, producing 40 answers across 20 questions and two engines
Detection that survives real-world naming. A mention counts after normalising case, plural forms and legal suffixes such as LLC and Inc — not a literal string match — and the same normalisation is applied to the subject and to every competitor equally.
Competitors observed, not assumed. The first pass returned 0% for every tracked name, including all three competitors. Rather than reporting that as a finding, we read what the engines actually volunteered and discovered the tracked set was an assumption that did not match the market. Re-running the same frozen questions against the businesses the engines genuinely named produced a different competitor set — All Seasons Heating & Cooling and Denver Heating & Cooling, neither on the original list. A competitor list is a hypothesis; what the engines repeatedly name is the observation.
Stated scope. The reading covers two open-weight engines on one frozen question set in one vertical and city. It is not a claim about other engines or markets, and the report says so plainly.

THE OUTCOMES
The audit produced a clear, defensible baseline: across all 40 answers the subject was named zero times, for 0% share of voice and no citations of its domain as a source, while two competitors were recommended repeatedly. A question-by-question breakdown shows exactly where a competitor was recommended and the subject was not.
Equally valuable was the methodological finding. An all-zero scoreboard across every tracked name is itself informative — it usually signals that the tracked set does not match who is actually being recommended, rather than that the market has no visible players. Reading the engines’ own answers instead of the operator’s assumptions is part of the method.
Because the question set is frozen, a later rerun is directly comparable, so any change in discoverability shows up as movement in the same figure rather than a new baseline.
The project demonstrates our ability to build measurement systems around AI models — treating their variability, their naming behaviour and the limits of a sample as engineering problems to handle explicitly rather than gloss over.
