How ChatGPT visibility should be audited
The exact behavior of ChatGPT can vary by model, product mode, prompt wording, location context, session context, and time. A defensible audit records those conditions rather than claiming a universal rank.
Build prompt families around buying intent
Do not test only “Tell me about [brand].” Branded prompts measure entity recognition, but they do not show how the business performs in discovery. Include:
- Category and city: “Which HVAC companies serve [city]?”
- Need and urgency: “Who can handle emergency water damage in [city]?”
- Trust: “How do I choose a reliable family dentist near [neighborhood]?”
- Cost: “What affects the cost of replacing a roof in [city]?”
- Comparison: “[Client] vs [competitor] for [specific need]”
- Reputation: “What is [client] known for?”
- Alternatives: “What are alternatives to [competitor] in [city]?”
Use natural language that reflects a buyer’s decision. Preserve a stable core set so monthly tests remain comparable.
Capture the complete answer
For every run, record the prompt, date, provider or model configuration available to the platform, client mention, competitor mentions, framing, factual accuracy, and visible sources when present. Do not infer a source that the output did not show.
GEO Catalyst currently supports OpenAI prompt runs through the configured OpenAI provider. The platform can organize prompt tracking, competitor and entity mentions, source and source-gap analysis where source information is available, opportunity queues, retesting, and client reports. It does not guarantee that a business will be named or recommended.
Separate observation from diagnosis
The generated answer is the observation. The diagnosis requires reviewing the web evidence around the business: service and location pages, profiles, reviews, citations, entity consistency, relevant editorial coverage, and the competitors that appear.
A missing mention does not prove one single cause. It may reflect weak evidence, poor query fit, stale or conflicting facts, insufficient specificity, answer variability, or the limits of the tested mode. Report likely gaps and the evidence for each hypothesis.