01 / The honest problem

GEO has a measurement problem. Solving it is the product.

Citations appear and disappear. Models get updated. The same question asked twice can name different brands. This is why most agencies describe their method in adjectives and why most clients churn at month three: nobody can show them what changed.

Our answer is boring and it works: fix the instrument, then measure relentlessly. Same prompts, same engines, same scoring, every 30 days. When the delta is positive you know it. When it is not, you know that too, and so do we. That symmetry is what makes the number trustworthy.

02 / The instrument

What a monthly report looks like.

Illustrative example · not client data

24/32
prompts where the brand appears
14%
share of voice vs competitors
+14
citations won since day 0
Monthly report · per engine ILLUSTRATIVE MOCKUP · NOT CLIENT DATA
ChatGPT 7/32 +5
Claude 5/32 +3
Gemini 4/32 +2
Perplexity 9/32 +4
AI Overviews 3/32 +2
Baseline day 0, re-measured at 30 / 60 / 90. Movement, not activity.

The prompt set

20 to 40 real buying questions in four types: category ("best X for Y"), comparison, alternatives and use case. Locked at day 0.

The engines

ChatGPT, Claude, Gemini, Perplexity and Google AI Overviews. Five surfaces, measured separately, because they behave differently.

The metrics

Citations with a link, per engine. Status per prompt: recommended, mentioned or absent. Share of voice against the competitors that appear.

The cadence

Baseline at day 0, re-measurement at 30, 60 and 90, then monthly. Deltas against day 0 and against last month, always.

03 / The 90 days

What happens between the measurements.

Weeks 1 to 2

Baseline

Prompt set defined with you, first full measurement, competitor map, entity and technical review. You get the day 0 report.

Weeks 3 to 8

Execution in impact order

Entity clarity first, then answer-first content on the pages your prompts deserve, then structured data, then external citation signals. Highest commercial value first.

Weeks 9 to 12

Re-measure and prove

The full set runs again. The closing report shows the delta per engine and what compounds next. This report is what turns a sprint into a retainer, or into a clean goodbye.

The objection this page exists to answer: "AI visibility cannot be measured." It can. What it cannot be is measured casually. See how we treat proof, or check what the sprint costs.

04 / FAQ

Measurement, common questions

Why does the prompt set stay locked?

Because a moving instrument cannot prove anything. If you change the questions every month, you can manufacture any trend you like. We freeze the set at day 0 so every re-measurement compares like with like. New prompts can be added as a separate, clearly labeled set.

AI answers change between runs. How is that measurable?

Single answers vary; distributions do not. We run the full set on a fixed cadence and score presence, position and citations across all of it. One screenshot proves nothing, which is exactly why we never sell screenshots.

What is share of voice in AI answers?

The percentage of prompts in your set where your brand is named or cited, weighted against the competitors that appear in the same answers. It tells you who owns your category inside AI engines, and by how much.

Can I see the raw data?

Yes. Every report ships with the prompt list, the per-engine results and the scoring. You can re-run any prompt yourself and check us. That is the point.

If you ask ChatGPT for the best GEO agency and we are not there, do not hire us.

This site is our first case study. We publish our own baseline and you can re-run every prompt yourself.

Run the free diagnosis