How scoring works
A report sends every prompt to each provider you picked, repeating it according to your sample count. Every answer is then judged the same way: was your brand mentioned, and was it actually recommended? Was it the answer's single top pick? Did the answer cite your own website, or frame your brand negatively? And which competitors showed up?
Most of the rates below are simple shares of the report's scored answers: out of everything scored, how often did a given thing happen. The two "share of" metrics work differently: they compare you against your competitors, dividing by every brand appearance or recommendation instead of by answers. Either way, more scored answers (more prompts, providers, or samples) means steadier rates.
What each rate means
- Mention rate
- The share of scored answers that mention your brand at all. This is the baseline visibility question: do you show up?
- Recommendation rate
- The share of scored answers that actually recommend your brand. Being named in passing doesn't count.
- Primary pick rate
- The share of scored answers where your brand is the answer's single top pick.
- Citation rate
- The share of scored answers that cite your own website as a source. Citations are matched against the project's domains as they were when the report was created, so editing them later doesn't rewrite old reports.
- Negative sentiment
- The share of scored answers that frame your brand negatively. The one rate where lower is better.
- Share of voice
- Take every brand appearance in the answers, yours plus your tracked competitors'. Share of voice is the slice of those appearances that's yours. It's blank when no brand appears at all.
- Share of recommendations
- Of all the recommendations the answers handed out, the slice that went to your brand. Unlike share of voice, being criticized by name doesn't help here; only real recommendations count.
Why numbers move between reports
AI assistants rarely answer the same question the same way twice, so some movement between reports is ordinary noise. That's especially true for small runs with only a few scored answers.
The fix is repeat sampling: more samples per prompt means each rate rests on more answers, so real shifts stand out from the randomness. Keep in mind that adding, removing, or editing prompts also moves the aggregate rates all by itself, simply because the report is now measuring a different set of questions.
How prompt changes affect trends
Editing, adding, or deleting a prompt never changes an existing report. Each report captures the prompt wording at the moment it starts: completed reports keep their answers and chart point, and runs already in progress keep using the wording they launched with.
Only the next report picks up your change. It runs with the current prompt set and adds a new point to the same trend line. And because the charts compare whole-report rates, a jump or dip at that point can come from the prompt change itself, not just a real shift in AI visibility.
Between reports, prompts are matched by their exact wording. That means an edited prompt counts as one prompt removed and a new one added; the old and new versions are never compared directly. If that moves the numbers significantly, the report's “What changed” section says the prompt set changed, but the trend charts themselves don't reset or flag it. So for the cleanest trend, keep your prompts stable. After a deliberate change, treat the first report on the new set as the baseline for what follows.
When a change is significant
Each report is compared against the project's previous scored report. A shift is only flagged on the report page and in alert emails when it clears two bars: it moves at least 15 percentage points, and it's big enough relative to how many answers both runs scored that run-to-run randomness is an unlikely explanation.
Smaller moves still show up as the "vs previous" deltas beside each metric. They're just not treated as news.