The first benchmark measured to the Standard.
The Gr8 Index is the public, periodic benchmark of how often South Africa's AI assistants recommend the brands in a category — and it is the first such benchmark anywhere to be run to the GR8ER2 AI Visibility Standard: a published Prompt Corpus, a disclosed brand universe, at least eight runs per question, every engine labelled, and a confidence band on every number. This document is the method of record. It is written so a stranger could reproduce the Index — or challenge it — line by line.
Built on: The GR8ER2 AI Visibility Standard v1.1 — every rule here instantiates that specification; where the Index adapts a rule for a multi-brand benchmark, the adaptation is stated and justified.
Evidence: every empirical figure carries a GR8ER2 Evidence Bank ID and is Verified against its primary source. No brand, rank or figure in a published Index is ever invented (M-6).
This method, and the Gr8 Index it governs, describe an independent analysis of information generated by publicly accessible AI assistants. Figures record how often each brand was named by the assistants on the dates stated — a point-in-time measurement of visibility that varies with time, phrasing, location and personalisation. They are prepared from publicly available information, their sources are identified wherever practicable, and every figure is reconstructable from logs retained under this method. They are not GR8ER2's assessment of any institution's products, services, financial standing or conduct, are not a ranking of quality, and are not financial, legal or purchasing advice. Any analysis or commentary is GR8ER2's interpretation of the measured information. Brand and institution names are used only to identify the entities measured; no affiliation or endorsement is implied.
What the Gr8 Index is — and how it relates to the Score and the Standard
Three terms, cleanly separated, so nothing is conflated downstream.
The specification: how any AI-visibility number must be defined, measured and disclosed to count as a finding. The Index obeys it; it does not amend it.
A single 0–100 number for one brand, against one Prompt Corpus, on a stated engine panel: 0.40·Mention Rate + 0.35·position-weighted Share of Voice + 0.25·Citation Rate, always shipped with its confidence band.
The published benchmark: every brand in a defined category's brand universe, each given its Gr8 Score under one frozen Corpus and one panel, ranked into a league table, for a defined market and cycle. The Index is what the market reads; the Score is the unit it is built from.
Why this Index exists, and why now. The only South African AI-visibility benchmark published to date runs three repetitions per question — below the independent academic floor of seven to eight EB-035. Its findings are not false; their confidence is simply wider than their headline admits. The Gr8 Index is the first South African benchmark built to meet that floor: eight runs per question, four engines, a published corpus and a disclosed brand universe. We are not claiming to beat anyone's number. We are publishing the first one that can be checked.
The Standard requires GR8ER2 to measure its own Gr8 Score every cycle, to the same rules as any brand — including when it is bad. The Index honours this literally: GR8ER2's own category (AI-visibility and digital-marketing agencies) is category one, and GR8ER2 is scored inside it by the same engine-derived rule as every competitor. If we fall below the inclusion threshold, we publish that too. A benchmark whose author exempts itself is not a benchmark.
Scope and engine panel (v1)
The four categories — South Africa
Four categories are measured in the first cycle, chosen for finite, recognisable brand universes, high stakes in an AI recommendation, and genuine buyer intent to ask an assistant. Each is measured identically.
| # | Category | Why it is in v1 |
|---|---|---|
| C1 | AI-visibility / digital-marketing / SEO agencies | GR8ER2's own category. Scored to the same rule as all others — the Standard's self-measurement commitment, made literal. |
| C2 | Short-term insurance (car & household) | A finite, instantly recognisable SA brand set; a considered purchase buyers increasingly research through AI. |
| C3 | Retail banking | Iconic, bounded SA brand set; everyday relevance and high switching stakes. |
| C4 | Medical aid / health insurance | A high-consideration purchase with a well-defined SA scheme set and strong intent to seek recommendations. |
The engine panel — four Core direct-API engines
v1 measures the four Core direct-API engines of the Standard, each queried through its official API with web search enabled where supported, each computed and reported separately (M-4), never blended silently.
| Engine | Channel | Surface measured |
|---|---|---|
| ChatGPT | OpenAI API | Assistant answer with web search enabled |
| Claude | Anthropic API | Assistant answer with web search enabled |
| Gemini | Google Generative Language API | Assistant answer — not Google AI Overviews / AI Mode (M-5) |
| Perplexity | Perplexity API | Answer with live retrieval |
Four engines with eight runs each clears the Standard's High confidence floor (≥3 engines and ≥8 runs). The panel, and any engine that fails or is excluded mid-cycle, is always named with the reason (R-3 graceful degradation): if Perplexity is unavailable, the Index runs and publishes on the remaining three — still High — and says so.
Which model, on each engine
Each engine is queried at the model tier a logged-out, everyday user is served by default — not the top-priced flagship — with web search enabled. The Index measures what an ordinary South African buyer actually sees, and an anonymous consumer is not served the most expensive model; choosing a representative tier is therefore a fidelity decision first and a cost decision second. The exact model identifier for every engine is logged per answer (M-8), published with the Index, and frozen for the cycle; changing a model is a version event (M-9), never silent.
What is deliberately held for v1.1, and disclosed, not hidden. Google's consumer AI — AI Overviews and AI Mode — is the surface most South Africans actually see, but it has no official API and must be read through a licensed SERP-data channel with ~1-minute token expiry and per-prompt "no surface served" logging (Standard M-7, Section 3 Google rules). Wiring it correctly is a few hours of build-and-test; rushing it would put an unvalidated number on our most visible surface — the one thing the Standard forbids. It is therefore added as a dated v1.1 panel expansion, published transparently as a panel change (M-4), never backfilled silently into a v1 number. The Standard's versioning rules exist precisely for this.
The brand universe — engine-derived, by published rule
The brand set is the denominator of Share of Voice. Whoever picks it decides who can win. So we do not pick it: the engines reveal the field, a published rule decides inclusion, and the whole set — plus the brands that fell below the line — is disclosed.
This is the single biggest methodological risk in any benchmark. A curated list smuggles the author's judgement into the result. The Index removes that judgement from the result and replaces it with a rule anyone can re-run.
The derivation, step by step
Run the nine brand-neutral Discovery prompts for the category at full protocol — 9 prompts × 8 runs × 4 engines = 288 qualifying answers per category, SA-pinned and logged-out (Section 4). Extract every brand/entity named, with its answer position, by named-entity recognition and a human review pass.
A brand enters the frozen tracked set if it is named in ≥5% of the category's Discovery answers (≥15 of 288) and appears on ≥2 of the 4 engines. The percentage floor removes one-off hallucinations; the two-engine floor removes single-engine artifacts. Both thresholds are fixed for v1 and printed with the Index.
Every brand named at least once but under threshold is listed in a "tail" appendix with its mention count. We never quietly drop a brand; a reader sees exactly where the line fell and who sat just under it. This is the discipline that proves the set was not curated.
Once derived, the set is locked for the cycle (C-2), versioned, and never silently edited (C-3). A brand that crosses the threshold next cycle joins then, annotated — not retroactively.
Entity resolution — the South African ambiguity problem
SA brand names collide across sectors, and a benchmark that miscounts them is wrong at the source. The resolution rules are fixed and published:
- Aliases are canonicalised: "FNB" → First National Bank; "Std Bank" → Standard Bank; "Absa", "Nedbank", "Capitec", "TymeBank" to their canonical entities. The canonical-name map is published with the Index.
- Multi-sector groups are resolved by category context. "Discovery" resolves to Discovery Bank in C3, Discovery Health in C4, and Discovery Insure in C2 — never pooled across categories. The same applies to Momentum, Old Mutual and Sanlam, which span life, health and short-term lines. A mention that cannot be resolved to the category's entity from its answer context is logged and adjudicated by a stated rule, and the adjudication log is kept.
- Company, not product: a product or plan name maps to its parent company entity.
- Global vs local: a global brand is included only if it genuinely serves South African buyers in that category and is named in SA-pinned answers.
GR8ER2's own place in C1, handled honestly. GR8ER2 is subject to B-2 exactly like every competitor. If it meets the threshold, it is ranked. If it does not, the Index states plainly that GR8ER2 is below the inclusion threshold in its own category this cycle, and reports GR8ER2's Gr8 Score from its branded prompts separately. For a company that sells visibility, publishing "we are not yet visible, measured to our own rule" is not a weakness — it is the single most credible thing in the document.
The Prompt Corpus
The Standard's reference corpus is twenty prompts (9 Discovery / 5 Comparison / 3 Problem / 3 Branded) for a single-brand audit. A benchmark measures many brands at once, so the Index adapts the shape while preserving the intent proportions — and the adaptation is documented, as the Standard's Section 2 requires.
The Index corpus structure (the documented adaptation)
| Tier | Count | Form | Feeds |
|---|---|---|---|
| Discovery | 9 | Brand-neutral, shared by the whole category | League table + the brand harvest (B-1) |
| Comparison | 5 | Category-level; two are leader-anchored, filled post-harvest | League table |
| Problem | 3 | Brand-neutral, shared by the whole category | League table |
| Branded | 3 | Templated and instantiated per brand in the set | Each brand's own Gr8 Score (not the ranking) |
The 17 shared, brand-neutral prompts (9+5+3) are the comparable, non-self-referential measure of who AI recommends in the open field — these drive the league table. The 3 Branded prompts are instantiated for each brand (e.g. "what is {brand}") and feed only that brand's complete Gr8 Score, never the ranking, because a brand being asked about itself cannot be a fair comparator against the field. Per brand this is still 9/5/3/3 = twenty prompts; the intent proportions are preserved exactly (Standard §2). Prompts are weighted flat within each tier — no search-volume weighting, for the reason EB-038 makes plain: volume weighting re-hides the very choice the Corpus exists to expose. The two leader-anchored Comparison prompts are instantiated once per category from the top two discovered brands, run as shared inputs to every brand's Share of Voice, and printed in full in the published Index.
v1 scope — branded prompts. Because the league table is driven entirely by the 17 shared, open-field prompts, the 3 Branded prompts are run in v1 for GR8ER2 only — satisfying the self-measurement commitment — while the shared-corpus Gr8 Score and the full league table are published for every brand in the universe. Branded-intent "complete" scores for the whole set are a dated extension in a later cycle. This is a disclosed scope decision, not a change to any definition: no brand's ranking depends on a branded prompt.
The four rules of a conformant Corpus (inherited, enforced)
C-1 Published — every prompt ships with the Index. C-2 Frozen per cycle — locked once the cycle begins. C-3 Versioned, never silently edited — superseded prompts marked and kept. C-4 Brand-neutral in Discovery and Problem — only Branded prompts name a brand.
The frozen v1 prompts for each category follow. Placeholders in braces are filled post-harvest and printed in the published Index.
C1 · AI-visibility / digital-marketing / SEO agencies
Discovery
1 best AI visibility agency in South Africa
2 who can help my business show up in ChatGPT and AI search in South Africa
3 top GEO or AEO agencies in South Africa
4 best digital marketing agency for AI search optimisation in South Africa
5 which South African agency should I hire to get my brand recommended by AI assistants
6 best SEO agency in South Africa 2026
7 leading generative engine optimisation specialists in South Africa
8 recommend an agency in South Africa to improve how often AI recommends my company
9 best AI search marketing companies in Cape Town or Johannesburg
Comparison
10 compare the top AI visibility agencies in South Africa
11 alternatives to {LEADER_1} for AI visibility in South Africa
12 {LEADER_1} vs {LEADER_2} for AI search optimisation
13 best alternatives to a traditional SEO agency for AI search in South Africa
14 which AI visibility consultancy in South Africa is the most reputable, and why
Problem
15 my business is not showing up when people ask ChatGPT for recommendations — who in South Africa can fix it
16 how much does it cost to improve AI search visibility in South Africa, and who offers it
17 we are losing leads to competitors that AI recommends — which South African agency handles this
Branded (per brand {B})
18 what is {B}
19 is {B} a good AI visibility or digital marketing agency
20 {B} reviews and reputation
C2 · Short-term insurance (car & household)
Discovery
1 best short-term insurance in South Africa
2 best car insurance in South Africa
3 best home or household insurance in South Africa
4 most recommended short-term insurer in South Africa 2026
5 best value car and home insurance in South Africa
6 which short-term insurance company should I choose in South Africa
7 top-rated insurers for car insurance in South Africa
8 best insurance company for claims in South Africa
9 affordable car insurance in South Africa — which company is best
Comparison
10 compare the best short-term insurers in South Africa
11 alternatives to {LEADER_1} for car insurance in South Africa
12 {LEADER_1} vs {LEADER_2} car insurance — which is better
13 best alternatives to the biggest insurers in South Africa
14 which South African insurer offers the best cover for the price
Problem
15 my car insurance premium is too high in South Africa — which insurer is cheaper
16 I need to insure my car and house in South Africa — which company is best
17 which insurer in South Africa is best for quick claims payouts
Branded (per brand {B})
18 what is {B}
19 is {B} a good insurance company in South Africa
20 {B} insurance reviews and claims reputation
C3 · Retail banking
Discovery
1 best bank in South Africa
2 which bank should I open an account with in South Africa
3 best bank account for low fees in South Africa
4 most recommended bank in South Africa 2026
5 best digital or online bank in South Africa
6 best bank for students or young professionals in South Africa
7 which South African bank has the best app and online banking
8 best bank for a small business in South Africa
9 safest and most reliable bank in South Africa
Comparison
10 compare the best banks in South Africa
11 alternatives to {LEADER_1} in South Africa
12 {LEADER_1} vs {LEADER_2} — which bank is better in South Africa
13 best alternatives to the big banks in South Africa
14 which South African bank offers the best value for money
Problem
15 I am paying too much in bank fees in South Africa — which bank is cheaper
16 I want to switch banks in South Africa — which one should I move to
17 which bank in South Africa is best for everyday banking
Branded (per brand {B})
18 what is {B}
19 is {B} a good bank in South Africa
20 {B} bank reviews and fees
C4 · Medical aid / health insurance
Discovery
1 best medical aid in South Africa
2 which medical aid should I join in South Africa
3 best value medical aid scheme in South Africa
4 most recommended medical aid in South Africa 2026
5 best medical aid for families in South Africa
6 affordable medical aid options in South Africa — which is best
7 best medical aid for young or single people in South Africa
8 which medical aid has the best hospital cover in South Africa
9 top-rated medical aid schemes in South Africa
Comparison
10 compare the best medical aids in South Africa
11 alternatives to {LEADER_1} medical aid in South Africa
12 {LEADER_1} vs {LEADER_2} medical aid — which is better
13 best alternatives to the biggest medical aids in South Africa
14 which medical aid in South Africa offers the best cover for the price
Problem
15 my medical aid is too expensive in South Africa — which scheme is cheaper
16 I need medical aid for my family in South Africa — which is best
17 which medical aid in South Africa is best for chronic illness cover
Branded (per brand {B})
18 what is {B}
19 is {B} a good medical aid in South Africa
20 {B} medical aid reviews and complaints
Corpus scope note (v1): prompts are English-only. South African buyers in these categories query AI predominantly in English, but Afrikaans and other official languages carry real volume in insurance, banking and medical aid. English-only is a stated v1 limitation, logged in the blind-spot register (Section 6) and scheduled for a future multilingual expansion — never implied to be the whole picture.
The measurement protocol, instantiated
Every rule of the Standard's protocol maps to a concrete Index procedure. None is skipped.
Each prompt is run eight times per engine per cycle — the independent floor EB-035. 17 shared prompts × 8 runs × 4 engines = 544 shared answers per category, plus 3 branded × 8 × 4 per included brand.
The unit is "named in 6 of 8 runs", never "named". Each brand's per-prompt hit rate and its spread are retained and reported.
On branded prompts, a name echoed while the engine states it cannot find the brand ("I couldn't find {B}") is scored as an absence, not a mention. Applied by the GR8ER2 engine; critical for a fair read on low-visibility brands, including GR8ER2 itself.
Scores are computed per engine, then combined with the panel named. Every league table shows the per-engine breakdown beneath the composite.
The Gemini API result is labelled as such and never implied to be Google AI Overviews or AI Mode. Those consumer surfaces enter only at v1.1 via the licensed SERP channel, separately labelled.
Every brand, rank and figure traces to a logged run. Non-appearances and zeros are published, not hidden. A brand named zero times is reported as zero.
Every run is executed logged-out, in a fresh session, with the SA market pinned and the timestamp, sampling parameters and run conditions logged. Honest constraint: the Core APIs do not all honour consumer-style geolocation. The Index pins the market three ways and logs which held per engine: (a) explicit "in South Africa" framing in every prompt; (b) any region/locale parameter the API exposes; (c) web-search region settings where available. Where true geo-pinning is unavailable on an engine, that is recorded as a condition, not silently assumed — and it is a first-class reason the Google consumer surface (which does geo-pin) matters for v1.1.
Each run records the exact model identifier and version. A provider model change between cycles is annotated on the trend line, never absorbed as if it were a change in a brand's visibility.
Cycle-to-cycle movement is reportable only across byte-identical protocol — same corpus, engines, run count, conditions — with both run counts printed. The v1 → v1.1 panel change (adding Google + confirming Perplexity) breaks and restates the series rather than comparing across it. No manufactured movement.
When an engine answers without retrieving the web, that is recorded; the per-engine no-retrieval rate is published as a first-class metric EB-029. Citation Rate is computed over all answers, not only those that cited — a conditional rate may not be called a citation rate. Mention and citation converge at different speeds; the eight-run floor is set by the slower one (citation).
Scoring and the league table
The three components, per brand, per engine
Mention Rate (MR) — share of qualifying shared answers in which the brand is genuinely named (after the Branded-Echo Guard). Share of Voice (SoV) — the brand's share of all tracked-brand attention, position-weighted: within each answer, distinct brands are ordered by first mention and weighted harmonically (position 1 = 1.0, 2 = 0.5, 3 ≈ 0.33, …); a brand's SoV is its summed weight over the summed weight of all tracked brands, averaged across answers, runs and engines. Citation Rate (CR) — share of answers citing the brand's own domain as a linked source, over all answers (M-10).
// each component normalised to 0–100 before weighting; per engine, then averaged across the panel; composite to the nearest whole number
Components are computed per engine, then combined by an equal-weighted mean across the panel — engines are not weighted against each other, because there is no defensible basis to say one surface "counts more", and the per-engine figures are always shown so any reader may re-weight. The weighting of the three components (40/35/25) is the Standard's published editorial judgement, frozen per version, not a claim of optimality. Changing it is a version change, never silent.
Citation Rate counts the source, not the sentence. A brand's own domain cited as a linked source counts toward its Citation Rate even when the brand is not named in the prose of that answer. The two components deliberately measure different things: Mention Rate and Share of Voice measure being named in the answer text, and are computed over the prose with URLs excluded — a brand name appearing inside a cited URL slug (“…bank-zero-beats-tymebank-and-discovery-bank…”) is not a mention and may not set a first-mention position for the harmonic SoV weighting. Citation Rate measures being pointed to, and is read from the answer's citation list. A brand can therefore legitimately post CR > 0 with MR = 0: the engine used it as a source without naming it. That is a real and reportable state — influence without attribution — and the Index states it rather than treating it as an inconsistency to be smoothed away.
Where no citation of a brand's own domain appears anywhere in the logged corpus, its Citation Rate is reported as N/A — not measurable, never as a measured zero, and no domain is inferred (M-6). A brand's domain is attached only when it is the most-cited host matching that brand's name among answers naming it; a host can belong to only one brand per category.
Sentiment is reported, not baked in. Whether a mention recommends, neutrally lists, or flags a brand is recorded by a stated LLM-judgement scheme and shown beside the Score — but kept out of the composite in v1, because folding a model judgement into the headline would trade away reproducibility. Readers see both the quantity and the quality of presence.
Confidence band and interval — on every brand
With four engines and eight runs, every brand's Score qualifies for the High band. Each Score ships with a confidence interval computed from the observed run-level variance, sense-checked against the study's benchmark standard error of 0.062 (95% interval ≈ ±0.121) at eight runs EB-035. A brand whose interval overlaps the next brand's is reported as a statistical tie, never a false-precision rank — the Index shows the overlap rather than inventing a gap it cannot support.
| Band | Requires | v1 status |
|---|---|---|
| High | ≥3 engines and ≥8 runs | Met (4 engines × 8 runs) |
| Indicative | ≥3 engines, 2–7 runs | Fallback only if runs are lost |
| Snapshot | Single run or <3 engines | Not published as an Index |
What the Index publishes per category
A ranked league table of every included brand's Gr8 Score with its band and interval; the per-engine breakdown beneath each; the sentiment split; the no-retrieval rate per engine; and, where two brands tie within their intervals, the tie shown as such.
Blind-spot register — every known failure mode, and its control
A benchmark is only as trustworthy as its list of ways to be wrong. Each is named, with the control that contains it. This register is published with the Index.
| # | Blind spot | Control in this method |
|---|---|---|
| 1 | Brand-set bias — the denominator decides who can win. | Engine-derived inclusion rule B-1–B-4; thresholds and the below-threshold tail published. |
| 2 | Self-measurement conflict — we are in C1. | Same rule applies to GR8ER2; below-threshold result published if it occurs; conflict disclosed; logged-out neutrality + M-3 guard. |
| 3 | Branded-echo inflation on low-visibility brands. | Branded-Echo Guard (M-3); branded prompts excluded from the ranking entirely. |
| 4 | Entity ambiguity — "Discovery", "Momentum", "Old Mutual" span sectors. | Category-context resolution + published canonical map + adjudication log (Section 2). |
| 5 | Non-retrieval — answers with no web search. | No-retrieval rate published per engine (M-10); CR computed over all answers. |
| 6 | Model drift mid-series. | Model IDs logged; drift annotated, never absorbed (M-8). |
| 7 | Geo-pinning limits on Core APIs. | Three-way pinning, per-engine success logged (M-7); a stated reason for the v1.1 Google surface. |
| 8 | Engine API migration / router risk — Perplexity retired its Sonar chat endpoint; its new Agent API also routes to OpenAI/Anthropic/Google models. | Adapter rebuilt on the current endpoint and pinned to perplexity/sonar (Perplexity's own answer model) so the panel never silently re-measures another engine (M-4). R-3 graceful degradation if any engine drops. Resolved 1 Oct 2026. |
| 9 | Google consumer surface absent in v1. | Disclosed as a v1.1 panel expansion; series broken and restated, not backfilled (M-9). |
| 10 | Language — English-only v1. | Stated scope limit; multilingual expansion scheduled; never implied to be complete. |
| 11 | Prompt-phrasing bias. | Full corpus published (C-1); flat intra-tier weighting; no volume weighting EB-038. |
| 12 | Day/temporal effects — sources churn daily EB-032. | Eight runs; run window and timestamps logged; consecutive-day overlap (~35% EB-034) acknowledged. |
| 13 | False precision in ranking. | Overlapping intervals reported as statistical ties, not ranked gaps. |
| 14 | Reputational / defamation risk — publishing that AI speaks poorly of a named brand. | Every statement attributed to the engine and conditions that produced it ("ChatGPT, SA, logged-out, {date}…"); sentiment reported neutrally and never editorialised; figures traceable (M-6). Legal review of the first public Index before release. |
| 15 | "Measured gap" against the 3-run benchmark. | We never treat a difference against an external, differently-run benchmark as a measured gap; cross-protocol comparisons are refused (M-9). |
| 16 | Citation provenance differs per engine — Gemini returns grounding-redirect URLs (vertexaisearch…/grounding-api-redirect/…) that carry no publisher identity. | The real publisher domain is recovered from the grounding metadata (web.title) and scored; the redirect URI is retained alongside for audit provenance. Prevents a false-zero Citation Rate masquerading as a measured zero (M-6). Verified on known-cited brands. |
Publication, cadence and governance
What ships with every published Index
- The ranked league table per category, each Score with its confidence band and interval, and ties shown as ties.
- The per-engine breakdown, the sentiment split, and the no-retrieval rate per engine.
- The full frozen Prompt Corpus, including the post-harvest leader instantiations (C-1).
- The frozen brand universe, the inclusion rule and thresholds, the below-threshold tail, and the canonical-name / adjudication notes.
- The engine panel, model IDs, run dates and conditions — and any engine lost or excluded, with the reason.
- The method version and a link to this document.
Cadence
Recommended: a quarterly public Index (stable enough to be newsworthy, frequent enough to show movement), with optional monthly internal runs for the ROI Lab. Each public cycle is a version; movement between cycles is reported only under M-9.
Home and governance
Proposed canonical home: gr8er2.com/index, a sibling to /standard, each Index cycle a dated, permanent entry. Governance follows the Standard's Section 8: the method is versioned, superseded elements are marked and kept, and a change to any definition, threshold, weight or panel is a version bump with a change log — never a silent edit. This document is the v1 method of record; it is itself under that rule.
Execution runbook & today's critical path
What has to happen, in order, to run the first cycle — and who does each part. The method above is complete; this is the operational path to a published number.
Dependencies to clear first
OpenAI, Anthropic, Google Generative Language, Perplexity — each with web search / retrieval enabled where offered. Deaven: confirm keys exist and have credit.
Perplexity resolved (endpoint migration, not billing; pinned to perplexity/sonar). Gemini: Google AI Studio now issues only AQ. authorization keys (the AIza format is deprecated) — the key type is correct. AQ. keys must be sent via the x-goog-api-key header (not a ?key= query param) with a current SDK/endpoint, or they fail auth; confirm the key's project has quota/billing enabled for the 429. Agent: fix the adapter's auth method. Fix before Gate A to re-derive the universe free from cache and keep the 4-engine panel; else run on three (still High), disclosed.
Extend the existing GR8ER2 engine (which already implements the Branded-Echo Guard) to run the Index: batch prompts × runs × engines, log conditions and model IDs, extract entities, score. Agent: confirm current engine capabilities before building new.
Run order
Run the 9 Discovery prompts per category at full protocol; extract entities; apply B-2; resolve entities (Section 2); freeze the set + tail; publish the rule. Human review of the entity list before freeze.
Fill {LEADER_1}/{LEADER_2} from the top two discovered brands per category; instantiate the 3 Branded prompts for each included brand; freeze the full corpus (C-2).
All shared + branded prompts × 8 runs × 4 engines, logged-out, SA-pinned, conditions + model IDs + timestamps logged. Record non-retrieval per answer.
Apply M-3 guard; compute MR, SoV (harmonic position-weighted), CR per engine; normalise; composite 40/35/25; average across panel; compute intervals; flag ties; code sentiment.
Human audit of a random sample of answers against the automated scoring; branded-echo spot-check; entity-resolution adjudication review; confirm no unsourced figure (M-6). Legal review of the public Index (blind spot 14) before release.
Build the Index report with everything in Section 7; version it; stage at /index.
Data schema (minimum fields to log per answer)
category · prompt_id · intent · run_no · engine · model_id · timestamp · region_pin_method · region_pin_confirmed · retrieved_web(bool) · raw_answer · brands_named[] · brand_positions[] · citations[] · sentiment_per_brand · echo_guard_flag
Division of labour. The method of record (this document), the corpus and the rules are complete and need no further input. What remains is operational: Deaven clears D-1/D-2; the dev agent confirms the engine (D-3) and executes the run order. The first published number depends on the measurement run finishing and passing QA — do not publish a partial run as an Index.
