Skip to main content

Public-web benchmark · baseline edition

The State of AI Discoverability

This inaugural authoritative cross-sectional baseline covers 3001 unique domains. The overall median is 73, with the interquartile range spanning 64 to 80. Strong access foundations coexist with much thinner survival through understanding and answerability.

2026-Q3 publication · captured · Methodology

3,001

domains analyzed

73/100

median score

64-80

middle-half range

Model-written, statistic-bound commentary

Key findings

The analysis is computed in code. GPT-5.6 Sol turns only those supplied metrics into commentary, and every numeric statement is bound to a named metric before publication.

Median 73

The overall distribution is broad.

Scores range from 36 to 98, while the middle quartile range runs from 64 to 80.

n=3,001 sites ·

Mean 39.6

Content Answerability is a comparatively weak scored dimension.

Its median is 40; Content Freshness & Authority is also low, with a mean of 34.7 and median of 25.

n=3,001 sites ·

Advantage 61.6 percentage points

Retrievable, self-contained sections are the clearest top-quartile separator.

The pass rate is 93.3 percent in the top quartile and 31.7 percent across the rest of the sample. This is an association, not evidence of causality.

n=3,001 sites ·

Correlation 0.454

HTML extractability and site architecture show the strongest supplied section-level association.

Their Spearman rank correlation indicates that these dimensions tend to co-occur, without establishing causality.

n=3,001 sites ·

A cumulative five-stage funnel

Where AI discoverability breaks down

The funnel represents cumulative population survival, not a score. Reachable retains 2963 domains, or 98.7 percent of the full population, after 38 leave at that stage. Permitted retains 2313, or 77.1 percent, after a further 650 leave. Readable retains 1834, or 61.1 percent, after 479 leave. Understandable retains 928, or 30.9 percent, after 906 leave. Answerable retains 538, or 17.9 percent, after 390 leave.

These percentages describe population survival through the funnel. They are not section scores and are not added to the audit score.

Reachable

98.7%

Permitted

77.1%

Readable

61.1%

Understandable

30.9%

Answerable

17.9%

Reachable

Can an automated reader retrieve the page?

2,963 remain

Permitted

Are search and answer crawlers allowed to access it?

2,313 remain; 89 unknown

Readable

Does the useful content survive different readers and rendering modes?

1,834 remain

Understandable

Is the entity presented consistently across visible and machine-readable sources?

928 remain

Answerable

Is the content divided into retrievable, self-contained answer units?

538 remain

Overall score distribution

The shape of the field

The middle half of the sample sits between 64 and 80, around a median of 73. This edition is the first authoritative baseline; earlier, much smaller samples are retained but not interpreted as change over time.

36MIN64P2573MEDIAN80P7585P9098MAX

Eight comparable scored dimensions

Where sites win and lose

Content Freshness & Authority has the lowest average score; Bot Access & Control Plane has the highest. The bars show dispersion, not just the mean.

Content Freshness & Authority

avg 35

p25 25 · median 25 · p75 48 · p90 66

Freshness and authority have a mean of 34.7 and median of 25, placing this among the weakest scored dimensions in the baseline.

Content Answerability

avg 40

p25 0 · median 40 · p75 69 · p90 83

Answerability remains limited, with a mean of 39.6 and median of 40. Self-contained sections show a 61.6 percentage-point top-quartile pass-rate advantage.

Entity Clarity

avg 70

p25 56 · median 70 · p75 85 · p90 90

Entity Clarity records a mean of 69.9 and median of 70. Consistent entity naming carries a 25.9 percentage-point top-quartile pass-rate advantage.

HTML Extractability & Main Content Clarity

avg 77

p25 70 · median 82 · p75 88 · p90 92

Extractability and main-content clarity have a mean of 77 and median of 81.7. A canonical URL shows a 27.9 percentage-point top-quartile pass-rate advantage.

Trust & Security

avg 78

p25 70 · median 79 · p75 86 · p90 90

Trust and security record a mean of 78.1 and median of 78.5. About or Company links show a 28.5 percentage-point top-quartile pass-rate advantage.

Site Architecture & Coverage

avg 84

p25 80 · median 93 · p75 100 · p90 100

Site architecture is broadly strong, with a mean of 83.9 and median of 93.2. Navigation present in served and rendered HTML shows a 26.2 percentage-point top-quartile advantage.

Fetch, Render, and URL Integrity

avg 87

p25 81 · median 90 · p75 93 · p90 97

Fetch, render, and URL integrity are comparatively strong, with a mean of 86.6 and median of 89.7.

Bot Access & Control Plane

avg 94

p25 100 · median 100 · p75 100 · p90 100

Access controls are broadly strong, with a mean of 94.2 and median of 100.

Observational measure

Structured Data

Structured Data is reported for diagnostic context but carries no overall-score weight, so it is not ranked against the eight scored dimensions. It was evaluated for 44.4% of this sample; its observed median was 100.

Check-level pass-rate lift

What separates the top quartile

These are the checks with the largest pass-rate gaps between sites at or above the overall seventy-fifth percentile and the rest of the sample. The comparison is descriptive, not causal.

CheckTop quartileOthersLift
Content is divided into retrievable, self-contained sectionsContent Answerability93.3%31.7%+61.6 pp
Title, primary heading, and opening content share a clear topicContent Answerability51.4%9.7%+41.7 pp
Social profile links presentEntity Clarity84%48.2%+35.8 pp
Exactly one H1HTML Extractability & Main Content Clarity74.6%45.6%+29 pp
About/Company link presentTrust & Security76.8%48.3%+28.5 pp
Canonical URL presentHTML Extractability & Main Content Clarity81.6%53.6%+27.9 pp
Important navigation is present in served and rendered HTMLSite Architecture & Coverage92.1%65.9%+26.2 pp
Entity name is consistent across visible and machine-readable sourcesEntity Clarity65.4%39.5%+25.9 pp

Relationships inside the audit

Which signals move together

Spearman rho compares ranked section scores; phi compares pass versus non-pass outcomes for pairs of checks. Strong relationships can reveal shared implementation patterns or overlapping measurement, but they do not prove that one signal causes another.

Section-score correlations

Pairrhon
HTML Extractability & Main Content Clarity + Site Architecture & Coverage0.4543001
Site Architecture & Coverage + Trust & Security0.4433001
HTML Extractability & Main Content Clarity + Trust & Security0.4313001
Entity Clarity + HTML Extractability & Main Content Clarity0.3863001
Entity Clarity + Trust & Security0.3513001
Entity Clarity + Site Architecture & Coverage0.3323001
Content Answerability + HTML Extractability & Main Content Clarity0.3293001
Fetch, Render, and URL Integrity + Site Architecture & Coverage0.3233001

Check-pass associations

Pairphin
Viewport configured for mobile devices + Viewport meta present1.0003001
Sitemap available + Sitemap looks like XML (not HTML)0.9992306
Last-modified date present (schema or meta) + Publication date present (schema or meta)0.8793001
Canonical URL matches the current page URL + Canonical URL matches site origin0.6251838
Privacy policy link present + Terms link present0.6003001
Main content requires JavaScript that fetch-only AI crawlers do not run + Sufficient on-page text0.5743001
Canonical URL present + Meta description missing0.5663001
About/Company link present + Privacy policy link present0.5153001

Only groups with at least 30 domains

Differences by inferred industry

Industry labels are inferred from public site signals and should be read as directional. Small groups are suppressed rather than over-interpreted.

Inferred industryDomainsMedianMiddle half
SaaS & CloudThe inferred SaaS & Cloud group contains 318 domains and has a median overall score of 80.3188072-83
EducationThe inferred Education group contains 193 domains and has a median overall score of 77.1937772-82
Marketing & AdvertisingThe inferred Marketing & Advertising group contains 93 domains and has a median overall score of 77.937767-82
Healthcare407668-81
Finance & Banking977566-80
Non-Profit557570-81
Government917464-80
TechnologyThe inferred Technology group contains 870 domains and has a median overall score of 73.8707360-81
Travel & Hospitality527363-81
Media & Entertainment5007266-77
Telecommunications387264-77
Other697166-79
E-CommerceThe inferred E-Commerce group contains 118 domains and has a median overall score of 68.1186857-75
GamingThe inferred Gaming group contains 116 domains and has a median overall score of 68.1166859-78

A clean quarterly publication

Methodology

The sample includes active, public-index-eligible domains from the automated benchmark crawler. Each domain contributes its latest completed root-page baseline audit using the current scoring system inside a 365-day window. Nyman Media domains are excluded. This is not a self-selected user-audit sample.

The funnel is deterministic and cumulative. A domain only remains at a stage if it clears that stage and every earlier one. The benchmark crawl collects audit evidence only; it does not pull customer outcomes, product usage, or payment data.

Statistics are deterministic and generated before commentary. The commentary model cannot change the sample, scores, correlations, or tables. See the full scoring and population methodology.

  • This is a cross-sectional baseline and does not support claims about change over time.
  • Each domain was evaluated in a single automated run.
  • The population is limited to active, public-index-eligible domains in the benchmark crawler.
  • Industry labels are inferred; observations are limited to supplied credible industries meeting the stated sample threshold.
  • No outcome data are included, so associations and separator gaps do not establish causality.
  • Structured Data is observational and is not ranked as a scored dimension.