Nobody gets an A
25 real websites, scored on the rubric this document describes. The results are the reason the rest of it is worth reading.
The set
The 25 domains are a fixed benchmark, not a survey. They were chosen to be hard to argue with — cloudflare.com, github.com, mozilla.org, shopify.com, stripe.com, wikipedia.org and nineteen others, ranging from engineering-led companies to ordinary small-business sites.
This file is not marketing material. It is the frozen rubric output the regression suite asserts against: change the scoring and the test fails. Which makes it the one dataset here that cannot be quietly tuned to flatter us — the file you would have to edit to improve these numbers is the file that would break the build.
Two of the 25 are ours, and they are the top two scores. We have left them in and named them — securafy.com and securafyai.com — because removing them after seeing the result is exactly the sort of thing this document exists to argue against. Every figure below holds without them.
What came back
Not one of the 25 earned an A. The best managed a B. The median composite was 74, the range 52 to 91, and 9 landed below 70 — the threshold at which this tool tells an agency there is a real conversation to have.
These are not neglected websites. Several are maintained by companies whose entire business is the web. The point is not that they are bad; it is that "well built" and "scores well across eight categories" turn out to be different things, and the gap between them is where the work is.
If sites built by engineering-led companies land on a C, a small business scoring a C is not behind. It is average — and average is a position, not a verdict.
| Domain | Score | Weakest category |
|---|---|---|
| github.com | 74 C | AEO 40 |
| cloudflare.com | 77 C | AEO 50 |
| mozilla.org | 77 C | AEO 50 |
| wikipedia.org | 61 D | AEO 30 |
| shopify.com | 71 C | AEO 55 |
Where they lose it
The weakest category across the set is AEO, median 55, with 8 of 25 scoring under 50. That is the newest thing being measured here and the least attended to anywhere: whether an AI assistant can actually read, understand and cite the page.
The strongest is Best Practices, median 95, with 24 of 25 at 90 or better. It is worth 4% of the composite — the smallest weight of the eight.
That pairing is the whole argument for weighting. A category almost everyone passes cannot distinguish anyone, and a rubric that let it count equally would be measuring conformity rather than quality.