Read this before interpreting the ledger
The authoritative hash-chained ledger classified 2 rows runner-completed, 6 partial, and 2 errors. Both runner-completed rows contain method-invalid console passes: the collector was unsupported and absence of surfaced errors was treated as evidence. Therefore this report publishes no scores, pass rates, rankings, or completed-quality claims. “Runner-completed” means only that the ledger classified the row completed.
Ledger disposition
No target was replaced or retried. GitHub is correctly reported as missing-report; permit expiry is not claimed because the retained timestamps disprove it. Facebook is an error because the runner invalidated completed cross-origin journey states.
Ten-row inventory
| Origin | Archetype | Disposition / reason | Coverage | Journey |
|---|---|---|---|---|
1. https://github.com | documentation | errormissing-report | 0 judged; 0 blocked; 0 not run; 58 missing | partial |
2. https://web.facebook.com | blocked-complex | errorrunner-error | 0 judged; 0 blocked; 0 not run; 58 missing | invalid |
3. https://www.google.com | search-portal | runner-completednot-applicable | 58 judged; 0 blocked; 0 not run; 0 missing | partial |
4. https://www.reddit.com | interaction-heavy | partialatomic-coverage-incomplete | 49 judged; 9 blocked; 0 not run; 0 missing | partial |
5. https://www.amazon.com | commerce | partialatomic-coverage-incomplete | 39 judged; 19 blocked; 0 not run; 0 missing | partial |
6. https://en.wikipedia.org | simple-static | runner-completednot-applicable | 58 judged; 0 blocked; 0 not run; 0 missing | partial |
7. https://www.nytimes.com | content-news | partialatomic-coverage-incomplete | 50 judged; 6 blocked; 2 not run; 0 missing | partial |
8. https://gemini.google.com | spa-product | partialatomic-coverage-incomplete | 36 judged; 22 blocked; 0 not run; 0 missing | partial |
9. https://www.netflix.com | media-heavy | partialatomic-coverage-incomplete | 50 judged; 8 blocked; 0 not run; 0 missing | partial |
10. https://www.apple.com | general | partialatomic-coverage-incomplete | 49 judged; 9 blocked; 0 not run; 0 missing | partial |
1. https://github.com
Ledger: error · documentation
Reason: missing-report. The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.
Recomputed coverage: expected 58; recorded 0; judged 0; blocked 0; not run 0; missing 58; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-like-url) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
No network or performance summary was derivable because no report was produced.

2. https://web.facebook.com
Ledger: error · blocked-complex
Reason: runner-error. The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.
Recomputed coverage: expected 58; recorded 0; judged 0; blocked 0; not run 0; missing 58; unknown 0; duplicates 0.
Journey actions
- baseline-load: invalidated-cross-origin-state (
completed-cross-origin-state) - bounded-scroll: invalidated-cross-origin-state (
completed-cross-origin-state)
Safe aggregates
No network or performance summary was derivable because no report was produced.

3. https://www.google.com
Ledger: runner-completed · search-portal
Reason: not-applicable. The ledger classified this row completed, but its console pass used an unsupported absence-of-error method.
Recomputed coverage: expected 58; recorded 58; judged 58; blocked 0; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-method-post) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 43
- Transfer
- 816,024 bytes
- Observed origins
- 1 first-party; 4 third-party
- Security headers present
- 2 of 6
- Cookies observed
- 2 total; 2 Secure; 2 HttpOnly
- Trace timings
- FCP 1,482.05 ms; LCP 1,499.69 ms; 1 long tasks

4. https://www.reddit.com
Ledger: partial · interaction-heavy
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after a humanity challenge and cross-origin navigation block.
Recomputed coverage: expected 58; recorded 58; judged 49; blocked 9; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: navigation-blocked (
cross-origin-document) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 8
- Transfer
- 514,675 bytes
- Observed origins
- 1 first-party; 3 third-party
- Security headers present
- 3 of 6
- Cookies observed
- 1 total; 1 Secure; 0 HttpOnly
- Trace timings
- FCP 677.84 ms; LCP 677.84 ms; 0 long tasks

5. https://www.amazon.com
Ledger: partial · commerce
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like POST traffic was blocked.
Recomputed coverage: expected 58; recorded 58; judged 39; blocked 19; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-method-post) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 43
- Transfer
- 1,101,460 bytes
- Observed origins
- 1 first-party; 3 third-party
- Security headers present
- 0 of 6
- Cookies observed
- 7 total; 3 Secure; 0 HttpOnly
- Trace timings
- FCP not observed; LCP not observed; 0 long tasks

6. https://en.wikipedia.org
Ledger: runner-completed · simple-static
Reason: not-applicable. The ledger classified this row completed, but its console pass used an unsupported absence-of-error method.
Recomputed coverage: expected 58; recorded 58; judged 58; blocked 0; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: performed (
same-origin-get-navigation-completed) - bounded-scroll: performed (
bounded-scroll-completed)
Safe aggregates
- Requests
- 40
- Transfer
- 682,131 bytes
- Observed origins
- 1 first-party; 2 third-party
- Security headers present
- 3 of 6
- Cookies observed
- 7 total; 6 Secure; 4 HttpOnly
- Trace timings
- FCP 1,123.66 ms; LCP 1,123.66 ms; 1 long tasks

7. https://www.nytimes.com
Ledger: partial · content-news
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like POST traffic was blocked; six checks were blocked and two were not run.
Recomputed coverage: expected 58; recorded 58; judged 50; blocked 6; not run 2; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-method-post) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 368
- Transfer
- 4,950,323 bytes
- Observed origins
- 1 first-party; 27 third-party
- Security headers present
- 5 of 6
- Cookies observed
- 14 total; 9 Secure; 3 HttpOnly
- Trace timings
- FCP 755.89 ms; LCP 4,760.96 ms; 13 long tasks

8. https://gemini.google.com
Ledger: partial · spa-product
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like OPTIONS traffic was blocked.
Recomputed coverage: expected 58; recorded 58; judged 36; blocked 22; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-method-options) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 100
- Transfer
- 5,424,867 bytes
- Observed origins
- 1 first-party; 8 third-party
- Security headers present
- 4 of 6
- Cookies observed
- 1 total; 1 Secure; 1 HttpOnly
- Trace timings
- FCP 6,623.01 ms; LCP 6,861.74 ms; 2 long tasks

9. https://www.netflix.com
Ledger: partial · media-heavy
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete under mutation and scope boundaries and unavailable diagnostic or account-flow evidence.
Recomputed coverage: expected 58; recorded 58; judged 50; blocked 8; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-like-url) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 53
- Transfer
- 2,579,592 bytes
- Observed origins
- 1 first-party; 11 third-party
- Security headers present
- 3 of 6
- Cookies observed
- 9 total; 3 Secure; 3 HttpOnly
- Trace timings
- FCP 2,462.92 ms; LCP 2,462.92 ms; 2 long tasks

10. https://www.apple.com
Ledger: partial · general
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete because exact-origin scope left primary, error, account, and diagnostic flows unavailable.
Recomputed coverage: expected 58; recorded 58; judged 49; blocked 9; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: performed (
same-origin-get-navigation-completed) - bounded-scroll: performed (
bounded-scroll-completed)
Safe aggregates
- Requests
- 51
- Transfer
- 1,965,289 bytes
- Observed origins
- 1 first-party; 1 third-party
- Security headers present
- 5 of 6
- Cookies observed
- 6 total; 5 Secure; 0 HttpOnly
- Trace timings
- FCP 6,310.82 ms; LCP 8,477.41 ms; 0 long tasks

Row evidence details
1. https://github.com
Ledger: error · documentation
Reason: missing-report. The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.
Recomputed coverage: expected 58; recorded 0; judged 0; blocked 0; not run 0; missing 58; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-like-url) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
No network or performance summary was derivable because no report was produced.

2. https://web.facebook.com
Ledger: error · blocked-complex
Reason: runner-error. The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.
Recomputed coverage: expected 58; recorded 0; judged 0; blocked 0; not run 0; missing 58; unknown 0; duplicates 0.
Journey actions
- baseline-load: invalidated-cross-origin-state (
completed-cross-origin-state) - bounded-scroll: invalidated-cross-origin-state (
completed-cross-origin-state)
Safe aggregates
No network or performance summary was derivable because no report was produced.

3. https://www.google.com
Ledger: runner-completed · search-portal
Reason: not-applicable. The ledger classified this row completed, but its console pass used an unsupported absence-of-error method.
Recomputed coverage: expected 58; recorded 58; judged 58; blocked 0; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-method-post) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 43
- Transfer
- 816,024 bytes
- Observed origins
- 1 first-party; 4 third-party
- Security headers present
- 2 of 6
- Cookies observed
- 2 total; 2 Secure; 2 HttpOnly
- Trace timings
- FCP 1,482.05 ms; LCP 1,499.69 ms; 1 long tasks

4. https://www.reddit.com
Ledger: partial · interaction-heavy
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after a humanity challenge and cross-origin navigation block.
Recomputed coverage: expected 58; recorded 58; judged 49; blocked 9; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: navigation-blocked (
cross-origin-document) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 8
- Transfer
- 514,675 bytes
- Observed origins
- 1 first-party; 3 third-party
- Security headers present
- 3 of 6
- Cookies observed
- 1 total; 1 Secure; 0 HttpOnly
- Trace timings
- FCP 677.84 ms; LCP 677.84 ms; 0 long tasks

5. https://www.amazon.com
Ledger: partial · commerce
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like POST traffic was blocked.
Recomputed coverage: expected 58; recorded 58; judged 39; blocked 19; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-method-post) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 43
- Transfer
- 1,101,460 bytes
- Observed origins
- 1 first-party; 3 third-party
- Security headers present
- 0 of 6
- Cookies observed
- 7 total; 3 Secure; 0 HttpOnly
- Trace timings
- FCP not observed; LCP not observed; 0 long tasks

6. https://en.wikipedia.org
Ledger: runner-completed · simple-static
Reason: not-applicable. The ledger classified this row completed, but its console pass used an unsupported absence-of-error method.
Recomputed coverage: expected 58; recorded 58; judged 58; blocked 0; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: performed (
same-origin-get-navigation-completed) - bounded-scroll: performed (
bounded-scroll-completed)
Safe aggregates
- Requests
- 40
- Transfer
- 682,131 bytes
- Observed origins
- 1 first-party; 2 third-party
- Security headers present
- 3 of 6
- Cookies observed
- 7 total; 6 Secure; 4 HttpOnly
- Trace timings
- FCP 1,123.66 ms; LCP 1,123.66 ms; 1 long tasks

7. https://www.nytimes.com
Ledger: partial · content-news
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like POST traffic was blocked; six checks were blocked and two were not run.
Recomputed coverage: expected 58; recorded 58; judged 50; blocked 6; not run 2; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-method-post) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 368
- Transfer
- 4,950,323 bytes
- Observed origins
- 1 first-party; 27 third-party
- Security headers present
- 5 of 6
- Cookies observed
- 14 total; 9 Secure; 3 HttpOnly
- Trace timings
- FCP 755.89 ms; LCP 4,760.96 ms; 13 long tasks

8. https://gemini.google.com
Ledger: partial · spa-product
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like OPTIONS traffic was blocked.
Recomputed coverage: expected 58; recorded 58; judged 36; blocked 22; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-method-options) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 100
- Transfer
- 5,424,867 bytes
- Observed origins
- 1 first-party; 8 third-party
- Security headers present
- 4 of 6
- Cookies observed
- 1 total; 1 Secure; 1 HttpOnly
- Trace timings
- FCP 6,623.01 ms; LCP 6,861.74 ms; 2 long tasks

9. https://www.netflix.com
Ledger: partial · media-heavy
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete under mutation and scope boundaries and unavailable diagnostic or account-flow evidence.
Recomputed coverage: expected 58; recorded 58; judged 50; blocked 8; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: mutation-blocked (
mutation-like-url) - bounded-scroll: skipped-after-blocker (
journey-aborted-after-blocker)
Safe aggregates
- Requests
- 53
- Transfer
- 2,579,592 bytes
- Observed origins
- 1 first-party; 11 third-party
- Security headers present
- 3 of 6
- Cookies observed
- 9 total; 3 Secure; 3 HttpOnly
- Trace timings
- FCP 2,462.92 ms; LCP 2,462.92 ms; 2 long tasks

10. https://www.apple.com
Ledger: partial · general
Reason: atomic-coverage-incomplete. Atomic coverage was incomplete because exact-origin scope left primary, error, account, and diagnostic flows unavailable.
Recomputed coverage: expected 58; recorded 58; judged 49; blocked 9; not run 0; missing 0; unknown 0; duplicates 0.
Journey actions
- baseline-load: performed (
same-origin-get-navigation-completed) - bounded-scroll: performed (
bounded-scroll-completed)
Safe aggregates
- Requests
- 51
- Transfer
- 1,965,289 bytes
- Observed origins
- 1 first-party; 1 third-party
- Security headers present
- 5 of 6
- Cookies observed
- 6 total; 5 Secure; 0 HttpOnly
- Trace timings
- FCP 6,310.82 ms; LCP 8,477.41 ms; 0 long tasks

All 580 tests and evidence
This explorer retains every catalog slot. An issue is a tested product failure. Blocked and not run mean incomplete collection. Unavailable means the site produced no atomic report, so no test result is claimed.
Exact totals: 174 pass; 147 issues; 68 not applicable; 73 blocked; 2 not run; 116 unavailable.
Showing all 580 site-check slots.
1. https://github.com · 58 slots · missing-report
Collection failure: these 58 slots were materialized from the catalog so the denominator remains visible. They were not tested and have no check-specific evidence.
Respect user preferences · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
respects-color-schemeHonours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
respects-reduced-motionHonours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
respects-contrastHonours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Implement natural interactions · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
view-transitionsState and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
scroll-driven-animationsScroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
physical-gesturesGesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Provide guided navigation · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
scroll-state-aware-chromeSticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
anchored-positioningTooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
directs-attentionNavigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Maximize content, reduce noise · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-intrusive-interruptionsNo intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
semantic-dismissible-primitivesOverlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
reduced-chromeMinimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Adapt to the form factor · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
responsive-no-horizontal-scrollLayout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
component-level-responsivenessComponents adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
input-modality-awareTouch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Support core task success · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
clear-purpose-and-primary-actionThe page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
primary-flow-completionThe representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
clear-system-state-and-recoveryLoading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Be fast and stable · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
good-core-web-vitalsCore Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
visual-stabilityNo cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
efficient-main-threadThe main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
efficient-resource-deliveryCritical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
trim-unused-and-duplicate-codeThe page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Be inclusive · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
names-roles-labelsInteractive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
sufficient-contrastText and essential UI meet WCAG colour-contrast minimums against their background. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
structure-and-focusHeading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
legible-textText is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
zoom-reflow-targets-and-mediaThe experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Follow best practices · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-console-errorsThe page loads without console errors or uncaught exceptions. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
sound-document-and-assetsValid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.) | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
browser-platform-hygieneThe page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Be discoverable · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
title-and-descriptionThe page has a unique, descriptive <title> and a meta description. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
crawlable-and-mobile-friendlyLinks are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
canonical-and-indexing-signalsPublic pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
structured-and-shareable-metadataWhere the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Be private and secure · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
secure-transport-and-headersServed over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
data-minimisation-and-third-partiesNo over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
in-context-permissions-and-modern-authPermission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
defensive-browser-policiesBrowser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Be resilient · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
progressive-enhancementCore content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
resilient-runtime-behaviourThe page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
offline-and-installableWhere the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.) | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
network-and-http-failure-statesHTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Be internationalised · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
lang-dir-and-logical-propertiesCorrect lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
locale-aware-dataDates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
time-zone-correctnessTime handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Be trustworthy · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-dark-patternsNo deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive). | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
humane-error-handlingForms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
trustworthy-input-assistanceInput is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
safe-commercial-and-account-flowsCheckout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Be sustainable · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
optimised-assetsImages and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
no-wasteful-workBackground work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
third-party-and-media-budgetThird-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Be agent ready · 2 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
structured-agent-capabilitiesWhere it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
on-device-inferenceOn-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
Be memory-efficient · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-leak-under-repeated-interactionRepeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
bounded-footprintHeap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
no-detached-dom-or-unbounded-listenersNo growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session. | Unavailable Source: missing-report; confidence: none | Implementation and methodCatalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry. |
2. https://web.facebook.com · 58 slots · runner-error
Collection failure: these 58 slots were materialized from the catalog so the denominator remains visible. They were not tested and have no check-specific evidence.
Respect user preferences · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
respects-color-schemeHonours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
respects-reduced-motionHonours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
respects-contrastHonours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Implement natural interactions · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
view-transitionsState and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
scroll-driven-animationsScroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
physical-gesturesGesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Provide guided navigation · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
scroll-state-aware-chromeSticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
anchored-positioningTooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
directs-attentionNavigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Maximize content, reduce noise · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-intrusive-interruptionsNo intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
semantic-dismissible-primitivesOverlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
reduced-chromeMinimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Adapt to the form factor · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
responsive-no-horizontal-scrollLayout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
component-level-responsivenessComponents adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
input-modality-awareTouch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Support core task success · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
clear-purpose-and-primary-actionThe page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
primary-flow-completionThe representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
clear-system-state-and-recoveryLoading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Be fast and stable · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
good-core-web-vitalsCore Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
visual-stabilityNo cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
efficient-main-threadThe main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
efficient-resource-deliveryCritical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
trim-unused-and-duplicate-codeThe page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Be inclusive · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
names-roles-labelsInteractive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
sufficient-contrastText and essential UI meet WCAG colour-contrast minimums against their background. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
structure-and-focusHeading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
legible-textText is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
zoom-reflow-targets-and-mediaThe experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Follow best practices · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-console-errorsThe page loads without console errors or uncaught exceptions. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
sound-document-and-assetsValid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.) | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
browser-platform-hygieneThe page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Be discoverable · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
title-and-descriptionThe page has a unique, descriptive <title> and a meta description. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
crawlable-and-mobile-friendlyLinks are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
canonical-and-indexing-signalsPublic pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
structured-and-shareable-metadataWhere the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Be private and secure · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
secure-transport-and-headersServed over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
data-minimisation-and-third-partiesNo over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
in-context-permissions-and-modern-authPermission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
defensive-browser-policiesBrowser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Be resilient · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
progressive-enhancementCore content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
resilient-runtime-behaviourThe page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
offline-and-installableWhere the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.) | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
network-and-http-failure-statesHTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Be internationalised · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
lang-dir-and-logical-propertiesCorrect lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
locale-aware-dataDates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
time-zone-correctnessTime handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Be trustworthy · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-dark-patternsNo deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive). | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
humane-error-handlingForms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
trustworthy-input-assistanceInput is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
safe-commercial-and-account-flowsCheckout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Be sustainable · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
optimised-assetsImages and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
no-wasteful-workBackground work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
third-party-and-media-budgetThird-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Be agent ready · 2 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
structured-agent-capabilitiesWhere it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
on-device-inferenceOn-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
Be memory-efficient · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-leak-under-repeated-interactionRepeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
bounded-footprintHeap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
no-detached-dom-or-unbounded-listenersNo growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session. | Unavailable Source: runner-error; confidence: none | Implementation and methodCatalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses. Method used or attempted: No check-specific method was executed because the site report is unavailable. | No atomic report or check-specific evidence was produced for this site-check slot. Reference: No artifact reference; unavailable | The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin. |
3. https://www.google.com · 58 slots · available
Method warning: this site has a report, but its no-console-errors pass is method-invalid and is labelled in the table.
Respect user preferences · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
respects-color-schemeHonours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method. Method used or attempted: Compared permit-bound light and dark screenshots. | The page rendered a light surface under prefers-color-scheme: light and a dark surface under dark, with legible consent UI in both. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
respects-reduced-motionHonours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses. Method used or attempted: Permit-bound evaluate probe under prefers-reduced-motion: reduce. | The media query matched and [host/path omitted]() returned zero active animations. Reference: other-private-evidence; retained-private | The retained evidence directly supported this check under the captured conditions. |
respects-contrastHonours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. Method used or attempted: Permit-bound forced-colors and prefers-contrast screenshot review. | Essential consent text, links, outline borders, language, and sign-in controls remained visible in forced colors. Reference: other-private-evidence; retained-private | The retained evidence directly supported this check under the captured conditions. |
Implement natural interactions · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
view-transitionsState and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps. | Not applicable Source: not-applicable; confidence: medium | Implementation and methodCatalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. Method used or attempted: Reviewed the static entry and journey states. | No same-document or route state change was safely exercised on this single search entry surface. Reference: No artifact reference; described-only | No applicable state or route transition was present in the bounded audited path. |
scroll-driven-animationsScroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. Method used or attempted: Layout, reduced-motion probe, and journey state review. | The entry document had no scrollable desktop surface and no scroll-linked animation. Reference: layout-summary, other-private-evidence; retained-private | There is no scroll-linked visual effect on this surface. |
physical-gesturesGesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. Method used or attempted: Reviewed DOM, layout, and bounded repeated native scrolling. | The page relies on native scrolling and native form controls; no custom gesture surface or pointer-driven replacement was observed. Reference: memory-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Provide guided navigation · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
scroll-state-aware-chromeSticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. Method used or attempted: Journey and layout review. | The desktop entry state has no meaningful scrolling chrome; the planned bounded scroll was skipped after the journey network mutation blocker. Reference: journey-summary, layout-summary; retained-private | No sticky or affixed chrome needs to react on this bounded single-screen path. |
anchored-positioningTooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned. | Not applicable Source: not-applicable; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. Method used or attempted: DOM review for popovers and anchored transient UI. | No tooltip or popover was open or required on the audited state; the consent surface is modal. Reference: page-probe-summary; retained-private | No anchored transient overlay was present on the audited state. |
directs-attentionNavigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. Method used or attempted: Visual hierarchy review of desktop, mobile, and journey baseline screenshots. | The centered Google mark, search field, and consent heading establish a clear reading order and visible current state. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
Maximize content, reduce noise · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-intrusive-interruptionsNo intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. Method used or attempted: First-load screenshot review at the journey viewport. | A consent wall obscures the search task on load; at 780x493 its action controls are below the visible viewport. Reference: screenshot; retained-private | The first-load consent wall obscures the core search task and its actions fall below a short desktop viewport.
|
semantic-dismissible-primitivesOverlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. Method used or attempted: DOM and role inspection. | The consent surface is exposed as a dialog and the DOM contains a dialog element rather than only an unstructured overlay div. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
reduced-chromeMinimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. Method used or attempted: Desktop and mobile screenshot review. | Outside the required consent decision, the search page is visually sparse and gives the central task most of the viewport. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
Adapt to the form factor · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
responsive-no-horizontal-scrollLayout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. Method used or attempted: Permit-bound 360x800 layout metrics and screenshot. | The response has no viewport meta tag; Chrome laid it out at 980 CSS px and scaled it to [host/path omitted] on the 360px device viewport. Reference: layout-summary, screenshot; retained-private | The page omits viewport metadata and scales a 980px layout down on mobile.
|
component-level-responsivenessComponents adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints. | Not applicable Source: not-applicable; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. Method used or attempted: CSS feature probe and component inventory review. | No reused component was observed in multiple container contexts; the page uses a single search layout. Reference: page-probe-summary; retained-private | The audited page has no evidenced multi-container reuse case requiring component-level adaptation. |
input-modality-awareTouch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. Method used or attempted: Programmatically focused the first 20 focusable controls and inspected outlines and target rectangles. | Focused links and controls reported outline-style none and :focus-visible false; several text links were only 24 to 26 CSS px tall. Reference: page-probe-summary; retained-private | Keyboard focus is not visibly exposed on sampled controls.
|
Support core task success · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
clear-purpose-and-primary-actionThe page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses. Method used or attempted: First viewport and semantic form review. | The Google brand, labelled Search combobox, and clearly named search actions communicate the page purpose. Reference: page-probe-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
primary-flow-completionThe representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses. Method used or attempted: Reviewed the already-replayed strict journey and baseline screenshot. | The representative path did not reach search: the baseline action was marked mutation-blocked after POST telemetry and the bounded scroll was skipped; the consent wall also obscured search controls. Reference: journey-summary, screenshot; retained-private | The first-load consent wall obscures the core search task and its actions fall below a short desktop viewport.
|
clear-system-state-and-recoveryLoading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses. Method used or attempted: Consent state and journey ledger review. | The page explains the consent state and offers Reject all, Accept all, More options, privacy, and terms paths; the journey ledger explicitly records its blocker. Reference: journey-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
Be fast and stable · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
good-core-web-vitalsCore Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses. Method used or attempted: Permit-bound trace and layout observers. | LCP was [host/path omitted] ms, CLS was 0, and total blocking time was [host/path omitted] ms; no INP sample was produced because policy prohibited form submission. Reference: layout-summary, performance-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
visual-stabilityNo cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. Method used or attempted: Layout shift observer during page load. | The layout primitive recorded CLS 0 with no shift entries. Reference: layout-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
efficient-main-threadThe main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. Method used or attempted: Permit-bound performance trace. | The trace recorded one [host/path omitted] ms long task and only [host/path omitted] ms total blocking time. Reference: performance-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
efficient-resource-deliveryCritical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses. Method used or attempted: HAR summary plus rendered DOM review. | The simple entry transferred 816,024 bytes across 43 requests; 715,201 bytes were script and parser-inserted stylesheets appeared as render-blocking candidates. Reference: network-summary, page-probe-summary; retained-private | Resource delivery is heavy for the entry task and includes parser-inserted blocking styles.
|
trim-unused-and-duplicate-codeThe page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value. | Failed / issue Source: issues; confidence: medium | Implementation and methodCatalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses. Method used or attempted: HAR resource-type and page-complexity comparison. | Eight scripts transferred 715,201 bytes for a roughly 502-node search entry. The payload is disproportionate, although this run did not collect code coverage to isolate exact unused ranges. Reference: network-summary, page-probe-summary; retained-private | JavaScript payload is disproportionate to the initial search surface.
|
Be inclusive · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
names-roles-labelsInteractive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses. Method used or attempted: DOM accessibility semantics and visible control review. | The primary search textarea is labelled Search with role combobox, the form has role search, visible submit controls are named, and the visible Google SVG is exposed as an image. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
sufficient-contrastText and essential UI meet WCAG colour-contrast minimums against their background. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. Method used or attempted: Normal dark and forced-colors screenshot review. | Primary and consent text remained legible against dark surfaces, and essential UI survived forced colors. Reference: other-private-evidence, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
structure-and-focusHeading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. Method used or attempted: Heading, landmark, and programmatic focus inspection. | A logical H1 and search landmark exist, but the focus probe found no visible outline or :focus-visible match on the sampled keyboard-reachable links and controls. Reference: page-probe-summary; retained-private | Keyboard focus is not visibly exposed on sampled controls.
|
legible-textText is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. Method used or attempted: Desktop, mobile, and contrast screenshot review. | Search and consent copy uses readable line lengths and spacing; no text clipping was visible inside the mobile full-page capture. Reference: other-private-evidence, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
zoom-reflow-targets-and-mediaThe experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses. Method used or attempted: Narrow viewport layout and screenshot review. | Without viewport metadata, the page renders a 980px layout scaled down to 360px, making text and controls unusually small instead of reflowing at device width. Reference: layout-summary, screenshot; retained-private | The page omits viewport metadata and scales a 980px layout down on mobile.
|
Follow best practices · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-console-errorsThe page loads without console errors or uncaught exceptions. | Pass Source: pass; confidence: lowMethod invalid: this console pass used absence of a surfaced error even though the authoritative console collector was unavailable. Treat it as unreliable evidence, not a valid pass. | Implementation and methodCatalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. Method used or attempted: Observed permit-bound Chrome CLI load, trace completion, and journey event ledger. | The independent loads completed without an uncaught-exception signal in CLI output; the earlier journey console collector itself was unavailable, so confidence is low. Reference: journey-summary, performance-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
sound-document-and-assetsValid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.) | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. Method used or attempted: DOM, document metadata, image, and layout inspection. | The document has an HTML doctype and UTF-8 charset; visible SVG imagery retained its aspect ratio and the layout recorded CLS 0. Empty image placeholders reported by the image scanner were not visible content. Reference: image-summary, layout-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
browser-platform-hygieneThe page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load. | Pass Source: pass; confidence: low | Implementation and methodCatalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. Method used or attempted: DOM and runtime surface review. | No geolocation or notification prompt appeared on load, paste prevention was not observed, and the page remained functional in the raw no-JavaScript crawler view. Reference: discoverability-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Be discoverable · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
title-and-descriptionThe page has a unique, descriptive <title> and a meta description. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. Method used or attempted: DOM metadata probe. | The page title is Google, but meta[name=description] is absent. Reference: page-probe-summary; retained-private | The public homepage has no meta description.
|
crawlable-and-mobile-friendlyLinks are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. Method used or attempted: DOM links, raw HTML comparison, and mobile layout inspection. | Most navigation links have real href values, but the viewport meta tag is absent and one More options anchor lacks href. Reference: layout-summary, page-probe-summary; retained-private | Missing viewport metadata undermines mobile crawlability and presentation.
|
canonical-and-indexing-signalsPublic pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses. Method used or attempted: Discoverability fetch and metadata inspection. | The raw document returned HTTP 200, no accidental noindex signal was present, and the canonical homepage URL remained on the reviewed origin. Reference: discoverability-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
structured-and-shareable-metadataWhere the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses. Method used or attempted: Entity and metadata review. | The entry is a utility search form, not an article, product, event, or other rich entity needing share cards or JSON-LD. Reference: page-probe-summary; retained-private | The generic search entry does not represent a rich content entity. |
Be private and secure · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
secure-transport-and-headersServed over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses. Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence. | Retained evidence records 2 of 6 baseline security-header categories present. Cookie-attribute review records 2 of 2 records with Secure and 2 with HttpOnly. Names, values, routes, and raw headers remain private. Reference: cookie-attribute-summary, security-header-summary; retained-private | The main document omits several baseline transport and content defenses.
|
data-minimisation-and-third-partiesNo over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses. Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence. | Retained evidence records 1 first-party origin categories and 4 third-party origin categories across 43 requests. Destinations, identifiers, routes, headers, and bodies remain private. Reference: cookie-attribute-summary, network-summary, screenshot, tracker-summary; retained-private | The initial page has a material third-party and long-lived identifier footprint.
|
in-context-permissions-and-modern-authPermission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone. | Not applicable Source: not-applicable; confidence: medium | Implementation and methodCatalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses. Method used or attempted: Initial-load prompt and auth-surface review. | No browser permission prompt appeared. Authentication is only a cross-origin Sign in link and was outside the permitted journey. Reference: page-probe-summary, screenshot; retained-private | The bounded audited path does not contain an authentication flow to assess passkeys or WebAuthn. |
defensive-browser-policiesBrowser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses. Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence. | Retained evidence records 2 of 6 baseline browser-policy header categories present. Values and raw headers remain private. Reference: security-header-summary; retained-private | Defensive browser policy coverage is incomplete.
|
Be resilient · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
progressive-enhancementCore content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. Method used or attempted: Permit-bound discoverability raw fetch and crawler screenshot. | With JavaScript disabled, the Google mark, search field, search buttons, navigation, and footer remain rendered and usable; the raw response was not a JS shell. Reference: discoverability-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
resilient-runtime-behaviourThe page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. Method used or attempted: Journey viewport screenshot and dialog rectangle review. | At 780x493 the consent dialog extends below the viewport and its decision controls are not initially visible, so the mandatory overlay is cut off in a common short viewport. Reference: screenshot; retained-private | The consent dialog is cut off in a short viewport.
|
offline-and-installableWhere the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.) | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. Method used or attempted: Manifest and service-worker registration probe. | No manifest or active service-worker registration exists. Search results intrinsically require a live network. Reference: page-probe-summary; retained-private | This live search portal is intrinsically online; installability and offline result retrieval are not reasonable requirements. |
network-and-http-failure-statesHTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states. | Not applicable Source: not-applicable; confidence: medium | Implementation and methodCatalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses. Method used or attempted: Site purpose and bounded policy review. | The audited entry is a network-dependent search portal and the strict journey did not permit simulated failing submissions. Reference: journey-summary; retained-private | Search result retrieval is intrinsically online and no safe failure-producing action was permitted in this run. |
Be internationalised · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
lang-dir-and-logical-propertiesCorrect lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. Method used or attempted: DOM and stylesheet feature probe. | The document declares lang=en-GB and stylesheet inspection found logical inline properties. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
locale-aware-dataDates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. Method used or attempted: Rendered content and control inventory review. | The bounded entry state displays no dates, numbers, currencies, durations, or calendar data. Reference: page-probe-summary; retained-private | No locale-sensitive data is rendered on this path. |
time-zone-correctnessTime handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. Method used or attempted: Rendered content and platform probe. | The page renders no events or time values requiring time-zone or DST handling. Reference: page-probe-summary; retained-private | No time concepts are present on this path. |
Be trustworthy · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-dark-patternsNo deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive). | Failed / issue Source: issues; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. Method used or attempted: Consent wall visual and copy review. | The mandatory consent interstitial blocks the core search task on first load and describes measurement and personalised advertising before the user can proceed, although Reject all and Accept all are given equal visual weight on the full mobile capture. Reference: screenshot; retained-private | A mandatory consent interstitial delays the primary task.
|
humane-error-handlingForms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. Method used or attempted: Form requirements inspection. | The search field is not required and no validation or error state exists on the untouched entry path. Reference: page-probe-summary; retained-private | No user-error validation state applies to the optional search field before submission. |
trustworthy-input-assistanceInput is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. Method used or attempted: Search control semantics inspection. | The primary control is a labelled search combobox named q with native form submission; no paste prevention or obstructive formatting was observed. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
safe-commercial-and-account-flowsCheckout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses. Method used or attempted: Consent choice visual review. | Reject all and Accept all are both present with equal visual treatment, More options is available, and privacy and terms links are visible. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
Be sustainable · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
optimised-assetsImages and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. Method used or attempted: Image and HAR resource review. | The seven image requests transferred only 972 bytes; the visible product mark is SVG and no oversized image was reported. Reference: image-summary, network-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
no-wasteful-workBackground work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work. | Failed / issue Source: issues; confidence: medium | Implementation and methodCatalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. Method used or attempted: HAR and page-complexity review. | A roughly 502-node search entry transfers 715,201 bytes of JavaScript across eight script requests, disproportionate to the initial task surface. Reference: network-summary, page-probe-summary; retained-private | JavaScript payload is disproportionate to the initial search surface.
|
third-party-and-media-budgetThird-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible. | Failed / issue Source: issues; confidence: medium | Implementation and methodCatalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses. Method used or attempted: HAR and tracker review. | Eight third-party requests transferred 92,305 bytes across five third-party origins on the initial search page, before the user performs a search. Reference: network-summary, tracker-summary; retained-private | The initial page has a material third-party and long-lived identifier footprint.
|
Be agent ready · 2 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
structured-agent-capabilitiesWhere it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. Method used or attempted: Raw HTML and semantic form inspection. | The primary capability is exposed as a standard server-rendered form with role search, a labelled q combobox, named submit controls, and real links, so agents need not infer a canvas-only interaction. Reference: discoverability-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
on-device-inferenceOn-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses. Method used or attempted: Page purpose and runtime capability review. | The entry is a remote web-search portal; no bounded local summarisation or inference task is exposed. Reference: page-probe-summary; retained-private | On-device inference does not improve the audited search-entry task. |
Be memory-efficient · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-leak-under-repeated-interactionRepeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses. Method used or attempted: Compared independent heap summaries before and after ten bounded native scroll down/up cycles. | Post-cycle heap self size decreased by 7,716 bytes and node count decreased by 35 versus baseline, with no retained-growth signal. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
bounded-footprintHeap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses. Method used or attempted: Heap summary and DOM-size review. | The page held about [host/path omitted] MB V8 self size and roughly 502 DOM elements, proportionate for this feature-rich search entry. Reference: memory-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
no-detached-dom-or-unbounded-listenersNo growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses. Method used or attempted: Compared heap constructor summaries across repeated bounded scroll. | The post-cycle heap did not grow overall and no Detached constructor appeared among the largest retained constructors. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
4. https://www.reddit.com · 58 slots · available
Respect user preferences · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
respects-color-schemeHonours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Desktop and dark/high-contrast screenshots are pixel-identical at 40,303 bytes; DOM reports color-scheme normal and fixed white surfaces. Reference: No artifact reference; described-only | Dark preference is ignored
|
respects-reduced-motionHonours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Reduced-motion video captured one static frame and the probe found zero active animations. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
respects-contrastHonours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Text and challenge controls remain visually legible in the high-contrast screenshot; no content disappears. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
Implement natural interactions · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
view-transitionsState and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. Method used or attempted: Permit-bound recon plus strict journey replay result | Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. Reference: No artifact reference; described-only | The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. |
scroll-driven-animationsScroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The response is exactly one viewport tall and has no scroll-driven content. Reference: No artifact reference; described-only | The response is exactly one viewport tall and has no scroll-driven content. |
physical-gesturesGesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The static challenge response exposes no gesture-driven UI. Reference: No artifact reference; described-only | The static challenge response exposes no gesture-driven UI. |
Provide guided navigation · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
scroll-state-aware-chromeSticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The response has no scroll range or chrome whose state could react to scrolling. Reference: No artifact reference; described-only | The response has no scroll range or chrome whose state could react to scrolling. |
anchored-positioningTooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | No page-owned tooltip, popover, or anchored overlay is exposed without performing the prohibited CAPTCHA interaction. Reference: No artifact reference; described-only | No page-owned tooltip, popover, or anchored overlay is exposed without performing the prohibited CAPTCHA interaction. |
directs-attentionNavigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Desktop and mobile screenshots provide a direct H1, explanatory copy, and centered CAPTCHA with clear visual hierarchy. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
Maximize content, reduce noise · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-intrusive-interruptionsNo intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. Method used or attempted: Permit-bound recon plus strict journey replay result | Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. Reference: No artifact reference; described-only | The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. |
semantic-dismissible-primitivesOverlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | No page-owned dismissible dialog or popover is present in the observed response. Reference: No artifact reference; described-only | No page-owned dismissible dialog or popover is present in the observed response. |
reduced-chromeMinimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Screenshots show minimal branding/footer chrome around the single challenge task. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
Adapt to the form factor · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
responsive-no-horizontal-scrollLayout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Layout reports 0 horizontal overflow at both 1440x1000 and 360x800. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
component-level-responsivenessComponents adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Desktop/mobile screenshots and DOM CSS show the main surface constrains to 480px and footer switches to a column below 600px. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
input-modality-awareTouch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Programmatic focus on the first link produced a native auto 1px focus outline. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
Support core task success · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
clear-purpose-and-primary-actionThe page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The H1 and copy explicitly state the purpose and point to the reCAPTCHA as the only primary action. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
primary-flow-completionThe representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses. Method used or attempted: Permit-bound recon plus strict journey replay result | Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. Reference: No artifact reference; described-only | The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. |
clear-system-state-and-recoveryLoading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses. Method used or attempted: Permit-bound recon plus strict journey replay result | Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. Reference: No artifact reference; described-only | The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. |
Be fast and stable · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
good-core-web-vitalsCore Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Trace reports LCP/FCP [host/path omitted], 0 long tasks, and 0ms total blocking time; layout reports CLS 0. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
visual-stabilityNo cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Both layout captures report CLS 0 with no layout shifts. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
efficient-main-threadThe main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Trace reports no long tasks and 0ms total blocking time over [host/path omitted]. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
efficient-resource-deliveryCritical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | HAR summary records 514,675 transferred bytes for the challenge; 346,849 bytes are third-party and a parser-inserted reCAPTCHA script has no async/defer. Reference: No artifact reference; described-only | The minimal wall transfers half a megabyte, mostly third-party JavaScript
|
trim-unused-and-duplicate-codeThe page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses. Method used or attempted: Permit-bound recon plus strict journey replay result | Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. Reference: No artifact reference; described-only | The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. |
Be inclusive · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
names-roles-labelsInteractive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The evaluated control inventory shows the 100x32 Reddit logo link has no text, aria-label, or title; the image audit reports its image has no alt. Reference: No artifact reference; described-only | The Reddit logo link has no accessible name
|
sufficient-contrastText and essential UI meet WCAG colour-contrast minimums against their background. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Visual review of desktop, mobile, and increased-contrast captures shows readable dark text and visible controls on white surfaces. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
structure-and-focusHeading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The probe found one H1 and a visible native focus outline, but zero main, nav, header, or footer landmark elements. Reference: No artifact reference; described-only | The page has no semantic landmarks
|
legible-textText is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Desktop/mobile screenshots show a 24px heading and 16px centered body copy with readable line lengths. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
zoom-reflow-targets-and-mediaThe experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | At 360px there is no overflow, but the evaluated footer links are only 14 CSS px tall, below a comfortable touch target. Reference: No artifact reference; described-only | Footer links have undersized touch targets
|
Follow best practices · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-console-errorsThe page loads without console errors or uncaught exceptions. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. Method used or attempted: Permit-bound recon plus strict journey replay result | Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. Reference: No artifact reference; described-only | The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. |
sound-document-and-assetsValid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.) | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The image audit reports the sole image lacks width and height attributes; DOM confirms it renders at 100x28 from a 300x84 intrinsic asset. Reference: No artifact reference; described-only | The logo image omits intrinsic dimensions
|
browser-platform-hygieneThe page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | DOM confirms HTML doctype, UTF-8, viewport metadata, and no permission prompt; the page uses HTTPS. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
Be discoverable · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
title-and-descriptionThe page has a unique, descriptive <title> and a meta description. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The rendered title is present, but the probe found no meta description. Reference: No artifact reference; described-only | The anti-bot response has no meta description
|
crawlable-and-mobile-friendlyLinks are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The discoverability primitive found only 4% rendered-word coverage in raw HTML; title and H1 are absent from raw HTML. Reference: No artifact reference; described-only | The rendered challenge is almost absent without JavaScript
|
canonical-and-indexing-signalsPublic pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The probe found no canonical link and no robots meta on a 200 response that renders an anti-bot interstitial. Reference: No artifact reference; described-only | Indexing and sharing metadata are missing
|
structured-and-shareable-metadataWhere the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The probe found zero JSON-LD blocks and zero Open Graph metadata. Reference: No artifact reference; described-only | Indexing and sharing metadata are missing
|
Be private and secure · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
secure-transport-and-headersServed over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses. Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence. | Retained evidence records 3 of 6 baseline security-header categories present. Cookie-attribute review records 1 of 1 records with Secure and 0 with HttpOnly. Names, values, routes, and raw headers remain private. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
data-minimisation-and-third-partiesNo over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses. Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence. | Retained evidence records 1 first-party origin categories and 3 third-party origin categories across 8 requests. Destinations, identifiers, routes, headers, and bodies remain private. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
in-context-permissions-and-modern-authPermission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses. Method used or attempted: Permit-bound recon plus strict journey replay result | Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. Reference: No artifact reference; described-only | The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. |
defensive-browser-policiesBrowser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses. Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence. | Retained evidence records 3 of 6 baseline browser-policy header categories present. Values and raw headers remain private. Reference: No artifact reference; described-only | Defensive response policies are incomplete
|
Be resilient · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
progressive-enhancementCore content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The no-JavaScript/raw comparison exposes only 4% of rendered content words and omits the rendered title and H1. Reference: No artifact reference; described-only | The rendered challenge is almost absent without JavaScript
|
resilient-runtime-behaviourThe page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The challenge rendered consistently across desktop, mobile, reduced-motion, and crawler captures without runtime-visible breakage. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
offline-and-installableWhere the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.) | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Installability is not applicable to this transient anti-bot response, and the intended Reddit application was blocked. Reference: No artifact reference; described-only | Installability is not applicable to this transient anti-bot response, and the intended Reddit application was blocked. |
network-and-http-failure-statesHTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses. Method used or attempted: Permit-bound recon plus strict journey replay result | Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. Reference: No artifact reference; described-only | The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. |
Be internationalised · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
lang-dir-and-logical-propertiesCorrect lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The document declares lang=en; observed text is consistently left-to-right and mobile CSS adapts without clipping. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
locale-aware-dataDates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The observed response presents no locale-sensitive dates, numbers, currency, or calendars. Reference: No artifact reference; described-only | The observed response presents no locale-sensitive dates, numbers, currency, or calendars. |
time-zone-correctnessTime handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The observed response presents no time-zone-dependent data or scheduling. Reference: No artifact reference; described-only | The observed response presents no time-zone-dependent data or scheduling. |
Be trustworthy · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-dark-patternsNo deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive). | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The wall plainly explains why the challenge exists and presents no deceptive choices, confirmshaming, or commercial pressure. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
humane-error-handlingForms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. Method used or attempted: Permit-bound recon plus strict journey replay result | Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. Reference: No artifact reference; described-only | The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable. |
trustworthy-input-assistanceInput is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The only page form control is reCAPTCHA implementation plumbing; there is no address, payment, sign-in, or signup field to autofill. Reference: No artifact reference; described-only | The only page form control is reCAPTCHA implementation plumbing; there is no address, payment, sign-in, or signup field to autofill. |
safe-commercial-and-account-flowsCheckout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | No commercial, subscription, checkout, or account-management flow is exposed on the observed response. Reference: No artifact reference; described-only | No commercial, subscription, checkout, or account-management flow is exposed on the observed response. |
Be sustainable · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
optimised-assetsImages and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The one page-owned visual is an inline SVG logo; no raster transfer is used for it. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
no-wasteful-workBackground work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Trace shows no long tasks and the reduced-motion capture produced no continuing animation activity. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
third-party-and-media-budgetThird-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | HAR summary records 67% of transfer bytes from third parties, dominated by 340,589 bytes of reCAPTCHA JavaScript for a minimal wall. Reference: No artifact reference; described-only | The minimal wall transfers half a megabyte, mostly third-party JavaScript
|
Be agent ready · 2 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
structured-agent-capabilitiesWhere it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The anti-bot response is not an agent-capability surface, and the intended Reddit content could not be reached. Reference: No artifact reference; described-only | The anti-bot response is not an agent-capability surface, and the intended Reddit content could not be reached. |
on-device-inferenceOn-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The observed challenge has no inference use case or built-in AI surface. Reference: No artifact reference; described-only | The observed challenge has no inference use case or built-in AI surface. |
Be memory-efficient · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-leak-under-repeated-interactionRepeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | The document has no scroll range and the strict journey did not permit a repeatable interaction; fabricating one would not test a real user action. Reference: No artifact reference; described-only | The document has no scroll range and the strict journey did not permit a repeatable interaction; fabricating one would not test a real user action. |
bounded-footprintHeap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | Baseline heap summary is 6,490,062 bytes for 105,620 nodes, proportionate to the challenge plus reCAPTCHA. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
no-detached-dom-or-unbounded-listenersNo growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable | After ten bounded scroll attempts, heap growth was only 11,487 bytes and 26 nodes; no Detached* constructor appears in either top-constructor summary. Reference: No artifact reference; described-only | The retained evidence directly supported this check under the captured conditions. |
5. https://www.amazon.com · 58 slots · available
Respect user preferences · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
respects-color-schemeHonours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The dark-mode probe reports body rgb(255, 255, 255), text rgb(15, 17, 17), and color-scheme normal while the dark media query matches. Reference: page-probe-summary, screenshot; retained-private | The homepage remains light when the user requests dark mode.
|
respects-reduced-motionHonours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | Reduced-motion emulation matched and the page exposed zero active animations with automatic scroll behavior. Reference: layout-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
respects-contrastHonours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | The high-contrast screenshot retained readable navigation, promotional text, and controls without disappearing content. Reference: other-private-evidence; retained-private | The retained evidence directly supported this check under the captured conditions. |
Implement natural interactions · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
view-transitionsState and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | The only replayed baseline triggered an automatic POST and the strict journey aborted; no safe state transition was available to inspect. Reference: No artifact reference; described-only | The only replayed baseline triggered an automatic POST and the strict journey aborted; no safe state transition was available to inspect. |
scroll-driven-animationsScroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | The bounded-scroll journey step was skipped after the baseline mutation blocker, so scroll-linked behavior could not be exercised safely. Reference: No artifact reference; described-only | The bounded-scroll journey step was skipped after the baseline mutation blocker, so scroll-linked behavior could not be exercised safely. |
physical-gesturesGesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | No permitted reversible gesture surface could be exercised after the strict journey aborted. Reference: No artifact reference; described-only | No permitted reversible gesture surface could be exercised after the strict journey aborted. |
Provide guided navigation · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
scroll-state-aware-chromeSticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | The planned bounded scroll was skipped after the automatic POST blocker. Reference: No artifact reference; described-only | The planned bounded scroll was skipped after the automatic POST blocker. |
anchored-positioningTooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Opening menus or overlays was outside the already replayed journey and no additional control interaction was permitted. Reference: No artifact reference; described-only | Opening menus or overlays was outside the already replayed journey and no additional control interaction was permitted. |
directs-attentionNavigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | The strict journey could not advance beyond baseline, so post-navigation attention movement was not observable. Reference: No artifact reference; described-only | The strict journey could not advance beyond baseline, so post-navigation attention movement was not observable. |
Maximize content, reduce noise · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-intrusive-interruptionsNo intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | Independent desktop and mobile load screenshots showed the page content directly with no modal interstitial obscuring it. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
semantic-dismissible-primitivesOverlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The DOM probe found many div role=dialog nodes but zero dialog, popover, or details elements. Reference: page-probe-summary; retained-private | Overlay semantics rely on custom role-based containers rather than native primitives.
|
reduced-chromeMinimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | Desktop and mobile screenshots devote the first viewport to two navigation rows and large promotional panels; the mobile capture reaches footer chrome with little product-detail content. Reference: screenshot; retained-private | Navigation, promotional chrome, location messaging, and dense merchandising compete with the shopping content.
|
Adapt to the form factor · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
responsive-no-horizontal-scrollLayout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The mobile layout reports no viewport meta and 620 px horizontal overflow at 360x800; desktop also reports 220 px overflow. The mobile screenshot is visibly clipped to a desktop-width surface. Reference: layout-summary, screenshot; retained-private | The entry page does not establish a mobile viewport and overflows horizontally.
|
component-level-responsivenessComponents adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Rendered evidence proves page-level overflow, but authored CSS and safe component resizing were unavailable, so container-level behavior could not be conclusively judged. Reference: No artifact reference; described-only | Rendered evidence proves page-level overflow, but authored CSS and safe component resizing were unavailable, so container-level behavior could not be conclusively judged. |
input-modality-awareTouch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The focus probe reports 0 px outline on the Amazon home link, delivery control, search field, search submit, language, account, orders, and cart controls. Reference: page-probe-summary; retained-private | Keyboard focus is not visibly exposed on several primary controls.
|
Support core task success · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
clear-purpose-and-primary-actionThe page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | The Amazon identity, prominent search control, category navigation, and shopping promotions make the commerce purpose and next action clear. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
primary-flow-completionThe representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Search, account, product, cart, and checkout routes were outside the exact-URL permit, and the strict baseline aborted on an automatic POST. Reference: No artifact reference; described-only | Search, account, product, cart, and checkout routes were outside the exact-URL permit, and the strict baseline aborted on an automatic POST. |
clear-system-state-and-recoveryLoading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Failure, empty, success, and retry states require prohibited route or mutation testing and were not safely reachable. Reference: No artifact reference; described-only | Failure, empty, success, and retry states require prohibited route or mutation testing and were not safely reachable. |
Be fast and stable · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
good-core-web-vitalsCore Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | The trace did not yield FCP or LCP and no permitted interaction was available for an INP-like measurement. Reference: No artifact reference; described-only | The trace did not yield FCP or LCP and no permitted interaction was available for an INP-like measurement. |
visual-stabilityNo cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | Desktop and mobile layout observers recorded CLS 0 with no shift entries during their capture windows. Reference: layout-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
efficient-main-threadThe main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | The performance trace recorded zero long tasks and 0 ms total blocking time over [host/path omitted] seconds. Reference: performance-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
efficient-resource-deliveryCritical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The HAR records 43 requests and 1,101,460 transferred bytes, including ten parser-inserted render-blocking candidates and a 286,614-byte WAF script. Reference: network-summary; retained-private | The landing load has a long render-blocking resource chain and a large anti-bot script cost.
|
trim-unused-and-duplicate-codeThe page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | No permit-aware code-coverage evidence was gathered, so shipped bytes cannot be classified as used or unused. Reference: No artifact reference; described-only | No permit-aware code-coverage evidence was gathered, so shipped bytes cannot be classified as used or unused. |
Be inclusive · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
names-roles-labelsInteractive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The accessibility probe found 56 unnamed controls among 268 controls, while the image audit found three images without alt in its captured variant. Reference: image-summary, page-probe-summary; retained-private | A material number of interactive controls lack a detectable accessible name.
|
sufficient-contrastText and essential UI meet WCAG colour-contrast minimums against their background. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | Desktop, mobile, and prefers-contrast screenshots show legible text and controls against their immediate surfaces; no visual contrast failure was observed. Reference: other-private-evidence, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
structure-and-focusHeading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The probe reports zero H1 elements and 0 px outlines on primary navigation and search controls. Reference: page-probe-summary; retained-private | The page has no H1 and several major controls suppress visible focus.
|
legible-textText is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | Captured headings, navigation labels, and body text are readable and not clipped vertically in the inspected states. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
zoom-reflow-targets-and-mediaThe experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | At 360x800 the layout viewport remains 980 px wide and horizontal overflow is 620 px, making content clipped and zoomed out. Reference: layout-summary, screenshot; retained-private | Narrow-screen reflow fails because the page omits a viewport declaration.
|
Follow best practices · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-console-errorsThe page loads without console errors or uncaught exceptions. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | The journey console dimension was blocked by the unavailable Stage 2 collector and no independent console-event capture was retained. Reference: No artifact reference; described-only | The journey console dimension was blocked by the unavailable Stage 2 collector and no independent console-event capture was retained. |
sound-document-and-assetsValid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.) | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The image audit found 11 of 12 captured images without width/height, two oversized images, seven without srcset, and eight legacy-format images. Reference: image-summary; retained-private | Image sizing and delivery are structurally fragile.
|
browser-platform-hygieneThe page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Deprecation, BFCache, source-map, prompt timing, and vulnerable-library checks require evidence not available within the bounded entry-only run. Reference: No artifact reference; described-only | Deprecation, BFCache, source-map, prompt timing, and vulnerable-library checks require evidence not available within the bounded entry-only run. |
Be discoverable · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
title-and-descriptionThe page has a unique, descriptive <title> and a meta description. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | The rendered probe found a descriptive title and a detailed meta description. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
crawlable-and-mobile-friendlyLinks are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The discoverability probe reports 1% raw-to-rendered content coverage and isJsShell=true; the page also lacks viewport meta. Reference: discoverability-summary, layout-summary; retained-private | The rendered landing page is not mobile-friendly and most content is absent from the raw HTML response.
|
canonical-and-indexing-signalsPublic pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | Discoverability evidence records fetchedStatus 202, 1% coverage, no title in raw HTML, and no raw meta description, despite a rendered canonical URL. Reference: discoverability-summary; retained-private | The non-JavaScript fetch returns an HTTP 202 challenge shell rather than indexable page content.
|
structured-and-shareable-metadataWhere the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The probe found two Open Graph tags but no JSON-LD blocks for the organization or commerce destination. Reference: page-probe-summary; retained-private | The homepage exposes limited entity metadata.
|
Be private and secure · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
secure-transport-and-headersServed over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses. Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence. | Retained evidence records 0 of 6 baseline security-header categories present. Cookie-attribute review records 3 of 7 records with Secure and 0 with HttpOnly. Names, values, routes, and raw headers remain private. Reference: cookie-attribute-summary, security-header-summary; retained-private | The captured main response lacks baseline defensive headers and several cookies have weak attributes.
|
data-minimisation-and-third-partiesNo over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses. Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence. | Retained evidence records 1 first-party origin categories and 3 third-party origin categories across 43 requests. Destinations, identifiers, routes, headers, and bodies remain private. Reference: network-summary, tracker-summary; retained-private | Most landing-load requests and bytes are delivered from third-party-classified origins.
|
in-context-permissions-and-modern-authPermission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Authentication is on an unpermitted path and no permission prompt fired on the captured landing state; modern-auth support could not be verified. Reference: No artifact reference; described-only | Authentication is on an unpermitted path and no permission prompt fired on the captured landing state; modern-auth support could not be verified. |
defensive-browser-policiesBrowser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses. Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence. | Retained evidence records 0 of 6 baseline browser-policy header categories present. Values and raw headers remain private. Reference: security-header-summary; retained-private | Browser-enforced defense headers are absent in the captured response.
|
Be resilient · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
progressive-enhancementCore content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The raw response has 14 content tokens versus 304 rendered tokens, 1% coverage, no raw title, and no raw description; the crawler screenshot is a challenge/blank shell. Reference: discoverability-summary; retained-private | Core landing content is effectively unavailable without JavaScript.
|
resilient-runtime-behaviourThe page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Menus and async state could not be exercised because the strict replay aborted after baseline network mutation. Reference: No artifact reference; described-only | Menus and async state could not be exercised because the strict replay aborted after baseline network mutation. |
offline-and-installableWhere the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.) | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. Method used or attempted: Applicability judgement from the audited commerce entry state. | Offline/installability is contextual and not essential to this public commerce landing page; progressive enhancement is judged separately. Reference: No artifact reference; described-only | Offline/installability is contextual and not essential to this public commerce landing page; progressive enhancement is judged separately. |
network-and-http-failure-statesHTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Network failure simulation and error routes were outside the bounded execution plan. Reference: No artifact reference; described-only | Network failure simulation and error routes were outside the bounded execution plan. |
Be internationalised · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
lang-dir-and-logical-propertiesCorrect lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | The rendered document declares lang=en-us and computes ltr direction; the captured English content follows that direction. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
locale-aware-dataDates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. Method used or attempted: Applicability judgement from the audited commerce entry state. | No date, currency, duration, or other locale-sensitive value was present in the audited entry state. Reference: No artifact reference; described-only | No date, currency, duration, or other locale-sensitive value was present in the audited entry state. |
time-zone-correctnessTime handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. Method used or attempted: Applicability judgement from the audited commerce entry state. | No time or event data was present in the audited entry state. Reference: No artifact reference; described-only | No time or event data was present in the audited entry state. |
Be trustworthy · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-dark-patternsNo deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive). | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | No consent wall, confirmshaming, forced continuity, disguised ad control, or purchase prompt appeared on the audited entry state. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
humane-error-handlingForms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Form mutation and submission are hard-denied by the journey protocol. Reference: No artifact reference; described-only | Form mutation and submission are hard-denied by the journey protocol. |
trustworthy-input-assistanceInput is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Account, address, and payment forms were outside the exact-URL boundary, so autofill assistance could not be assessed. Reference: No artifact reference; described-only | Account, address, and payment forms were outside the exact-URL boundary, so autofill assistance could not be assessed. |
safe-commercial-and-account-flowsCheckout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses. Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey. | Cart, checkout, purchase, authentication, account management, and cancellation are hard-denied or outside the exact-URL permit. Reference: No artifact reference; described-only | Cart, checkout, purchase, authentication, account management, and cancellation are hard-denied or outside the exact-URL permit. |
Be sustainable · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
optimised-assetsImages and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The image audit reports two oversized images, seven without srcset, eight legacy-format images, and no modern-format image in the captured set. Reference: image-summary; retained-private | Captured image delivery omits responsive sizing and modern formats.
|
no-wasteful-workBackground work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | The flow baseline was aborted after an automatic WAF verification POST attempt; the independent HAR still recorded 43 requests and [host/path omitted] MB transferred. Reference: journey-summary, network-summary; retained-private | The baseline immediately starts substantial challenge and telemetry work.
|
third-party-and-media-budgetThird-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses. Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited. | Third-party-classified traffic accounts for 794,498 bytes and 37 requests; scripts alone transfer 449,270 bytes. Reference: network-summary; retained-private | Third-party-classified delivery dominates the landing-page transfer budget.
|
Be agent ready · 2 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
structured-agent-capabilitiesWhere it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. Method used or attempted: Applicability judgement from the audited commerce entry state. | Agentic commerce is an emerging opportunity, but no declared agent-facing intent or surface was available; absence is not penalized. Reference: No artifact reference; described-only | Agentic commerce is an emerging opportunity, but no declared agent-facing intent or surface was available; absence is not penalized. |
on-device-inferenceOn-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses. Method used or attempted: Applicability judgement from the audited commerce entry state. | No on-device inference use case or declared intent was present on the commerce landing state. Reference: No artifact reference; described-only | No on-device inference use case or declared intent was present on the commerce landing state. |
Be memory-efficient · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-leak-under-repeated-interactionRepeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | After ten bounded scroll-to-500-and-back cycles, summarized heap size rose only 200,845 bytes ([host/path omitted]%) and node count 1,851 ([host/path omitted]%) across independent captures, with no unbounded-growth signal. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
bounded-footprintHeap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | The rendered commerce landing state used about [host/path omitted] MB summarized self-size; this is substantial but proportionate to a dense, script-heavy homepage and did not grow materially in the bounded repeat. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
no-detached-dom-or-unbounded-listenersNo growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses. Method used or attempted: Permit-bound evidence review for the entry state. | No Detached* constructor appeared among top heap constructors, and closure count changed by only 74 across the repeated-scroll pair; no accumulating pattern was observed. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
6. https://en.wikipedia.org · 58 slots · available
Method warning: this site has a report, but its no-console-errors pass is method-invalid and is labelled in the table.
Respect user preferences · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
respects-color-schemeHonours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | Desktop and prefers-color-scheme: dark screenshots are byte-identical; computed color-scheme is normal and the rendered body remains rgb(248,249,250). Reference: page-probe-summary, screenshot; retained-private | The page does not follow the system dark-color preference by default.
|
respects-reduced-motionHonours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Reduced-motion probe matched the preference, reported scroll-behavior auto, and found zero active animations; CSS also contains a reduced-motion rule. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
respects-contrastHonours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | The prefers-contrast: more screenshot retains visible dark text, blue links, borders, and controls across the full page. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Implement natural interactions · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
view-transitionsState and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | CSS probe found no view-transition rules or declarations; the MPA exposes hundreds of same-origin article links. Reference: journey-summary, page-probe-summary; retained-private | Same-origin document navigation has no View Transition enhancement.
|
scroll-driven-animationsScroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | The audited main page and bounded-scroll journey contain no parallax, scrollytelling, reveal, or other scroll-linked animation to implement. |
physical-gesturesGesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | The recorded bounded wheel scroll moved 246 CSS px using native scrolling, with no pointer interception, unexpected state change, or network request. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Provide guided navigation · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
scroll-state-aware-chromeSticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | The bounded-scroll screenshot shows content moving predictably without a static header covering the reading area; no JS scroll toggling was detected. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
anchored-positioningTooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | No tooltip, popover, or menu overlay was opened or present in the bounded journey, so there is no positioned overlay to assess. |
directs-attentionNavigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | The baseline has a clear selected Main Page tab, section headings, blue link affordances, and a visible skip link focus treatment. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Maximize content, reduce noise · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-intrusive-interruptionsNo intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Baseline and full-page screenshots contain no load-time modal, interstitial, consent wall, or banner obscuring content. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
semantic-dismissible-primitivesOverlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | DOM inspection found no active overlay in the audited state; existing controls use native form and button elements rather than a visible ad-hoc modal. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
reduced-chromeMinimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Screenshots show the viewport dominated by encyclopedia content with compact navigation and no decorative application frame. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Adapt to the form factor · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
responsive-no-horizontal-scrollLayout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | At requested 360x800, layout reports a 1120 CSS-pixel client width and the DOM declares meta viewport width=1120; the full-page mobile capture shows desktop columns and tiny text. Reference: layout-summary, page-probe-summary, screenshot; retained-private | The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.
|
component-level-responsivenessComponents adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | At requested 360x800, layout reports a 1120 CSS-pixel client width and the DOM declares meta viewport width=1120; the full-page mobile capture shows desktop columns and tiny text. Reference: layout-summary, page-probe-summary, screenshot; retained-private | The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.
|
input-modality-awareTouch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | The focus/target probe measured Donate, Create account, Log in and many content links at 15 to 17 px high, below a comfortable touch-target size. Reference: page-probe-summary, screenshot; retained-private | Many important navigation links have touch boxes only 15 to 17 CSS pixels high.
|
Support core task success · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
clear-purpose-and-primary-actionThe page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | The first viewport clearly says “Welcome to Wikipedia”, exposes search/navigation, and immediately presents the featured article. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
primary-flow-completionThe representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Both supplied journey actions completed: same-origin baseline load and bounded scroll, with 2 planned and 2 performed actions and no blocked action. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
clear-system-state-and-recoveryLoading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | The supplied non-mutating baseline and bounded-scroll journey has no loading, empty, error, success, or partial-completion state. |
Be fast and stable · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
good-core-web-vitalsCore Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Permit-bound trace measured LCP/FCP [host/path omitted] ms and TBT [host/path omitted] ms; layout measured CLS [host/path omitted], all within good lab ranges. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
visual-stabilityNo cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Layout observers measured CLS [host/path omitted] desktop and 0 mobile, with only one very small desktop shift. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
efficient-main-threadThe main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Trace found one [host/path omitted] ms long task and [host/path omitted] ms total blocking time, while the page reached LCP in [host/path omitted] s. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
efficient-resource-deliveryCritical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | Image audit found 16 below-fold images without loading=lazy, 17 legacy-format images, 3 oversized images, and one missing dimension; HAR lists 10 images without cache headers, including a 129 KB portrait. Reference: image-summary, layout-summary, network-summary; retained-private | Below-fold imagery is eagerly loaded and several image delivery details waste bytes or weaken stability.
|
trim-unused-and-duplicate-codeThe page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | HAR transferred 682,131 bytes total with 281,579 script bytes and no third-party tracker package; this is proportionate for the feature-rich reference landing page. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Be inclusive · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
names-roles-labelsInteractive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Probe found zero unlabeled visible controls; all images carry an alt attribute and the page exposes 13 landmarks. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
sufficient-contrastText and essential UI meet WCAG colour-contrast minimums against their background. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Desktop and high-contrast screenshots retain legible dark text, blue links, controls, and section boundaries; no content disappears in the preference condition. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
structure-and-focusHeading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Visible heading sequence is one H1 followed by eight H2s with no skipped levels; first keyboard target has a 2 px solid focus outline. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
legible-textText is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | At requested 360x800, layout reports a 1120 CSS-pixel client width and the DOM declares meta viewport width=1120; the full-page mobile capture shows desktop columns and tiny text. Reference: layout-summary, page-probe-summary, screenshot; retained-private | The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.
|
zoom-reflow-targets-and-mediaThe experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | At requested 360x800, layout reports a 1120 CSS-pixel client width and the DOM declares meta viewport width=1120; the full-page mobile capture shows desktop columns and tiny text. The focus/target probe measured Donate, Create account, Log in and many content links at 15 to 17 px high, below a comfortable touch-target size. Reference: layout-summary, page-probe-summary, screenshot; retained-private | The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down. Many important navigation links have touch boxes only 15 to 17 CSS pixels high.
|
Follow best practices · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-console-errorsThe page loads without console errors or uncaught exceptions. | Pass Source: pass; confidence: mediumMethod invalid: this console pass used absence of a surfaced error even though the authoritative console collector was unavailable. Treat it as unreliable evidence, not a valid pass. | Implementation and methodCatalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | All permit-bound DOM, layout, trace, HAR, evaluate, and journey navigations completed without an uncaught exception surfaced by the harness. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
sound-document-and-assetsValid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.) | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | Image audit found 16 below-fold images without loading=lazy, 17 legacy-format images, 3 oversized images, and one missing dimension; HAR lists 10 images without cache headers, including a 129 KB portrait. Reference: image-summary, layout-summary, network-summary; retained-private | Below-fold imagery is eagerly loaded and several image delivery details waste bytes or weaken stability.
|
browser-platform-hygieneThe page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | The document uses HTML5, no inline event-handler attributes were found, and no permission prompt appeared during repeated permit-bound loads. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Be discoverable · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
title-and-descriptionThe page has a unique, descriptive <title> and a meta description. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | DOM and discoverability probes both report metaDescription null, although the title survives without JavaScript. Reference: discoverability-summary, page-probe-summary; retained-private | The Main Page has no meta description.
|
crawlable-and-mobile-friendlyLinks are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | At requested 360x800, layout reports a 1120 CSS-pixel client width and the DOM declares meta viewport width=1120; the full-page mobile capture shows desktop columns and tiny text. Reference: layout-summary, page-probe-summary, screenshot; retained-private | The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.
|
canonical-and-indexing-signalsPublic pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Fetch returned 200 after the canonical root redirect; rendered DOM declares canonical [origin omitted] and robots allows previews. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
structured-and-shareable-metadataWhere the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | The page contains one JSON-LD block plus matching og:title and og:image metadata for the visible Main Page content. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Be private and secure · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
secure-transport-and-headersServed over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses. Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence. | Retained evidence records 3 of 6 baseline security-header categories present. Cookie-attribute review records 6 of 7 records with Secure and 4 with HttpOnly. Names, values, routes, and raw headers remain private. Reference: cookie-attribute-summary, security-header-summary; retained-private | Browser defenses and one session-like cookie are weaker than expected.
|
data-minimisation-and-third-partiesNo over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses. Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence. | Retained evidence records 1 first-party origin categories and 2 third-party origin categories across 40 requests. Destinations, identifiers, routes, headers, and bodies remain private. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
in-context-permissions-and-modern-authPermission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | The public anonymous main-page journey requests no browser permission and contains no authentication step. |
defensive-browser-policiesBrowser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses. Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence. | Retained evidence records 3 of 6 baseline browser-policy header categories present. Values and raw headers remain private. Reference: cookie-attribute-summary, security-header-summary; retained-private | Browser defenses and one session-like cookie are weaker than expected.
|
Be resilient · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
progressive-enhancementCore content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Discoverability measured 98% content coverage without JavaScript, isJsShell false, and matching browser/crawler views. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
resilient-runtime-behaviourThe page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | The recorded load and scroll retained the same URL, 2,375-node state and 142 listeners, with no overlay clipping or state loss. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
offline-and-installableWhere the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.) | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | Wikipedia Main Page is a server-rendered reference document rather than an installable application; PWA installability is not needed for this path. |
network-and-http-failure-statesHTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | The reviewed exact-URL journey contains no client-side mutation or async task with an application-owned retry/error state; testing a different error URL was explicitly prohibited. |
Be internationalised · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
lang-dir-and-logical-propertiesCorrect lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | DOM reports html lang=en and dir=ltr; CSS inspection found logical inline/block properties and the page links to 347 language editions. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
locale-aware-dataDates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | This English reference landing page does not perform locale-sensitive number, currency, duration, or calendar calculations in the supplied journey. |
time-zone-correctnessTime handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | The supplied journey exposes no scheduled event or time-zone-sensitive transaction. |
Be trustworthy · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-dark-patternsNo deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive). | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Screenshots show no consent wall, disguised ad, confirmshaming, forced continuity, or hidden commitment; controls and links use direct labels. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
humane-error-handlingForms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | The bounded non-mutating journey does not submit a form or create an invalid-input state. |
trustworthy-input-assistanceInput is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | The only relevant visible input is encyclopedia search, for which address, payment, sign-in, and sign-up autocomplete tokens do not apply. |
safe-commercial-and-account-flowsCheckout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | No checkout, subscription, consent, authentication, or account-management flow appears in the supplied anonymous journey. |
Be sustainable · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
optimised-assetsImages and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited. | Image audit found 16 below-fold images without loading=lazy, 17 legacy-format images, 3 oversized images, and one missing dimension; HAR lists 10 images without cache headers, including a 129 KB portrait. Reference: image-summary, layout-summary, network-summary; retained-private | Below-fold imagery is eagerly loaded and several image delivery details waste bytes or weaken stability.
|
no-wasteful-workBackground work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | After the bounded scroll, no network requests occurred; trace and journey metrics show no ongoing animation or increasing task work. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
third-party-and-media-budgetThird-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | HAR totals 682 KB, has no audio/video/autoplay media and no known trackers; third-party bytes are 253 KB of Wikimedia assets/auth. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Be agent ready · 2 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
structured-agent-capabilitiesWhere it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | No developer intent to expose action tools to agents was declared; high server-rendered content coverage already supports read-only agents without making emerging WebMCP mandatory. |
on-device-inferenceOn-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses. Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey. | No applicable surface exists in the audited path. Reference: No artifact reference; described-only | The static encyclopedia landing page has no summarisation or generation interaction for which built-in on-device inference is necessary. |
Be memory-efficient · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-leak-under-repeated-interactionRepeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | After ten scroll-down/up repetitions, heap self size fell from 13,165,795 to 12,918,423 bytes and heap node count fell from 190,046 to 167,942. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
bounded-footprintHeap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Baseline heap was [host/path omitted] MB self size; journey JSHeapUsedSize was about [host/path omitted] MB for a content-rich 2,375-element page. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
no-detached-dom-or-unbounded-listenersNo growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses. Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence. | Heap constructor summaries show no growing Detached* population; flow listeners stayed at 142 and DOM node count at 2,375 before/after scroll. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
7. https://www.nytimes.com · 58 slots · available
Respect user preferences · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
respects-color-schemeHonours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method. Method used or attempted: multi-modal evidence review | Dark emulation retained a white surface. Reference: page-probe-summary, screenshot; retained-private | The homepage stays light when the user requests dark mode.
|
respects-reduced-motionHonours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses. Method used or attempted: permit-bound evidence review | The reduced-motion query matched; only three finished 300ms animations remained and the CSS capture contains reduced-motion handling. Reference: page-probe-summary, video; retained-private | The retained evidence directly supported this check under the captured conditions. |
respects-contrastHonours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. Method used or attempted: permit-bound evidence review | Forced-colors/high-contrast capture preserved visible text, controls, borders, and links. Reference: other-private-evidence; retained-private | The retained evidence directly supported this check under the captured conditions. |
Implement natural interactions · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
view-transitionsState and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps. | Blocked Source: blocked; confidence: low | Implementation and methodCatalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. Method used or attempted: strict journey replay review | The strict supplied journey aborted before any state or route transition could be exercised. Reference: journey-summary; retained-private | A blocked POST caused the journey to stop before a transition state. |
scroll-driven-animationsScroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. Method used or attempted: permit-bound evidence review | No scroll-linked animation or scrollytelling surface is exposed in the exact homepage state. Reference: page-probe-summary; retained-private | The audited state has no scroll-linked visual effect requiring a declarative timeline. |
physical-gesturesGesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. Method used or attempted: permit-bound evidence review | No gesture-driven removal, pull, swipe, or comparable physical interaction is exposed on the exact homepage state. Reference: screenshot; retained-private | The audited homepage state has no representative gesture surface. |
Provide guided navigation · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
scroll-state-aware-chromeSticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. Method used or attempted: strict journey replay review | The supplied bounded-scroll state was skipped, so responsive scroll chrome could not be observed in that reviewed path. Reference: journey-summary; retained-private | The strict journey stopped before its bounded-scroll action. |
anchored-positioningTooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned. | Blocked Source: blocked; confidence: low | Implementation and methodCatalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. Method used or attempted: permit-bound evidence review | No tooltip or menu could be safely opened in the supplied strict journey after it aborted. Reference: journey-summary; retained-private | The strict journey stopped before overlay interactions. |
directs-attentionNavigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved. | Blocked Source: blocked; confidence: low | Implementation and methodCatalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. Method used or attempted: permit-bound evidence review | No in-page jump or disclosure state was reached after the journey aborted. Reference: journey-summary; retained-private | The strict journey stopped before a focus-moving navigation action. |
Maximize content, reduce noise · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-intrusive-interruptionsNo intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. Method used or attempted: multi-modal evidence review | The consent dialog obscures nearly all primary content. Reference: screenshot; retained-private | A privacy alert dialog obscures the news on every fresh load.
|
semantic-dismissible-primitivesOverlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. Method used or attempted: permit-bound evidence review | The privacy surface uses a real dialog plus role=alertdialog and aria-modal=true, with explicit actions. Reference: page-probe-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
reduced-chromeMinimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. Method used or attempted: permit-bound evidence review | Below the modal, the publication layout is content-dense and uses restrained separators rather than a heavy app frame. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
Adapt to the form factor · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
responsive-no-horizontal-scrollLayout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. Method used or attempted: permit-bound evidence review | Desktop and 360px layouts reported zero horizontal overflow and a valid viewport meta tag. Reference: layout-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
component-level-responsivenessComponents adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. Method used or attempted: multi-modal evidence review | Captured CSS contains no container-query declarations. Reference: page-probe-summary; retained-private | No container-query implementation was found in the captured homepage CSS.
|
input-modality-awareTouch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. Method used or attempted: multi-modal evidence review | Visible focus and target-size probes found undersized and focus-invisible controls. Reference: page-probe-summary; retained-private | Many visible controls are undersized and lose a visible keyboard focus indicator.
|
Support core task success · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
clear-purpose-and-primary-actionThe page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses. Method used or attempted: multi-modal evidence review | Consent management displaces the news purpose in the first viewport. Reference: screenshot; retained-private | The first viewport prioritises consent management over the publication and its current news.
|
primary-flow-completionThe representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses. Method used or attempted: strict journey replay review | The supplied baseline action encountered a blocked POST and the bounded scroll was skipped, so the reviewed primary journey did not complete. Reference: journey-summary; retained-private | The strict journey result is partial and action 2 is skipped-after-blocker. |
clear-system-state-and-recoveryLoading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses. Method used or attempted: permit-bound evidence review | The privacy state is explicit and offers Accept all, Reject all, and Manage preferences without hiding the rejection path. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
Be fast and stable · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
good-core-web-vitalsCore Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses. Method used or attempted: multi-modal evidence review | LCP was [host/path omitted] and TBT 994ms. Reference: performance-summary; retained-private | Observed load performance misses the good LCP range and has substantial blocking work.
|
visual-stabilityNo cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. Method used or attempted: multi-modal evidence review | Mobile CLS was [host/path omitted] and most images omitted dimensions. Reference: image-summary, layout-summary; retained-private | The mobile homepage has poor observed layout stability.
|
efficient-main-threadThe main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. Method used or attempted: multi-modal evidence review | Trace captured 13 long tasks and 994ms TBT. Reference: performance-summary; retained-private | Heavy scripting blocks the main thread during load.
|
efficient-resource-deliveryCritical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses. Method used or attempted: multi-modal evidence review | HAR captured 368 requests and [host/path omitted]. Reference: network-summary; retained-private | The homepage load is exceptionally request-heavy for a content entry page.
|
trim-unused-and-duplicate-codeThe page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses. Method used or attempted: multi-modal evidence review | 259 script requests transferred [host/path omitted]. Reference: network-summary; retained-private | JavaScript and third-party code dominate transfer cost.
|
Be inclusive · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
names-roles-labelsInteractive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses. Method used or attempted: multi-modal evidence review | Empty interactive anchors and null image alternatives were observed. Reference: image-summary, page-probe-summary; retained-private | The homepage includes interactive links with no accessible text and editorial images with absent descriptions.
|
sufficient-contrastText and essential UI meet WCAG colour-contrast minimums against their background. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. Method used or attempted: permit-bound evidence review | The default consent surface is high contrast and forced-colors preserves all essential text and controls. Reference: other-private-evidence, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
structure-and-focusHeading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. Method used or attempted: multi-modal evidence review | Many focused interactive elements expose transparent or none outlines. Reference: page-probe-summary; retained-private | Keyboard focus is not consistently visible.
|
legible-textText is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. Method used or attempted: permit-bound evidence review | Desktop and mobile captures show readable type, comfortable line lengths, and no clipping in the privacy content. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
zoom-reflow-targets-and-mediaThe experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses. Method used or attempted: multi-modal evidence review | The mobile layout reflows, but several controls are 12px to 32px high and carousel buttons are 24px square. Reference: layout-summary, page-probe-summary; retained-private | Several primary and repeated controls miss an inclusive target-size floor.
|
Follow best practices · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-console-errorsThe page loads without console errors or uncaught exceptions. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. Method used or attempted: strict journey capture review | The supplied journey explicitly reports the console collector as blocked, and no pre-navigation console capture was available. Reference: journey-summary; retained-private | Required Stage 2 console collector was unavailable. |
sound-document-and-assetsValid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.) | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. Method used or attempted: multi-modal evidence review | Document basics pass, but 20 of 25 images omit intrinsic dimensions. Reference: image-summary, page-probe-summary; retained-private | Most rendered images omit intrinsic dimensions.
|
browser-platform-hygieneThe page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. Method used or attempted: multi-modal evidence review | Three requests for the same VTT resource returned 404. Reference: network-summary; retained-private | The load includes repeated failed media-caption requests.
|
Be discoverable · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
title-and-descriptionThe page has a unique, descriptive <title> and a meta description. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. Method used or attempted: permit-bound evidence review | A descriptive title and meta description are present in raw and rendered HTML. Reference: discoverability-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
crawlable-and-mobile-friendlyLinks are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. Method used or attempted: permit-bound evidence review | The page has a viewport meta tag, 95% no-JS content coverage, and 392 of 412 anchors expose hrefs. Reference: discoverability-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
canonical-and-indexing-signalsPublic pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses. Method used or attempted: permit-bound evidence review | HTTP status is 200; canonical, nine hreflang links, and no accidental noindex were observed. Reference: discoverability-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
structured-and-shareable-metadataWhere the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses. Method used or attempted: permit-bound evidence review | The rendered homepage exposes two JSON-LD blocks and five Open Graph tags. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Be private and secure · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
secure-transport-and-headersServed over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses. Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence. | Retained evidence records 5 of 6 baseline security-header categories present. Cookie-attribute review records 9 of 14 records with Secure and 3 with HttpOnly. Names, values, routes, and raw headers remain private. Reference: cookie-attribute-summary, security-header-summary; retained-private | Transport is strong, but CSP and cookie posture remain permissive.
|
data-minimisation-and-third-partiesNo over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses. Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence. | Retained evidence records 1 first-party origin categories and 27 third-party origin categories across 368 requests. Destinations, identifiers, routes, headers, and bodies remain private. Reference: network-summary, screenshot, tracker-summary; retained-private | Large advertising and measurement traffic begins before a privacy choice is made.
|
in-context-permissions-and-modern-authPermission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses. Method used or attempted: permit-bound evidence review | No browser permission prompt appeared on load; an authentication form was not part of the exact permitted page state. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
defensive-browser-policiesBrowser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses. Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence. | Retained evidence records 5 of 6 baseline browser-policy header categories present. Values and raw headers remain private. Reference: security-header-summary; retained-private | Defensive policy coverage is incomplete.
|
Be resilient · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
progressive-enhancementCore content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. Method used or attempted: permit-bound evidence review | Raw HTML returned 200 with 95% rendered-content coverage, including title, H1, and description. Reference: discoverability-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
resilient-runtime-behaviourThe page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. Method used or attempted: permit-bound evidence review | The privacy dialog remains contained and usable at desktop and 360px without overflow. Reference: layout-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
offline-and-installableWhere the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.) | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. Method used or attempted: permit-bound evidence review | The audited surface is a public news homepage, not an app-like offline workflow; progressive enhancement is judged separately and passes. Reference: discoverability-summary; retained-private | Offline installability is not required for this publication homepage archetype. |
network-and-http-failure-statesHTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states. | Not run Source: not-run; confidence: high | Implementation and methodCatalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses. Method used or attempted: permit-bound evidence review | No alternate URL or failure route was permitted, and injecting a network failure was outside the supplied bounded journey. Reference: journey-summary; retained-private | The exact-URL and strict-journey boundary did not permit a representative error-state path. |
Be internationalised · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
lang-dir-and-logical-propertiesCorrect lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. Method used or attempted: permit-bound evidence review | The page declares lang=en, offers language choices, exposes nine hreflang links, and captured CSS uses inline logical properties. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
locale-aware-dataDates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. Method used or attempted: permit-bound evidence review | The UK-observed page localises subscription currency to GBP and provides language selection. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
time-zone-correctnessTime handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. Method used or attempted: permit-bound evidence review | No event scheduling or time-zone-sensitive task is present in the exact homepage state. Reference: screenshot; retained-private | The audited homepage state exposes no time-zone-sensitive flow. |
Be trustworthy · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-dark-patternsNo deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive). | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. Method used or attempted: permit-bound evidence review | Accept all and Reject all have equal visual prominence, Manage preferences is explicit, and policy links are visible. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
humane-error-handlingForms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. Method used or attempted: permit-bound evidence review | No user-submitted form or invalid-input state is present in the exact homepage state. Reference: page-probe-summary; retained-private | The only visible choices are consent actions, not data-entry validation. |
trustworthy-input-assistanceInput is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. Method used or attempted: permit-bound evidence review | No sign-in, address, payment, or comparable text input is present in the audited state. Reference: page-probe-summary; retained-private | No user text-entry flow is present on the exact homepage state. |
safe-commercial-and-account-flowsCheckout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity. | Not run Source: not-run; confidence: high | Implementation and methodCatalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses. Method used or attempted: permit-bound evidence review | A subscription link is present, but cross-origin or different-path commercial navigation was expressly outside the exact-URL boundary. Reference: screenshot; retained-private | The commercial flow could not be opened without leaving the only permitted URL. |
Be sustainable · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
optimised-assetsImages and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. Method used or attempted: multi-modal evidence review | All 25 observed images use legacy JPG or PNG formats. Reference: image-summary; retained-private | Image delivery relies entirely on legacy formats in the observed state.
|
no-wasteful-workBackground work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. Method used or attempted: multi-modal evidence review | The initial state transfers [host/path omitted] and executes 259 script requests before content access. Reference: network-summary, performance-summary; retained-private | The initial page performs disproportionate background and third-party work.
|
third-party-and-media-budgetThird-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses. Method used or attempted: multi-modal evidence review | Third parties account for [host/path omitted] and 114 requests. Reference: network-summary, tracker-summary; retained-private | Third-party scripts consume more than half of transferred bytes.
|
Be agent ready · 2 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
structured-agent-capabilitiesWhere it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. Method used or attempted: permit-bound evidence review | The publication exposes crawlable structured content but no declared agent-action surface; emerging WebMCP capability is treated as an opportunity, not a failure. Reference: discoverability-summary, page-probe-summary; retained-private | No agent-facing transactional capability is declared for this public homepage. |
on-device-inferenceOn-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses. Method used or attempted: permit-bound evidence review | No on-device inference feature is part of the homepage purpose. Reference: page-probe-summary; retained-private | Built-in AI is not necessary for the audited homepage task. |
Be memory-efficient · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-leak-under-repeated-interactionRepeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses. Method used or attempted: permit-bound evidence review | After ten bounded down/up scroll cycles, heap self-size fell from [host/path omitted] to [host/path omitted] and node count fell from [host/path omitted] to [host/path omitted]. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
bounded-footprintHeap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses. Method used or attempted: multi-modal evidence review | Both snapshots retain roughly 116MB and [host/path omitted] heap nodes. Reference: memory-summary; retained-private | The baseline memory footprint is large for a news homepage.
|
no-detached-dom-or-unbounded-listenersNo growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses. Method used or attempted: permit-bound evidence review | The post-scroll snapshot did not grow: nodes, closures, objects, arrays, and total self-size all decreased; Detached constructors were not among the retained leaders. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
8. https://gemini.google.com · 58 slots · available
Respect user preferences · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
respects-color-schemeHonours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
respects-reduced-motionHonours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The permit-bound reduced-motion probe reported prefers-reduced-motion: reduce=true but 9 running animations, including two with 1,498,500 ms durations. Reference: other-private-evidence; retained-private | Long-running motion remains active when reduced motion is requested.
|
respects-contrastHonours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
Implement natural interactions · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
view-transitionsState and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
scroll-driven-animationsScroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
physical-gesturesGesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
Provide guided navigation · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
scroll-state-aware-chromeSticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
anchored-positioningTooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
directs-attentionNavigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
Maximize content, reduce noise · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-intrusive-interruptionsNo intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | Desktop, mobile, and pre-replayed journey screenshots show a modal covering the prompt and main content. On 360x800, the dialog fills most of the viewport and initially exposes only Read more while the decisive controls are below the fold. Reference: screenshot; retained-private | The consent dialog obscures the primary assistant on first load.
|
semantic-dismissible-primitivesOverlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: page-probe-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
reduced-chromeMinimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: page-probe-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
Adapt to the form factor · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
responsive-no-horizontal-scrollLayout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: layout-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
component-level-responsivenessComponents adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
input-modality-awareTouch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: journey-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Support core task success · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
clear-purpose-and-primary-actionThe page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: page-probe-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
primary-flow-completionThe representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
clear-system-state-and-recoveryLoading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
Be fast and stable · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
good-core-web-vitalsCore Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The permit-bound trace measured FCP 6,623 ms and LCP 6,862 ms, well outside the good LCP range; it also recorded 108 ms total blocking time. Reference: performance-summary; retained-private | Cold-load rendering is too slow.
|
visual-stabilityNo cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: layout-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
efficient-main-threadThe main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The permit-bound trace measured FCP 6,623 ms and LCP 6,862 ms, well outside the good LCP range; it also recorded 108 ms total blocking time. Reference: performance-summary; retained-private | Cold-load rendering is too slow.
|
efficient-resource-deliveryCritical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The HAR summary recorded 100 requests and 5,424,867 transferred bytes: [host/path omitted] MB of scripts and [host/path omitted] MB of fonts. Two parser-inserted high-priority scripts were render-blocking candidates. Reference: network-summary, other-private-evidence; retained-private | The initial app shell transfers a disproportionate payload.
|
trim-unused-and-duplicate-codeThe page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The HAR captured 51 script requests totalling 3,787,375 transferred bytes before any assistant task was performed, including individual bundles over 1 MB and 700 KB. Reference: other-private-evidence; retained-private | The signed-out shell eagerly loads a large script surface.
|
Be inclusive · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
names-roles-labelsInteractive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The images primitive found all 5 images without width/height attributes and 4 without alt attributes; the narrow layout still measured CLS [host/path omitted]. Reference: image-summary, layout-summary; retained-private | Rendered images omit intrinsic dimensions and several omit alt text.
|
sufficient-contrastText and essential UI meet WCAG colour-contrast minimums against their background. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
structure-and-focusHeading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: page-probe-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
legible-textText is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
zoom-reflow-targets-and-mediaThe experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
Follow best practices · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-console-errorsThe page loads without console errors or uncaught exceptions. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
sound-document-and-assetsValid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.) | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The images primitive found all 5 images without width/height attributes and 4 without alt attributes; the narrow layout still measured CLS [host/path omitted]. Reference: image-summary, layout-summary; retained-private | Rendered images omit intrinsic dimensions and several omit alt text.
|
browser-platform-hygieneThe page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
Be discoverable · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
title-and-descriptionThe page has a unique, descriptive <title> and a meta description. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
crawlable-and-mobile-friendlyLinks are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | Discoverability evidence measured only 2% rendered-word coverage in raw HTML and classified the page as isJsShell=true; the crawler screenshot contains no useful assistant content. Reference: discoverability-summary; retained-private | The public entry page is effectively a JavaScript shell for non-JS crawlers.
|
canonical-and-indexing-signalsPublic pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
structured-and-shareable-metadataWhere the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content. | Not applicable Source: not-applicable; confidence: medium | Implementation and methodCatalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses. Method used or attempted: Applicability judgement from the signed-out Gemini app archetype | The audited entry state does not present this contextual surface. Reference: No artifact reference; described-only | This contextual check is not applicable to the exact signed-out entry state and bounded non-mutating journey. |
Be private and secure · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
secure-transport-and-headersServed over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses. Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence. | Retained evidence records 4 of 6 baseline security-header categories present. Cookie-attribute review records 1 of 1 records with Secure and 1 with HttpOnly. Names, values, routes, and raw headers remain private. Reference: security-header-summary; retained-private | The document security policy leaves avoidable browser-defense gaps.
|
data-minimisation-and-third-partiesNo over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses. Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence. | Retained evidence records 1 first-party origin categories and 8 third-party origin categories across 100 requests. Destinations, identifiers, routes, headers, and bodies remain private. Reference: other-private-evidence, screenshot, tracker-summary; retained-private | The signed-out baseline contacts a broad third-party-origin surface.
|
in-context-permissions-and-modern-authPermission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
defensive-browser-policiesBrowser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses. Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence. | Retained evidence records 4 of 6 baseline browser-policy header categories present. Values and raw headers remain private. Reference: security-header-summary; retained-private | The document security policy leaves avoidable browser-defense gaps.
|
Be resilient · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
progressive-enhancementCore content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The raw document returned HTTP 200 but exposed only 2% of rendered content words without JavaScript and was classified as a JS shell. Reference: discoverability-summary; retained-private | Core public content does not progressively enhance without JavaScript.
|
resilient-runtime-behaviourThe page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
offline-and-installableWhere the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.) | Not applicable Source: not-applicable; confidence: medium | Implementation and methodCatalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. Method used or attempted: Applicability judgement from the signed-out Gemini app archetype | The audited entry state does not present this contextual surface. Reference: No artifact reference; described-only | This contextual check is not applicable to the exact signed-out entry state and bounded non-mutating journey. |
network-and-http-failure-statesHTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
Be internationalised · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
lang-dir-and-logical-propertiesCorrect lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
locale-aware-dataDates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly. | Not applicable Source: not-applicable; confidence: medium | Implementation and methodCatalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. Method used or attempted: Applicability judgement from the signed-out Gemini app archetype | The audited entry state does not present this contextual surface. Reference: No artifact reference; described-only | This contextual check is not applicable to the exact signed-out entry state and bounded non-mutating journey. |
time-zone-correctnessTime handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous. | Not applicable Source: not-applicable; confidence: medium | Implementation and methodCatalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. Method used or attempted: Applicability judgement from the signed-out Gemini app archetype | The audited entry state does not present this contextual surface. Reference: No artifact reference; described-only | This contextual check is not applicable to the exact signed-out entry state and bounded non-mutating journey. |
Be trustworthy · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-dark-patternsNo deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive). | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe | Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition. Reference: page-probe-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
humane-error-handlingForms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
trustworthy-input-assistanceInput is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
safe-commercial-and-account-flowsCheckout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
Be sustainable · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
optimised-assetsImages and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The images primitive found all 5 images without width/height attributes and 4 without alt attributes; the narrow layout still measured CLS [host/path omitted]. Reference: image-summary, layout-summary; retained-private | Rendered images omit intrinsic dimensions and several omit alt text.
|
no-wasteful-workBackground work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
third-party-and-media-budgetThird-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The baseline transferred [host/path omitted] MB, including [host/path omitted] MB from non-entry origins, to render a consent-covered assistant shell; scripts and fonts dominate. Reference: other-private-evidence, screenshot; retained-private | Initial resource use is disproportionate to the visible signed-out state.
|
Be agent ready · 2 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
structured-agent-capabilitiesWhere it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | The discoverability probe found only 2% raw-HTML word coverage and no useful crawler-visible task surface, leaving non-JS agents to infer the product from metadata. Reference: discoverability-summary; retained-private | The AI assistant entry surface is not robustly machine-readable without executing the application.
|
on-device-inferenceOn-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server. | Not applicable Source: not-applicable; confidence: medium | Implementation and methodCatalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses. Method used or attempted: Applicability judgement from the signed-out Gemini app archetype | The audited entry state does not present this contextual surface. Reference: No artifact reference; described-only | This contextual check is not applicable to the exact signed-out entry state and bounded non-mutating journey. |
Be memory-efficient · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-leak-under-repeated-interactionRepeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
bounded-footprintHeap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses. Method used or attempted: Permit-bound raw-CDP evidence and model judgement | A single permit-bound heap summary contained 763,142 nodes, 3,168,144 edges, and 36,441,757 self-size bytes before any assistant interaction. Reference: memory-summary; retained-private | The signed-out, consent-covered shell has a large baseline heap footprint.
|
no-detached-dom-or-unbounded-listenersNo growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses. Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state | The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary. Reference: journey-summary; retained-private | Exact execution policy and the partial journey prevented direct evidence for this check. |
9. https://www.netflix.com · 58 slots · available
Respect user preferences · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
respects-color-schemeHonours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method. Method used or attempted: Direct permitted evidence review against pinned guidance. | Dark and default screenshots are byte-identical, while computed color-scheme is normal. The dark visual design is hard-coded rather than preference-driven. Reference: page-probe-summary, screenshot; retained-private | The page does not react to the user color-scheme preference
|
respects-reduced-motionHonours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Reduced-motion evaluate probe reports zero animations, while the default probe reports ten running 4,000ms animations. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
respects-contrastHonours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Forced-colors/prefers-contrast screenshot keeps text, form borders, buttons, and dismiss control visible. Reference: other-private-evidence; retained-private | The retained evidence directly supported this check under the captured conditions. |
Implement natural interactions · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
view-transitionsState and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The retained DOM/CSS contains no view-transition declarations despite interactive FAQ, carousel, signup, and navigation surfaces. Reference: page-probe-summary; retained-private | State and route changes do not expose View Transition support
|
scroll-driven-animationsScroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. Method used or attempted: Contextual applicability judgement against the audited landing state. | No parallax, scrollytelling, or entry/exit reveal motion was observed; the horizontal carousel uses native scroll snap without scroll-linked animation. Reference: No artifact reference; described-only | No parallax, scrollytelling, or entry/exit reveal motion was observed; the horizontal carousel uses native scroll snap without scroll-linked animation. |
physical-gesturesGesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Retained CSS uses mandatory horizontal scroll-snap and logical scroll margins for the trending carousel. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Provide guided navigation · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
scroll-state-aware-chromeSticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The document is 3,266px tall at desktop and 3,970px at mobile, but the retained CSS has no animation-timeline or scroll-state query markers and the planned scroll state was skipped. Reference: journey-summary, layout-summary; retained-private | The long landing page has no scroll-state-aware guidance
|
anchored-positioningTooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. Method used or attempted: Contextual applicability judgement against the audited landing state. | No tooltip, popover, or menu overlay requiring anchored positioning exists in the audited landing state. Reference: No artifact reference; described-only | No tooltip, popover, or menu overlay requiring anchored positioning exists in the audited landing state. |
directs-attentionNavigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The carousel uses scroll-snap alignment and the visual hierarchy clearly highlights the current hero action. Reference: other-private-evidence, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Maximize content, reduce noise · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-intrusive-interruptionsNo intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Desktop and mobile screenshots show the cookie banner over the hero; on 360x800 it occupies roughly the top quarter of the viewport before the core task. Reference: other-private-evidence, screenshot; retained-private | The cookie banner obscures primary content on load
|
semantic-dismissible-primitivesOverlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The live DOM reports zero dialog, popover, and details elements while the cookie preference overlay and six FAQ disclosures are present. Reference: other-private-evidence, page-probe-summary; retained-private | Dismissible UI is implemented without native overlay primitives
|
reduced-chromeMinimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The first viewport is dominated by content and the core signup action, with minimal persistent chrome. Reference: other-private-evidence; retained-private | The retained evidence directly supported this check under the captured conditions. |
Adapt to the form factor · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
responsive-no-horizontal-scrollLayout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Layout evidence reports zero horizontal overflow at both 780x493 and 360x800, with viewport meta present. Reference: layout-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
component-level-responsivenessComponents adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The retained authored CSS contains no @container rule and computed container-type is normal on inspected components. Reference: page-probe-summary; retained-private | Components rely on viewport styling rather than container queries
|
input-modality-awareTouch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The focused Sign In link has a 2px solid outline; primary input and button are 56px tall and work without hover. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Support core task success · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
clear-purpose-and-primary-actionThe page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The headline, £[host/path omitted] starting price, cancellation promise, email field, and Get Started action are prominent in the first viewport. Reference: other-private-evidence; retained-private | The retained evidence directly supported this check under the captured conditions. |
primary-flow-completionThe representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses. Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy. | Signup requires entering and submitting an email, a hard-denied mutating journey action. It was not executed. Reference: No artifact reference; described-only | Signup requires entering and submitting an email, a hard-denied mutating journey action. It was not executed. |
clear-system-state-and-recoveryLoading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses. Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy. | Invalid submission and network-failure states require prohibited mutation/failure injection and were not executed. Reference: No artifact reference; described-only | Invalid submission and network-failure states require prohibited mutation/failure injection and were not executed. |
Be fast and stable · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
good-core-web-vitalsCore Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Trace reports LCP 2,[host/path omitted] and TBT [host/path omitted]; layout reports CLS [host/path omitted] desktop and 0 mobile. No field INP was claimed. Reference: layout-summary, performance-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
visual-stabilityNo cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Observed CLS is [host/path omitted] desktop and 0 mobile, both in the good range. Reference: layout-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
efficient-main-threadThe main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Trace reports two long tasks, longest [host/path omitted], with [host/path omitted] total blocking time. Reference: performance-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
efficient-resource-deliveryCritical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | HAR summary reports 53 requests and [host/path omitted] MB transferred, including [host/path omitted] MB scripts and 681 KB fonts. A 753 KB framework script and parser-inserted OneTrust script are high-priority candidates. Reference: network-summary, performance-summary; retained-private | Initial delivery is heavier than the acquisition task requires
|
trim-unused-and-duplicate-codeThe page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The HAR attributes [host/path omitted] MB of [host/path omitted] MB to non-main origins, including 341 KB reCAPTCHA and 352 KB OneTrust resources before signup interaction. Reference: network-summary; retained-private | Third-party and support code dominate initial transfer
|
Be inclusive · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
names-roles-labelsInteractive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Both email inputs have programmatic labels and autocomplete=email; visible controls have names. Missing alt values apply to decorative hero/logo images. Reference: image-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
sufficient-contrastText and essential UI meet WCAG colour-contrast minimums against their background. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Default, mobile, and forced-colors screenshots keep essential text and controls visibly distinct from their backgrounds. Reference: other-private-evidence, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
structure-and-focusHeading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The probe finds an H1 followed directly by H3 and no main or nav landmark, although keyboard focus is visibly outlined. Reference: page-probe-summary; retained-private | The document hierarchy and landmarks are incomplete
|
legible-textText is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Desktop and mobile screenshots show readable type, sensible wrapping, and no clipped primary copy. Reference: other-private-evidence, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
zoom-reflow-targets-and-mediaThe experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The viewport meta is width=device-width, initial-scale=[host/path omitted], minimum-scale=[host/path omitted], maximum-scale=[host/path omitted]. The maximum-scale restriction prevents user scaling in affected browsers. Reference: page-probe-summary, screenshot; retained-private | The viewport metadata disables pinch zoom
|
Follow best practices · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-console-errorsThe page loads without console errors or uncaught exceptions. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy. | The imported V1 journey explicitly records console capture as blocked, and no permit-bound console collector primitive was available. Reference: No artifact reference; described-only | The imported V1 journey explicitly records console capture as blocked, and no permit-bound console collector primitive was available. |
sound-document-and-assetsValid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.) | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The image audit reports 4 of 4 images without width/height and 2 oversized images; the logo natural width is 370px for an 89px display width. Reference: image-summary; retained-private | Images omit intrinsic dimensions and some are oversized
|
browser-platform-hygieneThe page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy. | BFCache, notification prompt, vulnerable library, and paste-prevention coverage was not available from the retained bounded paths. Reference: No artifact reference; described-only | BFCache, notification prompt, vulnerable library, and paste-prevention coverage was not available from the retained bounded paths. |
Be discoverable · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
title-and-descriptionThe page has a unique, descriptive <title> and a meta description. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The probe records a descriptive localized title and meta description. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
crawlable-and-mobile-friendlyLinks are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | All 30 links have hrefs and non-empty text; viewport meta is present; raw HTML returns 200. Reference: discoverability-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
canonical-and-indexing-signalsPublic pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The page succeeds with HTTP 200 at [route omitted] but the DOM probe finds no canonical URL and source inspection finds no hreflang markers. Reference: discoverability-summary, page-probe-summary; retained-private | The localized landing page has no canonical or hreflang signal
|
structured-and-shareable-metadataWhere the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The DOM includes Open Graph and Twitter metadata, but contains no JSON-LD for the organization/service represented by the page. Reference: page-probe-summary; retained-private | Share metadata exists but structured entity metadata is absent
|
Be private and secure · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
secure-transport-and-headersServed over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses. Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence. | Retained evidence records 3 of 6 baseline security-header categories present. Cookie-attribute review records 3 of 9 records with Secure and 3 with HttpOnly. Names, values, routes, and raw headers remain private. Reference: cookie-attribute-summary, security-header-summary; retained-private | Security headers and cookie flags have material gaps
|
data-minimisation-and-third-partiesNo over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses. Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence. | Retained evidence records 1 first-party origin categories and 11 third-party origin categories across 53 requests. Destinations, identifiers, routes, headers, and bodies remain private. Reference: network-summary, tracker-summary; retained-private | Extensive cross-origin code and telemetry load before task interaction
|
in-context-permissions-and-modern-authPermission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses. Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy. | Authentication was outside the permitted path; passkey support and prompt timing could not be exercised. Reference: No artifact reference; described-only | Authentication was outside the permitted path; passkey support and prompt timing could not be exercised. |
defensive-browser-policiesBrowser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses. Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence. | Retained evidence records 3 of 6 baseline browser-policy header categories present. Values and raw headers remain private. Reference: security-header-summary; retained-private | Several browser-enforced policies are absent
|
Be resilient · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
progressive-enhancementCore content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Discoverability evidence shows 90% rendered-word coverage in raw HTML, with title, H1, and description present and no JS shell. Reference: discoverability-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
resilient-runtime-behaviourThe page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The cookie banner and hero remain contained at desktop and mobile widths, and the carousel uses native scroll snap. Reference: other-private-evidence, page-probe-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
offline-and-installableWhere the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.) | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. Method used or attempted: Contextual applicability judgement against the audited landing state. | The public acquisition page is not an installed app surface; offline signup cannot complete meaningfully. Reference: No artifact reference; described-only | The public acquisition page is not an installed app surface; offline signup cannot complete meaningfully. |
network-and-http-failure-statesHTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses. Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy. | Network failure and error-route injection were outside the exact-URL bounded run. Reference: No artifact reference; described-only | Network failure and error-route injection were outside the exact-URL bounded run. |
Be internationalised · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
lang-dir-and-logical-propertiesCorrect lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | html lang is en but dir is empty; retained CSS contains many physical left/right declarations, although some logical scroll-margin-inline is used. Reference: page-probe-summary; retained-private | The localized page does not declare direction and still ships substantial physical-direction CSS
|
locale-aware-dataDates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The GB page renders pound pricing and retained source includes Intl usage rather than only hand-written locale formatting. Reference: other-private-evidence, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
time-zone-correctnessTime handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. Method used or attempted: Contextual applicability judgement against the audited landing state. | The audited landing page presents no dates, events, recurrence, or time-zone-sensitive data. Reference: No artifact reference; described-only | The audited landing page presents no dates, events, recurrence, or time-zone-sensitive data. |
Be trustworthy · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-dark-patternsNo deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive). | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Visible pricing and cancellation terms are adjacent to the CTA; Reject and Accept consent actions have equal prominence. Reference: other-private-evidence, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
humane-error-handlingForms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy. | Invalid form submission is a prohibited mutating action and was not executed. Reference: No artifact reference; described-only | Invalid form submission is a prohibited mutating action and was not executed. |
trustworthy-input-assistanceInput is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Both email fields use type=email, programmatic labels, and autocomplete=email. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
safe-commercial-and-account-flowsCheckout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses. Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy. | Subscription and account flows are hard-denied by the strict journey policy and were not executed. Reference: No artifact reference; described-only | Subscription and account flows are hard-denied by the strict journey policy and were not executed. |
Be sustainable · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
optimised-assetsImages and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Image evidence reports 2 oversized images, 2 legacy-format images, and missing intrinsic dimensions on all four img elements. Reference: image-summary; retained-private | Several first-view assets are not optimally sized or encoded
|
no-wasteful-workBackground work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | The initial load transfers a 341 KB reCAPTCHA script, 352 KB of OneTrust resources, and six logging requests before the signup action is used. Reference: network-summary, tracker-summary; retained-private | Task-specific work is eagerly loaded before user intent
|
third-party-and-media-budgetThird-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | HAR summary reports [host/path omitted] MB cross-origin transfer, 681 KB fonts, and [host/path omitted] MB scripts for a static acquisition first view. Reference: network-summary; retained-private | The initial third-party and font budget is disproportionate
|
Be agent ready · 2 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
structured-agent-capabilitiesWhere it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. Method used or attempted: Contextual applicability judgement against the audited landing state. | Agent capabilities are emerging and no agent-facing surface is intended for this consumer acquisition state. Reference: No artifact reference; described-only | Agent capabilities are emerging and no agent-facing surface is intended for this consumer acquisition state. |
on-device-inferenceOn-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses. Method used or attempted: Contextual applicability judgement against the audited landing state. | No observed landing-page task benefits from on-device inference. Reference: No artifact reference; described-only | No observed landing-page task benefits from on-device inference. |
Be memory-efficient · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-leak-under-repeated-interactionRepeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | After ten bounded scroll cycles, heap self size increased only 5,079 bytes while node and edge counts decreased. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
bounded-footprintHeap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | Baseline heap is [host/path omitted] MB self size for the media landing page and remains stable after repeated scrolling. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
no-detached-dom-or-unbounded-listenersNo growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses. Method used or attempted: Direct permitted evidence review against pinned guidance. | No Detached constructor appears in either top-constructor summary; node count falls from 621,070 to 620,495 after repeated scrolling. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
10. https://www.apple.com · 58 slots · available
Respect user preferences · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
respects-color-schemeHonours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method. Method used or attempted: Planned from a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corrob | Dark-emulation screenshot remained visually identical to the light baseline, and the CSS probe found no color-scheme or prefers-color-scheme signal. Reference: page-probe-summary, screenshot; retained-private | No user-preference dark theme
|
respects-reduced-motionHonours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses. Method used or attempted: Planned from a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model c | Under prefers-reduced-motion: reduce, 18 animations remained with the same 240 ms durations and the CSS probe found no reduced-motion rule. Reference: page-probe-summary, video; retained-private | Reduced-motion preference does not reduce animation
|
respects-contrastHonours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. Method used or attempted: Planned from a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses. | Forced-colors and prefers-contrast capture preserved visible text, controls, outlines, and button boundaries. Reference: other-private-evidence; retained-private | The retained evidence directly supported this check under the captured conditions. |
Implement natural interactions · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
view-transitionsState and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps. | Failed / issue Source: issues; confidence: medium | Implementation and methodCatalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. Method used or attempted: Planned from a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses. | The CSS probe found no view-transition usage despite multiple stateful galleries and navigational surfaces. Reference: page-probe-summary; retained-private | State changes do not use View Transitions
|
scroll-driven-animationsScroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners. | Failed / issue Source: issues; confidence: medium | Implementation and methodCatalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. Method used or attempted: Planned from source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses. | The page contains scroll and gallery motion, but the CSS probe found no animation-timeline, scroll-timeline, or view-timeline signal. Reference: page-probe-summary, screenshot; retained-private | Scroll-linked experiences do not use declarative timelines
|
physical-gesturesGesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. Method used or attempted: Planned from CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses. | The bounded wheel journey produced a native 246 CSS px scroll with no horizontal movement or gesture trapping. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
Provide guided navigation · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
scroll-state-aware-chromeSticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling. | Failed / issue Source: issues; confidence: medium | Implementation and methodCatalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. Method used or attempted: Planned from a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses. | Baseline and scrolled journey screenshots show unchanged static header chrome, and the CSS probe found no scroll-state or timeline signal. Reference: page-probe-summary, screenshot; retained-private | Header chrome does not respond to scroll state
|
anchored-positioningTooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned. | Failed / issue Source: issues; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. Method used or attempted: Planned from CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses. | Header menus and overlay surfaces are present, but the CSS probe found no anchor-name, position-anchor, or position-try usage. Reference: page-probe-summary; retained-private | Overlay positioning does not use CSS anchors
|
directs-attentionNavigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. Method used or attempted: Planned from CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses. | The initial hero has clear hierarchy and the bounded scroll carries attention directly from the CTA area into the product visual without disorientation. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
Maximize content, reduce noise · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-intrusive-interruptionsNo intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. Method used or attempted: Planned from a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses. | A fixed country or region chooser is present on load and consumes a substantial portion of the first viewport before the product content. Reference: page-probe-summary, screenshot; retained-private | Locale chooser dominates the first viewport
|
semantic-dismissible-primitivesOverlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. Method used or attempted: Planned from DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses. | The load-time locale surface is a custom fixed ASIDE; the probe found zero dialog, popover, or details primitives. Reference: page-probe-summary; retained-private | Locale overlay uses custom fixed chrome
|
reduced-chromeMinimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. Method used or attempted: Planned from a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses. | At 780x493, the locale chooser plus navigation occupies about 183 px, roughly 37% of the first viewport, before primary content. Reference: screenshot; retained-private | Locale chooser dominates the first viewport
|
Adapt to the form factor · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
responsive-no-horizontal-scrollLayout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. Method used or attempted: Planned from layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses. | At 360x800, layout reported scrollWidth 360, clientWidth 360, zero horizontal overflow, and a valid viewport meta tag. Reference: layout-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
component-level-responsivenessComponents adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints. | Failed / issue Source: issues; confidence: medium | Implementation and methodCatalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. Method used or attempted: Planned from CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses. | The responsive page uses no detectable @container or container-type rules, so reused components rely on page-level adaptation only. Reference: layout-summary, page-probe-summary, screenshot; retained-private | Components do not use container queries
|
input-modality-awareTouch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. Method used or attempted: Planned from a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses. | The focused locale control has a visible 2 px outline, but the locale close button measures only 18x18 CSS px. Reference: page-probe-summary; retained-private | Locale close target is too small
|
Support core task success · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
clear-purpose-and-primary-actionThe page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses. Method used or attempted: Planned from screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action | The first viewport names the iPhone offer and exposes Learn more and Shop iPhone actions with strong visual hierarchy. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
primary-flow-completionThe representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses. Method used or attempted: Planned from run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model | The reviewed boundary permits only the exact homepage URL, so following Shop iPhone or another primary action was prohibited. Reference: No artifact reference; described-only | Only the exact input URL is permitted. |
clear-system-state-and-recoveryLoading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses. Method used or attempted: Planned from exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain | No error, offline, empty, or completion state could be reached without network mutation or another URL. Reference: No artifact reference; described-only | The strict journey and exact-URL boundary prohibit the required failure-state exercise. |
Be fast and stable · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
good-core-web-vitalsCore Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses. Method used or attempted: Planned from Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal fir | The trace measured LCP 8477 ms and FCP 6311 ms; mobile layout measured CLS [host/path omitted], outside the good thresholds. Reference: layout-summary, performance-summary, screenshot; retained-private | Slow LCP and unstable mobile load
|
visual-stabilityNo cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. Method used or attempted: Planned from the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses. | Layout measured CLS [host/path omitted] on mobile and [host/path omitted] on desktop, including a single mobile shift of [host/path omitted]. Reference: layout-summary, screenshot; retained-private | Slow LCP and unstable mobile load
|
efficient-main-threadThe main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. Method used or attempted: Planned from the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses. | The trace found zero long tasks and 0 ms total blocking time during the captured load. Reference: performance-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
efficient-resource-deliveryCritical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses. Method used or attempted: Planned from a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discover | The HAR recorded 51 requests and [host/path omitted] MB transferred, including 10 parser-inserted render-blocking candidates and [host/path omitted] MB of fonts. Reference: network-summary; retained-private | Heavy render-blocking delivery path
|
trim-unused-and-duplicate-codeThe page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value. | Blocked Source: blocked; confidence: medium | Implementation and methodCatalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses. Method used or attempted: Planned from Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model | The HAR identified script weight but no code-coverage evidence was available to distinguish used from unused or duplicated code. Reference: No artifact reference; described-only | The permitted evidence primitives in this run did not expose JavaScript or CSS coverage. |
Be inclusive · 5 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
names-roles-labelsInteractive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses. Method used or attempted: Planned from axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The | The DOM probe found labelled search controls, extensive landmarks, and alt attributes on all img elements; visible primary actions have descriptive names. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
sufficient-contrastText and essential UI meet WCAG colour-contrast minimums against their background. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. Method used or attempted: Planned from axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses. | Normal and forced-colors screenshots retain legible text and clear essential-control boundaries. Reference: other-private-evidence, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
structure-and-focusHeading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. Method used or attempted: Planned from axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses. | Focus is visibly outlined, but the heading inventory contains a generic H1 of Apple and one empty H2 among 55 headings. Reference: page-probe-summary; retained-private | Heading structure is weak
|
legible-textText is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. Method used or attempted: Planned from a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses. | Desktop, mobile, and journey captures show unclipped, comfortably spaced hero and supporting text. Reference: layout-summary, screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
zoom-reflow-targets-and-mediaThe experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses. Method used or attempted: Planned from Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The mo | The 360 px page reflows without overflow, but the locale close control is only 18x18 CSS px, below an adequate touch target. Reference: layout-summary, page-probe-summary, screenshot; retained-private | Locale close target is too small
|
Follow best practices · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-console-errorsThe page loads without console errors or uncaught exceptions. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. Method used or attempted: Planned from capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses. | The supplied strict journey recorded its required console collector as blocked, and no equivalent Runtime/Log stream was retained. Reference: No artifact reference; described-only | Required Stage 2 console collector was unavailable in the supplied V1 journey. |
sound-document-and-assetsValid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.) | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. Method used or attempted: Planned from a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses. | Doctype and UTF-8 are valid, but all 52 img elements lack width and height attributes; measured mobile CLS is [host/path omitted]. Reference: image-summary, layout-summary, screenshot; retained-private | Images do not reserve intrinsic space
|
browser-platform-hygieneThe page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load. | Blocked Source: blocked; confidence: medium | Implementation and methodCatalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. Method used or attempted: Planned from Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses. | No deprecation, BFCache, source-map, vulnerable-library, or paste-prevention diagnostic was retained. Reference: No artifact reference; described-only | The bounded run did not include the required inspector and BFCache diagnostics. |
Be discoverable · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
title-and-descriptionThe page has a unique, descriptive <title> and a meta description. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. Method used or attempted: Planned from a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses. | The title is only Apple and there is no meta description, although Open Graph description metadata exists. Reference: page-probe-summary; retained-private | Homepage metadata is too generic
|
crawlable-and-mobile-friendlyLinks are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. Method used or attempted: Planned from a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses. | The probe found a viewport meta tag and 347 real href links; discoverability found 96% raw-HTML content coverage. Reference: discoverability-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
canonical-and-indexing-signalsPublic pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses. Method used or attempted: Planned from inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, [host/path omitted] and [host/path omitted]; Lighthouse SEO audits cover several of these. The model chooses. | The document returned 200, declares the homepage canonical, provides 137 hreflang links, and has no noindex robots meta. Reference: discoverability-summary, page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
structured-and-shareable-metadataWhere the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses. Method used or attempted: Planned from inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse [host/path omitted] blocks. The model c | The page exposes three JSON-LD blocks plus Open Graph title, description, and image matching the visible brand content. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Be private and secure · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
secure-transport-and-headersServed over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses. Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence. | Retained evidence records 5 of 6 baseline security-header categories present. Cookie-attribute review records 5 of 6 records with Secure and 0 with HttpOnly. Names, values, routes, and raw headers remain private. Reference: cookie-attribute-summary, security-header-summary; retained-private | CSP and cookie transport defenses are weakened
|
data-minimisation-and-third-partiesNo over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses. Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence. | Retained evidence records 1 first-party origin categories and 1 third-party origin categories across 51 requests. Destinations, identifiers, routes, headers, and bodies remain private. Reference: cookie-attribute-summary, network-summary, security-header-summary; retained-private | Analytics and identifiers activate on baseline load
|
in-context-permissions-and-modern-authPermission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses. Method used or attempted: Planned from a probe for permission requests fired on load; source inspection for passkey / WebAuthn / [host/path omitted] usage in auth flows. The model chooses. | No browser permission prompt appeared on load; authentication is not part of the audited homepage state. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
defensive-browser-policiesBrowser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses. Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence. | Retained evidence records 5 of 6 baseline browser-policy header categories present. Values and raw headers remain private. Reference: cookie-attribute-summary, security-header-summary; retained-private | Permissions policy is absent
|
Be resilient · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
progressive-enhancementCore content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. Method used or attempted: Planned from load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses. | Discoverability measured 96% of rendered content in raw HTML, with title and H1 present and no empty JavaScript shell. Reference: discoverability-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
resilient-runtime-behaviourThe page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly. | Blocked Source: blocked; confidence: medium | Implementation and methodCatalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. Method used or attempted: Planned from exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses. | The supplied journey only loaded and scrolled; menu-edge placement and repeated async overlay state were not exercised. Reference: No artifact reference; described-only | Strict supplied journey did not include disclosure or menu actions. |
offline-and-installableWhere the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.) | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. Method used or attempted: Planned from a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses. | The audited surface is a public product-marketing and commerce homepage, not an installable web application; no offline app intent was declared. Reference: No artifact reference; described-only | Offline installability is contextual and does not apply to this audited surface. |
network-and-http-failure-statesHTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses. Method used or attempted: Planned from simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The mo | The exact-URL policy and strict journey did not permit simulated HTTP failures or navigation to an error route. Reference: No artifact reference; described-only | Failure injection and alternate URLs were outside the execution boundary. |
Be internationalised · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
lang-dir-and-logical-propertiesCorrect lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. Method used or attempted: Planned from a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses. | The document declares en-US and ltr, uses logical CSS properties, and provides 137 hreflang alternatives including RTL locales. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
locale-aware-dataDates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. Method used or attempted: Planned from source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses. | The page presents region-specific content and a country chooser, and exposes extensive locale-specific alternate URLs. Reference: page-probe-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
time-zone-correctnessTime handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. Method used or attempted: Planned from source inspection for time-zone-aware date handling vs naive local Date math. The model chooses. | No time, event, recurrence, or time-zone-sensitive data appears in the audited homepage and journey states. Reference: No artifact reference; described-only | The audited surface contains no time-zone-sensitive content. |
Be trustworthy · 4 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-dark-patternsNo deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive). | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. Method used or attempted: Planned from a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses. | The locale prompt states its purpose plainly, offers a close control, and uses neutral Continue wording without confirmshaming. Reference: screenshot; retained-private | The retained evidence directly supported this check under the captured conditions. |
humane-error-handlingForms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. Method used or attempted: Planned from exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses. | The page contains search, but strict journey policy prohibited input mutation and submission, so validation and recovery could not be tested. Reference: No artifact reference; described-only | Form input and submission are forbidden by the strict journey policy. |
trustworthy-input-assistanceInput is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. Method used or attempted: Planned from source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses. | No address, payment, sign-in, or sign-up fields appear in the audited state; the only field is site search. Reference: No artifact reference; described-only | Autofill assistance for transactional input does not apply to the audited homepage state. |
safe-commercial-and-account-flowsCheckout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity. | Blocked Source: blocked; confidence: high | Implementation and methodCatalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses. Method used or attempted: Planned from walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support | Commercial links are visible, but checkout, subscription, authentication, and account routes could not be entered under the exact-URL boundary. Reference: No artifact reference; described-only | Only the exact homepage URL is permitted. |
Be sustainable · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
optimised-assetsImages and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. Method used or attempted: Planned from inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses. | Image inspection found 52 legacy-format images, 31 below-fold images without lazy loading, and 23 images without responsive srcset. Reference: image-summary, layout-summary, screenshot; retained-private | Image delivery is not resource-efficient
|
no-wasteful-workBackground work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. Method used or attempted: Planned from a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses. | Initial load shipped a 104 KB analytics script plus data-relay scripts and contacted securemetrics before user interaction. Reference: network-summary; retained-private | Non-essential analytics work starts immediately
|
third-party-and-media-budgetThird-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible. | Failed / issue Source: issues; confidence: high | Implementation and methodCatalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses. Method used or attempted: Planned from a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation ke | The homepage transferred [host/path omitted] MB, with [host/path omitted] MB of fonts, 424 KB of scripts, and three hero media requests including two aborted requests. Reference: network-summary; retained-private | Font, script, and media budget is excessive
|
Be agent ready · 2 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
structured-agent-capabilitiesWhere it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. Method used or attempted: Planned from source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses. | No agent-facing intent is declared for this audited marketing homepage; raw content remains highly machine-readable at 96% coverage. Reference: No artifact reference; described-only | This emerging capability is contextual and no agent-facing surface was declared. |
on-device-inferenceOn-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server. | Not applicable Source: not-applicable; confidence: high | Implementation and methodCatalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses. Method used or attempted: Planned from source inspection for built-in AI (language model / summariser) usage. The model chooses. | No page feature calls for on-device inference in the audited homepage and journey states. Reference: No artifact reference; described-only | On-device inference is contextual and no relevant task is present. |
Be memory-efficient · 3 tests
| Test | Verdict | How it was implemented | Evidence | Why it passed, failed, or was incomplete |
|---|---|---|---|---|
no-leak-under-repeated-interactionRepeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends. | Pass Source: pass; confidence: high | Implementation and methodCatalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses. Method used or attempted: Planned from compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repe | After ten scroll-to-500-and-back cycles, heap self-size rose only 9,131 bytes, about [host/path omitted]%, while closure count fell by four. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
bounded-footprintHeap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses. Method used or attempted: Planned from a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus [host/path omitted] (Nodes, JSHeapUsedSize) give the current footprint to judge aga | The baseline heap was [host/path omitted] MB self-size with 297,024 nodes, proportionate to the media-rich homepage and stable after exercise. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
no-detached-dom-or-unbounded-listenersNo growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session. | Pass Source: pass; confidence: medium | Implementation and methodCatalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses. Method used or attempted: Planned from the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer count | Neither heap summary surfaced a Detached constructor among top populations; closures decreased and arrays were effectively flat after repetition. Reference: memory-summary; retained-private | The retained evidence directly supported this check under the captured conditions. |
Methodology
- Keep the reviewed convenience cohort fixed at ten origins, with no substitutions or automatic retries.
- Verify the hash-chained event ledger and use it as the sole disposition authority. Stale
site-runpending states and conflicting report states are ignored. - Recompute coverage from the 58 atomic check outcomes. A missing report produces 58 missing rows.
- Allowlist only origins, categorical outcomes, counts, bytes, timing aggregates, header presence, and cookie-attribute counts. URLs are reduced to origins before counting.
- Deterministically re-encode selected media, strip metadata, verify hashes and dimensions, and admit it only after OCR and visual privacy review.
Network and trace facts describe one retained collection under its recorded conditions. They are not lab scores and are not comparable performance rankings.
Limitations
- This is a fixed, reviewed convenience pilot, not a representative sample. No population inference is supported.
- Observations are limited to public landing surfaces, reachable states, strict mutation containment, and the captured window.
- Authentication walls, humanity challenges, exact-origin scope, and unavailable Stage 2 collectors limited journey and check coverage.
- The ledger is authoritative for disposition, but its two completed classifications do not establish valid completion quality because both console methods were unsupported.
- Cleanup evidence is an operational assertion rather than authenticated per-session teardown receipts; additional audit profiles observed in private logs were outside the ten-profile cleanup list.
Public evidence and provenance
- Sanitized aggregate JSON
- All 580 sanitized test outcomes
- Strict test-outcome schema
- Public evidence manifest
- Media transform and review receipts
- Source and provenance hashes
- Sanitized cleanup summary
- Strict aggregate schema
Runner commit: e644b7c7ce7c1e82cf797bc8e7affb9fea03ff37
Pilot manifest SHA-256: d7eb927eac2089b2cea449071e50e8709d7328dea6bbba7ccf8feb9e5f309355
Source manifest SHA-256: af3d02a5a1466181c5900104795e25cd8d3838375702520260cdaede5078791d
Policy SHA-256: a8e7148c1d81bc57ca653517f6b57beabd483c9be9e5bf0abafc048f9014b5e0
Ledger SHA-256: 9098cd36c049e8a30d1b419dbd7b2cd26bd9246e25ec4682ecf20ee077f76bc9
Publication authorization: explicit authorization for this sanitized derivative was granted after the run. The private start receipt recorded the status at collection start; it does not negate later authorization. No relay identifiers are published.