Evidence report · fixed denominator 10 sites × 58 tests

Fixed-10 Web Uplift evidence report

Inspect every test, verdict, method, sanitized evidence summary, and failure or collection reason from the retained pilot. This is not a site ranking, quality league table, or replacement for the main State of the Web report.

Read this before interpreting the ledger

The authoritative hash-chained ledger classified 2 rows runner-completed, 6 partial, and 2 errors. Both runner-completed rows contain method-invalid console passes: the collector was unsupported and absence of surfaced errors was treated as evidence. Therefore this report publishes no scores, pass rates, rankings, or completed-quality claims. “Runner-completed” means only that the ledger classified the row completed.

Ledger disposition

10fixed origins
2ledger runner-completed
6partial
2errors

No target was replaced or retried. GitHub is correctly reported as missing-report; permit expiry is not claimed because the retained timestamps disprove it. Facebook is an error because the runner invalidated completed cross-origin journey states.

Ten-row inventory

Authoritative ledger disposition, recomputed atomic coverage, and journey status for the fixed ten origins
OriginArchetypeDisposition / reasonCoverageJourney
1. https://github.comdocumentationerror
missing-report
0 judged; 0 blocked; 0 not run; 58 missingpartial
2. https://web.facebook.comblocked-complexerror
runner-error
0 judged; 0 blocked; 0 not run; 58 missinginvalid
3. https://www.google.comsearch-portalrunner-completed
not-applicable
58 judged; 0 blocked; 0 not run; 0 missingpartial
4. https://www.reddit.cominteraction-heavypartial
atomic-coverage-incomplete
49 judged; 9 blocked; 0 not run; 0 missingpartial
5. https://www.amazon.comcommercepartial
atomic-coverage-incomplete
39 judged; 19 blocked; 0 not run; 0 missingpartial
6. https://en.wikipedia.orgsimple-staticrunner-completed
not-applicable
58 judged; 0 blocked; 0 not run; 0 missingpartial
7. https://www.nytimes.comcontent-newspartial
atomic-coverage-incomplete
50 judged; 6 blocked; 2 not run; 0 missingpartial
8. https://gemini.google.comspa-productpartial
atomic-coverage-incomplete
36 judged; 22 blocked; 0 not run; 0 missingpartial
9. https://www.netflix.commedia-heavypartial
atomic-coverage-incomplete
50 judged; 8 blocked; 0 not run; 0 missingpartial
10. https://www.apple.comgeneralpartial
atomic-coverage-incomplete
49 judged; 9 blocked; 0 not run; 0 missingpartial

1. https://github.com

Ledger: error · documentation

Reason: missing-report. The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Recomputed coverage: expected 58; recorded 0; judged 0; blocked 0; not run 0; missing 58; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-like-url)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

No network or performance summary was derivable because no report was produced.

Signed-out GitHub landing surface with an empty email field.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

2. https://web.facebook.com

Ledger: error · blocked-complex

Reason: runner-error. The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Recomputed coverage: expected 58; recorded 0; judged 0; blocked 0; not run 0; missing 58; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: invalidated-cross-origin-state (completed-cross-origin-state)
  2. bounded-scroll: invalidated-cross-origin-state (completed-cross-origin-state)

Safe aggregates

No network or performance summary was derivable because no report was produced.

Blank dark Facebook journey baseline frame.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

3. https://www.google.com

Ledger: runner-completed · search-portal

Reason: not-applicable. The ledger classified this row completed, but its console pass used an unsupported absence-of-error method.

Recomputed coverage: expected 58; recorded 58; judged 58; blocked 0; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-method-post)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
43
Transfer
816,024 bytes
Observed origins
1 first-party; 4 third-party
Security headers present
2 of 6
Cookies observed
2 total; 2 Secure; 2 HttpOnly
Trace timings
FCP 1,482.05 ms; LCP 1,499.69 ms; 1 long tasks
Signed-out Google consent surface.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

4. https://www.reddit.com

Ledger: partial · interaction-heavy

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after a humanity challenge and cross-origin navigation block.

Recomputed coverage: expected 58; recorded 58; judged 49; blocked 9; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: navigation-blocked (cross-origin-document)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
8
Transfer
514,675 bytes
Observed origins
1 first-party; 3 third-party
Security headers present
3 of 6
Cookies observed
1 total; 1 Secure; 0 HttpOnly
Trace timings
FCP 677.84 ms; LCP 677.84 ms; 0 long tasks
Reddit humanity challenge page.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.
A silent 0.1-second retained frame of the Reddit humanity challenge under reduced-motion capture. It has no audio; this caption describes the full visual content. First, middle, and last frame OCR and visual privacy review passed.

5. https://www.amazon.com

Ledger: partial · commerce

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like POST traffic was blocked.

Recomputed coverage: expected 58; recorded 58; judged 39; blocked 19; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-method-post)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
43
Transfer
1,101,460 bytes
Observed origins
1 first-party; 3 third-party
Security headers present
0 of 6
Cookies observed
7 total; 3 Secure; 0 HttpOnly
Trace timings
FCP not observed; LCP not observed; 0 long tasks
Blank white Amazon journey baseline frame.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

6. https://en.wikipedia.org

Ledger: runner-completed · simple-static

Reason: not-applicable. The ledger classified this row completed, but its console pass used an unsupported absence-of-error method.

Recomputed coverage: expected 58; recorded 58; judged 58; blocked 0; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: performed (same-origin-get-navigation-completed)
  2. bounded-scroll: performed (bounded-scroll-completed)

Safe aggregates

Requests
40
Transfer
682,131 bytes
Observed origins
1 first-party; 2 third-party
Security headers present
3 of 6
Cookies observed
7 total; 6 Secure; 4 HttpOnly
Trace timings
FCP 1,123.66 ms; LCP 1,123.66 ms; 1 long tasks
Signed-out English Wikipedia landing surface.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

7. https://www.nytimes.com

Ledger: partial · content-news

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like POST traffic was blocked; six checks were blocked and two were not run.

Recomputed coverage: expected 58; recorded 58; judged 50; blocked 6; not run 2; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-method-post)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
368
Transfer
4,950,323 bytes
Observed origins
1 first-party; 27 third-party
Security headers present
5 of 6
Cookies observed
14 total; 9 Secure; 3 HttpOnly
Trace timings
FCP 755.89 ms; LCP 4,760.96 ms; 13 long tasks
New York Times privacy preferences surface with no filled controls.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.
A silent 2.3-second reduced-motion scroll from a generic privacy panel to public news content. It has no audio; this caption describes the full visual content. First, middle, and last frame OCR and visual privacy review passed.

8. https://gemini.google.com

Ledger: partial · spa-product

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like OPTIONS traffic was blocked.

Recomputed coverage: expected 58; recorded 58; judged 36; blocked 22; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-method-options)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
100
Transfer
5,424,867 bytes
Observed origins
1 first-party; 8 third-party
Security headers present
4 of 6
Cookies observed
1 total; 1 Secure; 1 HttpOnly
Trace timings
FCP 6,623.01 ms; LCP 6,861.74 ms; 2 long tasks
Signed-out Gemini consent surface.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

9. https://www.netflix.com

Ledger: partial · media-heavy

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete under mutation and scope boundaries and unavailable diagnostic or account-flow evidence.

Recomputed coverage: expected 58; recorded 58; judged 50; blocked 8; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-like-url)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
53
Transfer
2,579,592 bytes
Observed origins
1 first-party; 11 third-party
Security headers present
3 of 6
Cookies observed
9 total; 3 Secure; 3 HttpOnly
Trace timings
FCP 2,462.92 ms; LCP 2,462.92 ms; 2 long tasks
Signed-out Netflix landing surface with an empty email field.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

10. https://www.apple.com

Ledger: partial · general

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete because exact-origin scope left primary, error, account, and diagnostic flows unavailable.

Recomputed coverage: expected 58; recorded 58; judged 49; blocked 9; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: performed (same-origin-get-navigation-completed)
  2. bounded-scroll: performed (bounded-scroll-completed)

Safe aggregates

Requests
51
Transfer
1,965,289 bytes
Observed origins
1 first-party; 1 third-party
Security headers present
5 of 6
Cookies observed
6 total; 5 Secure; 0 HttpOnly
Trace timings
FCP 6,310.82 ms; LCP 8,477.41 ms; 0 long tasks
Apple region selector above a public iPhone landing surface.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.
A silent 2.5-second reduced-motion scroll from a region selector through public product content. It has no audio; this caption describes the full visual content. First, middle, and last frame OCR and visual privacy review passed.

Row evidence details

1. https://github.com

Ledger: error · documentation

Reason: missing-report. The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Recomputed coverage: expected 58; recorded 0; judged 0; blocked 0; not run 0; missing 58; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-like-url)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

No network or performance summary was derivable because no report was produced.

Signed-out GitHub landing surface with an empty email field.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

2. https://web.facebook.com

Ledger: error · blocked-complex

Reason: runner-error. The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Recomputed coverage: expected 58; recorded 0; judged 0; blocked 0; not run 0; missing 58; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: invalidated-cross-origin-state (completed-cross-origin-state)
  2. bounded-scroll: invalidated-cross-origin-state (completed-cross-origin-state)

Safe aggregates

No network or performance summary was derivable because no report was produced.

Blank dark Facebook journey baseline frame.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

3. https://www.google.com

Ledger: runner-completed · search-portal

Reason: not-applicable. The ledger classified this row completed, but its console pass used an unsupported absence-of-error method.

Recomputed coverage: expected 58; recorded 58; judged 58; blocked 0; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-method-post)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
43
Transfer
816,024 bytes
Observed origins
1 first-party; 4 third-party
Security headers present
2 of 6
Cookies observed
2 total; 2 Secure; 2 HttpOnly
Trace timings
FCP 1,482.05 ms; LCP 1,499.69 ms; 1 long tasks
Signed-out Google consent surface.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

4. https://www.reddit.com

Ledger: partial · interaction-heavy

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after a humanity challenge and cross-origin navigation block.

Recomputed coverage: expected 58; recorded 58; judged 49; blocked 9; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: navigation-blocked (cross-origin-document)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
8
Transfer
514,675 bytes
Observed origins
1 first-party; 3 third-party
Security headers present
3 of 6
Cookies observed
1 total; 1 Secure; 0 HttpOnly
Trace timings
FCP 677.84 ms; LCP 677.84 ms; 0 long tasks
Reddit humanity challenge page.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.
A silent 0.1-second retained frame of the Reddit humanity challenge under reduced-motion capture. It has no audio; this caption describes the full visual content. First, middle, and last frame OCR and visual privacy review passed.

5. https://www.amazon.com

Ledger: partial · commerce

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like POST traffic was blocked.

Recomputed coverage: expected 58; recorded 58; judged 39; blocked 19; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-method-post)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
43
Transfer
1,101,460 bytes
Observed origins
1 first-party; 3 third-party
Security headers present
0 of 6
Cookies observed
7 total; 3 Secure; 0 HttpOnly
Trace timings
FCP not observed; LCP not observed; 0 long tasks
Blank white Amazon journey baseline frame.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

6. https://en.wikipedia.org

Ledger: runner-completed · simple-static

Reason: not-applicable. The ledger classified this row completed, but its console pass used an unsupported absence-of-error method.

Recomputed coverage: expected 58; recorded 58; judged 58; blocked 0; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: performed (same-origin-get-navigation-completed)
  2. bounded-scroll: performed (bounded-scroll-completed)

Safe aggregates

Requests
40
Transfer
682,131 bytes
Observed origins
1 first-party; 2 third-party
Security headers present
3 of 6
Cookies observed
7 total; 6 Secure; 4 HttpOnly
Trace timings
FCP 1,123.66 ms; LCP 1,123.66 ms; 1 long tasks
Signed-out English Wikipedia landing surface.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

7. https://www.nytimes.com

Ledger: partial · content-news

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like POST traffic was blocked; six checks were blocked and two were not run.

Recomputed coverage: expected 58; recorded 58; judged 50; blocked 6; not run 2; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-method-post)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
368
Transfer
4,950,323 bytes
Observed origins
1 first-party; 27 third-party
Security headers present
5 of 6
Cookies observed
14 total; 9 Secure; 3 HttpOnly
Trace timings
FCP 755.89 ms; LCP 4,760.96 ms; 13 long tasks
New York Times privacy preferences surface with no filled controls.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.
A silent 2.3-second reduced-motion scroll from a generic privacy panel to public news content. It has no audio; this caption describes the full visual content. First, middle, and last frame OCR and visual privacy review passed.

8. https://gemini.google.com

Ledger: partial · spa-product

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete after mutation-like OPTIONS traffic was blocked.

Recomputed coverage: expected 58; recorded 58; judged 36; blocked 22; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-method-options)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
100
Transfer
5,424,867 bytes
Observed origins
1 first-party; 8 third-party
Security headers present
4 of 6
Cookies observed
1 total; 1 Secure; 1 HttpOnly
Trace timings
FCP 6,623.01 ms; LCP 6,861.74 ms; 2 long tasks
Signed-out Gemini consent surface.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

9. https://www.netflix.com

Ledger: partial · media-heavy

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete under mutation and scope boundaries and unavailable diagnostic or account-flow evidence.

Recomputed coverage: expected 58; recorded 58; judged 50; blocked 8; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: mutation-blocked (mutation-like-url)
  2. bounded-scroll: skipped-after-blocker (journey-aborted-after-blocker)

Safe aggregates

Requests
53
Transfer
2,579,592 bytes
Observed origins
1 first-party; 11 third-party
Security headers present
3 of 6
Cookies observed
9 total; 3 Secure; 3 HttpOnly
Trace timings
FCP 2,462.92 ms; LCP 2,462.92 ms; 2 long tasks
Signed-out Netflix landing surface with an empty email field.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.

10. https://www.apple.com

Ledger: partial · general

Reason: atomic-coverage-incomplete. Atomic coverage was incomplete because exact-origin scope left primary, error, account, and diagnostic flows unavailable.

Recomputed coverage: expected 58; recorded 58; judged 49; blocked 9; not run 0; missing 0; unknown 0; duplicates 0.

Journey actions

  1. baseline-load: performed (same-origin-get-navigation-completed)
  2. bounded-scroll: performed (bounded-scroll-completed)

Safe aggregates

Requests
51
Transfer
1,965,289 bytes
Observed origins
1 first-party; 1 third-party
Security headers present
5 of 6
Cookies observed
6 total; 5 Secure; 0 HttpOnly
Trace timings
FCP 6,310.82 ms; LCP 8,477.41 ms; 0 long tasks
Apple region selector above a public iPhone landing surface.
Sanitized baseline screenshot. Metadata stripped; OCR and visual privacy review passed.
A silent 2.5-second reduced-motion scroll from a region selector through public product content. It has no audio; this caption describes the full visual content. First, middle, and last frame OCR and visual privacy review passed.

All 580 tests and evidence

This explorer retains every catalog slot. An issue is a tested product failure. Blocked and not run mean incomplete collection. Unavailable means the site produced no atomic report, so no test result is claimed.

Exact totals: 174 pass; 147 issues; 68 not applicable; 73 blocked; 2 not run; 116 unavailable.

Showing all 580 site-check slots.

1. https://github.com · 58 slots · missing-report

Collection failure: these 58 slots were materialized from the catalog so the denominator remains visible. They were not tested and have no check-specific evidence.

Respect user preferences · 3 tests
Respect user preferences checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
respects-color-scheme

Honours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

respects-reduced-motion

Honours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

respects-contrast

Honours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Implement natural interactions · 3 tests
Implement natural interactions checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
view-transitions

State and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

scroll-driven-animations

Scroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

physical-gestures

Gesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Provide guided navigation · 3 tests
Provide guided navigation checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
scroll-state-aware-chrome

Sticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

anchored-positioning

Tooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

directs-attention

Navigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Maximize content, reduce noise · 3 tests
Maximize content, reduce noise checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-intrusive-interruptions

No intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

semantic-dismissible-primitives

Overlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

reduced-chrome

Minimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Adapt to the form factor · 3 tests
Adapt to the form factor checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
responsive-no-horizontal-scroll

Layout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

component-level-responsiveness

Components adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

input-modality-aware

Touch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Support core task success · 3 tests
Support core task success checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
clear-purpose-and-primary-action

The page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

primary-flow-completion

The representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

clear-system-state-and-recovery

Loading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Be fast and stable · 5 tests
Be fast and stable checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
good-core-web-vitals

Core Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

visual-stability

No cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

efficient-main-thread

The main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

efficient-resource-delivery

Critical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

trim-unused-and-duplicate-code

The page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Be inclusive · 5 tests
Be inclusive checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
names-roles-labels

Interactive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

sufficient-contrast

Text and essential UI meet WCAG colour-contrast minimums against their background.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

structure-and-focus

Heading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

legible-text

Text is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

zoom-reflow-targets-and-media

The experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Follow best practices · 3 tests
Follow best practices checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-console-errors

The page loads without console errors or uncaught exceptions.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

sound-document-and-assets

Valid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.)

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

browser-platform-hygiene

The page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Be discoverable · 4 tests
Be discoverable checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
title-and-description

The page has a unique, descriptive <title> and a meta description.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

crawlable-and-mobile-friendly

Links are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

canonical-and-indexing-signals

Public pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

structured-and-shareable-metadata

Where the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Be private and secure · 4 tests
Be private and secure checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
secure-transport-and-headers

Served over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

data-minimisation-and-third-parties

No over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

in-context-permissions-and-modern-auth

Permission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

defensive-browser-policies

Browser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Be resilient · 4 tests
Be resilient checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
progressive-enhancement

Core content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

resilient-runtime-behaviour

The page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

offline-and-installable

Where the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.)

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

network-and-http-failure-states

HTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Be internationalised · 3 tests
Be internationalised checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
lang-dir-and-logical-properties

Correct lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

locale-aware-data

Dates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

time-zone-correctness

Time handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Be trustworthy · 4 tests
Be trustworthy checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-dark-patterns

No deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive).

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

humane-error-handling

Forms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

trustworthy-input-assistance

Input is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

safe-commercial-and-account-flows

Checkout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Be sustainable · 3 tests
Be sustainable checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
optimised-assets

Images and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

no-wasteful-work

Background work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

third-party-and-media-budget

Third-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Be agent ready · 2 tests
Be agent ready checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
structured-agent-capabilities

Where it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

on-device-inference

On-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

Be memory-efficient · 3 tests
Be memory-efficient checks for https://github.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-leak-under-repeated-interaction

Repeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

bounded-footprint

Heap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

no-detached-dom-or-unbounded-listeners

No growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session.

Unavailable
Source: missing-report; confidence: none
Implementation and method

Catalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records missing-report. Permit expiry is not claimed: the retained timestamps place the start before expiry.

2. https://web.facebook.com · 58 slots · runner-error

Collection failure: these 58 slots were materialized from the catalog so the denominator remains visible. They were not tested and have no check-specific evidence.

Respect user preferences · 3 tests
Respect user preferences checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
respects-color-scheme

Honours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

respects-reduced-motion

Honours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

respects-contrast

Honours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Implement natural interactions · 3 tests
Implement natural interactions checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
view-transitions

State and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

scroll-driven-animations

Scroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

physical-gestures

Gesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Provide guided navigation · 3 tests
Provide guided navigation checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
scroll-state-aware-chrome

Sticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

anchored-positioning

Tooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

directs-attention

Navigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Maximize content, reduce noise · 3 tests
Maximize content, reduce noise checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-intrusive-interruptions

No intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

semantic-dismissible-primitives

Overlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

reduced-chrome

Minimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Adapt to the form factor · 3 tests
Adapt to the form factor checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
responsive-no-horizontal-scroll

Layout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

component-level-responsiveness

Components adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

input-modality-aware

Touch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Support core task success · 3 tests
Support core task success checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
clear-purpose-and-primary-action

The page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

primary-flow-completion

The representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

clear-system-state-and-recovery

Loading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Be fast and stable · 5 tests
Be fast and stable checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
good-core-web-vitals

Core Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

visual-stability

No cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

efficient-main-thread

The main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

efficient-resource-delivery

Critical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

trim-unused-and-duplicate-code

The page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Be inclusive · 5 tests
Be inclusive checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
names-roles-labels

Interactive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

sufficient-contrast

Text and essential UI meet WCAG colour-contrast minimums against their background.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

structure-and-focus

Heading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

legible-text

Text is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

zoom-reflow-targets-and-media

The experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Follow best practices · 3 tests
Follow best practices checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-console-errors

The page loads without console errors or uncaught exceptions.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

sound-document-and-assets

Valid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.)

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

browser-platform-hygiene

The page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Be discoverable · 4 tests
Be discoverable checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
title-and-description

The page has a unique, descriptive <title> and a meta description.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

crawlable-and-mobile-friendly

Links are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

canonical-and-indexing-signals

Public pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

structured-and-shareable-metadata

Where the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Be private and secure · 4 tests
Be private and secure checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
secure-transport-and-headers

Served over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

data-minimisation-and-third-parties

No over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

in-context-permissions-and-modern-auth

Permission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

defensive-browser-policies

Browser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Be resilient · 4 tests
Be resilient checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
progressive-enhancement

Core content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

resilient-runtime-behaviour

The page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

offline-and-installable

Where the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.)

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

network-and-http-failure-states

HTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Be internationalised · 3 tests
Be internationalised checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
lang-dir-and-logical-properties

Correct lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

locale-aware-data

Dates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

time-zone-correctness

Time handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Be trustworthy · 4 tests
Be trustworthy checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-dark-patterns

No deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive).

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

humane-error-handling

Forms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

trustworthy-input-assistance

Input is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

safe-commercial-and-account-flows

Checkout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Be sustainable · 3 tests
Be sustainable checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
optimised-assets

Images and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

no-wasteful-work

Background work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

third-party-and-media-budget

Third-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Be agent ready · 2 tests
Be agent ready checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
structured-agent-capabilities

Where it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

on-device-inference

On-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

Be memory-efficient · 3 tests
Be memory-efficient checks for https://web.facebook.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-leak-under-repeated-interaction

Repeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

bounded-footprint

Heap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

no-detached-dom-or-unbounded-listeners

No growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session.

Unavailable
Source: runner-error; confidence: none
Implementation and method

Catalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses.

Method used or attempted: No check-specific method was executed because the site report is unavailable.

No atomic report or check-specific evidence was produced for this site-check slot.

Reference: No artifact reference; unavailable

The authoritative ledger records runner-error. The runner invalidated completed journey states that crossed away from the reviewed origin.

3. https://www.google.com · 58 slots · available

Method warning: this site has a report, but its no-console-errors pass is method-invalid and is labelled in the table.

Respect user preferences · 3 tests
Respect user preferences checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
respects-color-scheme

Honours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method.

Method used or attempted: Compared permit-bound light and dark screenshots.

The page rendered a light surface under prefers-color-scheme: light and a dark surface under dark, with legible consent UI in both.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

respects-reduced-motion

Honours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses.

Method used or attempted: Permit-bound evaluate probe under prefers-reduced-motion: reduce.

The media query matched and [host/path omitted]() returned zero active animations.

Reference: other-private-evidence; retained-private

The retained evidence directly supported this check under the captured conditions.

respects-contrast

Honours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Method used or attempted: Permit-bound forced-colors and prefers-contrast screenshot review.

Essential consent text, links, outline borders, language, and sign-in controls remained visible in forced colors.

Reference: other-private-evidence; retained-private

The retained evidence directly supported this check under the captured conditions.

Implement natural interactions · 3 tests
Implement natural interactions checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
view-transitions

State and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps.

Not applicable
Source: not-applicable; confidence: medium
Implementation and method

Catalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

Method used or attempted: Reviewed the static entry and journey states.

No same-document or route state change was safely exercised on this single search entry surface.

Reference: No artifact reference; described-only

No applicable state or route transition was present in the bounded audited path.

scroll-driven-animations

Scroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

Method used or attempted: Layout, reduced-motion probe, and journey state review.

The entry document had no scrollable desktop surface and no scroll-linked animation.

Reference: layout-summary, other-private-evidence; retained-private

There is no scroll-linked visual effect on this surface.

physical-gestures

Gesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

Method used or attempted: Reviewed DOM, layout, and bounded repeated native scrolling.

The page relies on native scrolling and native form controls; no custom gesture surface or pointer-driven replacement was observed.

Reference: memory-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Provide guided navigation · 3 tests
Provide guided navigation checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
scroll-state-aware-chrome

Sticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Method used or attempted: Journey and layout review.

The desktop entry state has no meaningful scrolling chrome; the planned bounded scroll was skipped after the journey network mutation blocker.

Reference: journey-summary, layout-summary; retained-private

No sticky or affixed chrome needs to react on this bounded single-screen path.

anchored-positioning

Tooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned.

Not applicable
Source: not-applicable; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Method used or attempted: DOM review for popovers and anchored transient UI.

No tooltip or popover was open or required on the audited state; the consent surface is modal.

Reference: page-probe-summary; retained-private

No anchored transient overlay was present on the audited state.

directs-attention

Navigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

Method used or attempted: Visual hierarchy review of desktop, mobile, and journey baseline screenshots.

The centered Google mark, search field, and consent heading establish a clear reading order and visible current state.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

Maximize content, reduce noise · 3 tests
Maximize content, reduce noise checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-intrusive-interruptions

No intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

Method used or attempted: First-load screenshot review at the journey viewport.

A consent wall obscures the search task on load; at 780x493 its action controls are below the visible viewport.

Reference: screenshot; retained-private

The first-load consent wall obscures the core search task and its actions fall below a short desktop viewport.

  • site-03-finding-01 · high · The first-load consent wall obscures the core search task and its actions fall below a short desktop viewport.
semantic-dismissible-primitives

Overlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

Method used or attempted: DOM and role inspection.

The consent surface is exposed as a dialog and the DOM contains a dialog element rather than only an unstructured overlay div.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

reduced-chrome

Minimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

Method used or attempted: Desktop and mobile screenshot review.

Outside the required consent decision, the search page is visually sparse and gives the central task most of the viewport.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

Adapt to the form factor · 3 tests
Adapt to the form factor checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
responsive-no-horizontal-scroll

Layout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

Method used or attempted: Permit-bound 360x800 layout metrics and screenshot.

The response has no viewport meta tag; Chrome laid it out at 980 CSS px and scaled it to [host/path omitted] on the 360px device viewport.

Reference: layout-summary, screenshot; retained-private

The page omits viewport metadata and scales a 980px layout down on mobile.

  • site-03-finding-02 · high · The page omits viewport metadata and scales a 980px layout down on mobile.
component-level-responsiveness

Components adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints.

Not applicable
Source: not-applicable; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

Method used or attempted: CSS feature probe and component inventory review.

No reused component was observed in multiple container contexts; the page uses a single search layout.

Reference: page-probe-summary; retained-private

The audited page has no evidenced multi-container reuse case requiring component-level adaptation.

input-modality-aware

Touch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

Method used or attempted: Programmatically focused the first 20 focusable controls and inspected outlines and target rectangles.

Focused links and controls reported outline-style none and :focus-visible false; several text links were only 24 to 26 CSS px tall.

Reference: page-probe-summary; retained-private

Keyboard focus is not visibly exposed on sampled controls.

  • site-03-finding-03 · high · Keyboard focus is not visibly exposed on sampled controls.
Support core task success · 3 tests
Support core task success checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
clear-purpose-and-primary-action

The page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses.

Method used or attempted: First viewport and semantic form review.

The Google brand, labelled Search combobox, and clearly named search actions communicate the page purpose.

Reference: page-probe-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

primary-flow-completion

The representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses.

Method used or attempted: Reviewed the already-replayed strict journey and baseline screenshot.

The representative path did not reach search: the baseline action was marked mutation-blocked after POST telemetry and the bounded scroll was skipped; the consent wall also obscured search controls.

Reference: journey-summary, screenshot; retained-private

The first-load consent wall obscures the core search task and its actions fall below a short desktop viewport.

  • site-03-finding-01 · high · The first-load consent wall obscures the core search task and its actions fall below a short desktop viewport.
clear-system-state-and-recovery

Loading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses.

Method used or attempted: Consent state and journey ledger review.

The page explains the consent state and offers Reject all, Accept all, More options, privacy, and terms paths; the journey ledger explicitly records its blocker.

Reference: journey-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

Be fast and stable · 5 tests
Be fast and stable checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
good-core-web-vitals

Core Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses.

Method used or attempted: Permit-bound trace and layout observers.

LCP was [host/path omitted] ms, CLS was 0, and total blocking time was [host/path omitted] ms; no INP sample was produced because policy prohibited form submission.

Reference: layout-summary, performance-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

visual-stability

No cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Method used or attempted: Layout shift observer during page load.

The layout primitive recorded CLS 0 with no shift entries.

Reference: layout-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

efficient-main-thread

The main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

Method used or attempted: Permit-bound performance trace.

The trace recorded one [host/path omitted] ms long task and only [host/path omitted] ms total blocking time.

Reference: performance-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

efficient-resource-delivery

Critical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses.

Method used or attempted: HAR summary plus rendered DOM review.

The simple entry transferred 816,024 bytes across 43 requests; 715,201 bytes were script and parser-inserted stylesheets appeared as render-blocking candidates.

Reference: network-summary, page-probe-summary; retained-private

Resource delivery is heavy for the entry task and includes parser-inserted blocking styles.

  • site-03-finding-04 · medium · Resource delivery is heavy for the entry task and includes parser-inserted blocking styles.
trim-unused-and-duplicate-code

The page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value.

Failed / issue
Source: issues; confidence: medium
Implementation and method

Catalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses.

Method used or attempted: HAR resource-type and page-complexity comparison.

Eight scripts transferred 715,201 bytes for a roughly 502-node search entry. The payload is disproportionate, although this run did not collect code coverage to isolate exact unused ranges.

Reference: network-summary, page-probe-summary; retained-private

JavaScript payload is disproportionate to the initial search surface.

  • site-03-finding-05 · medium · JavaScript payload is disproportionate to the initial search surface.
Be inclusive · 5 tests
Be inclusive checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
names-roles-labels

Interactive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses.

Method used or attempted: DOM accessibility semantics and visible control review.

The primary search textarea is labelled Search with role combobox, the form has role search, visible submit controls are named, and the visible Google SVG is exposed as an image.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

sufficient-contrast

Text and essential UI meet WCAG colour-contrast minimums against their background.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Method used or attempted: Normal dark and forced-colors screenshot review.

Primary and consent text remained legible against dark surfaces, and essential UI survived forced colors.

Reference: other-private-evidence, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

structure-and-focus

Heading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Method used or attempted: Heading, landmark, and programmatic focus inspection.

A logical H1 and search landmark exist, but the focus probe found no visible outline or :focus-visible match on the sampled keyboard-reachable links and controls.

Reference: page-probe-summary; retained-private

Keyboard focus is not visibly exposed on sampled controls.

  • site-03-finding-03 · high · Keyboard focus is not visibly exposed on sampled controls.
legible-text

Text is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Method used or attempted: Desktop, mobile, and contrast screenshot review.

Search and consent copy uses readable line lengths and spacing; no text clipping was visible inside the mobile full-page capture.

Reference: other-private-evidence, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

zoom-reflow-targets-and-media

The experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses.

Method used or attempted: Narrow viewport layout and screenshot review.

Without viewport metadata, the page renders a 980px layout scaled down to 360px, making text and controls unusually small instead of reflowing at device width.

Reference: layout-summary, screenshot; retained-private

The page omits viewport metadata and scales a 980px layout down on mobile.

  • site-03-finding-02 · high · The page omits viewport metadata and scales a 980px layout down on mobile.
Follow best practices · 3 tests
Follow best practices checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-console-errors

The page loads without console errors or uncaught exceptions.

Pass
Source: pass; confidence: low

Method invalid: this console pass used absence of a surfaced error even though the authoritative console collector was unavailable. Treat it as unreliable evidence, not a valid pass.

Implementation and method

Catalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

Method used or attempted: Observed permit-bound Chrome CLI load, trace completion, and journey event ledger.

The independent loads completed without an uncaught-exception signal in CLI output; the earlier journey console collector itself was unavailable, so confidence is low.

Reference: journey-summary, performance-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

sound-document-and-assets

Valid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.)

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Method used or attempted: DOM, document metadata, image, and layout inspection.

The document has an HTML doctype and UTF-8 charset; visible SVG imagery retained its aspect ratio and the layout recorded CLS 0. Empty image placeholders reported by the image scanner were not visible content.

Reference: image-summary, layout-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

browser-platform-hygiene

The page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load.

Pass
Source: pass; confidence: low
Implementation and method

Catalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

Method used or attempted: DOM and runtime surface review.

No geolocation or notification prompt appeared on load, paste prevention was not observed, and the page remained functional in the raw no-JavaScript crawler view.

Reference: discoverability-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Be discoverable · 4 tests
Be discoverable checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
title-and-description

The page has a unique, descriptive <title> and a meta description.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

Method used or attempted: DOM metadata probe.

The page title is Google, but meta[name=description] is absent.

Reference: page-probe-summary; retained-private

The public homepage has no meta description.

  • site-03-finding-06 · medium · The public homepage has no meta description.
crawlable-and-mobile-friendly

Links are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

Method used or attempted: DOM links, raw HTML comparison, and mobile layout inspection.

Most navigation links have real href values, but the viewport meta tag is absent and one More options anchor lacks href.

Reference: layout-summary, page-probe-summary; retained-private

Missing viewport metadata undermines mobile crawlability and presentation.

  • site-03-finding-07 · high · Missing viewport metadata undermines mobile crawlability and presentation.
canonical-and-indexing-signals

Public pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses.

Method used or attempted: Discoverability fetch and metadata inspection.

The raw document returned HTTP 200, no accidental noindex signal was present, and the canonical homepage URL remained on the reviewed origin.

Reference: discoverability-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

structured-and-shareable-metadata

Where the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses.

Method used or attempted: Entity and metadata review.

The entry is a utility search form, not an article, product, event, or other rich entity needing share cards or JSON-LD.

Reference: page-probe-summary; retained-private

The generic search entry does not represent a rich content entity.

Be private and secure · 4 tests
Be private and secure checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
secure-transport-and-headers

Served over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses.

Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence.

Retained evidence records 2 of 6 baseline security-header categories present. Cookie-attribute review records 2 of 2 records with Secure and 2 with HttpOnly. Names, values, routes, and raw headers remain private.

Reference: cookie-attribute-summary, security-header-summary; retained-private

The main document omits several baseline transport and content defenses.

  • site-03-finding-08 · high · The main document omits several baseline transport and content defenses.
data-minimisation-and-third-parties

No over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses.

Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence.

Retained evidence records 1 first-party origin categories and 4 third-party origin categories across 43 requests. Destinations, identifiers, routes, headers, and bodies remain private.

Reference: cookie-attribute-summary, network-summary, screenshot, tracker-summary; retained-private

The initial page has a material third-party and long-lived identifier footprint.

  • site-03-finding-09 · medium · The initial page has a material third-party and long-lived identifier footprint.
in-context-permissions-and-modern-auth

Permission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone.

Not applicable
Source: not-applicable; confidence: medium
Implementation and method

Catalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses.

Method used or attempted: Initial-load prompt and auth-surface review.

No browser permission prompt appeared. Authentication is only a cross-origin Sign in link and was outside the permitted journey.

Reference: page-probe-summary, screenshot; retained-private

The bounded audited path does not contain an authentication flow to assess passkeys or WebAuthn.

defensive-browser-policies

Browser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses.

Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence.

Retained evidence records 2 of 6 baseline browser-policy header categories present. Values and raw headers remain private.

Reference: security-header-summary; retained-private

Defensive browser policy coverage is incomplete.

  • site-03-finding-10 · high · Defensive browser policy coverage is incomplete.
Be resilient · 4 tests
Be resilient checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
progressive-enhancement

Core content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Method used or attempted: Permit-bound discoverability raw fetch and crawler screenshot.

With JavaScript disabled, the Google mark, search field, search buttons, navigation, and footer remain rendered and usable; the raw response was not a JS shell.

Reference: discoverability-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

resilient-runtime-behaviour

The page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

Method used or attempted: Journey viewport screenshot and dialog rectangle review.

At 780x493 the consent dialog extends below the viewport and its decision controls are not initially visible, so the mandatory overlay is cut off in a common short viewport.

Reference: screenshot; retained-private

The consent dialog is cut off in a short viewport.

  • site-03-finding-11 · high · The consent dialog is cut off in a short viewport.
offline-and-installable

Where the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.)

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

Method used or attempted: Manifest and service-worker registration probe.

No manifest or active service-worker registration exists. Search results intrinsically require a live network.

Reference: page-probe-summary; retained-private

This live search portal is intrinsically online; installability and offline result retrieval are not reasonable requirements.

network-and-http-failure-states

HTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states.

Not applicable
Source: not-applicable; confidence: medium
Implementation and method

Catalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses.

Method used or attempted: Site purpose and bounded policy review.

The audited entry is a network-dependent search portal and the strict journey did not permit simulated failing submissions.

Reference: journey-summary; retained-private

Search result retrieval is intrinsically online and no safe failure-producing action was permitted in this run.

Be internationalised · 3 tests
Be internationalised checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
lang-dir-and-logical-properties

Correct lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

Method used or attempted: DOM and stylesheet feature probe.

The document declares lang=en-GB and stylesheet inspection found logical inline properties.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

locale-aware-data

Dates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

Method used or attempted: Rendered content and control inventory review.

The bounded entry state displays no dates, numbers, currencies, durations, or calendar data.

Reference: page-probe-summary; retained-private

No locale-sensitive data is rendered on this path.

time-zone-correctness

Time handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

Method used or attempted: Rendered content and platform probe.

The page renders no events or time values requiring time-zone or DST handling.

Reference: page-probe-summary; retained-private

No time concepts are present on this path.

Be trustworthy · 4 tests
Be trustworthy checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-dark-patterns

No deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive).

Failed / issue
Source: issues; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

Method used or attempted: Consent wall visual and copy review.

The mandatory consent interstitial blocks the core search task on first load and describes measurement and personalised advertising before the user can proceed, although Reject all and Accept all are given equal visual weight on the full mobile capture.

Reference: screenshot; retained-private

A mandatory consent interstitial delays the primary task.

  • site-03-finding-12 · medium · A mandatory consent interstitial delays the primary task.
humane-error-handling

Forms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

Method used or attempted: Form requirements inspection.

The search field is not required and no validation or error state exists on the untouched entry path.

Reference: page-probe-summary; retained-private

No user-error validation state applies to the optional search field before submission.

trustworthy-input-assistance

Input is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

Method used or attempted: Search control semantics inspection.

The primary control is a labelled search combobox named q with native form submission; no paste prevention or obstructive formatting was observed.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

safe-commercial-and-account-flows

Checkout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses.

Method used or attempted: Consent choice visual review.

Reject all and Accept all are both present with equal visual treatment, More options is available, and privacy and terms links are visible.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

Be sustainable · 3 tests
Be sustainable checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
optimised-assets

Images and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Method used or attempted: Image and HAR resource review.

The seven image requests transferred only 972 bytes; the visible product mark is SVG and no oversized image was reported.

Reference: image-summary, network-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

no-wasteful-work

Background work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work.

Failed / issue
Source: issues; confidence: medium
Implementation and method

Catalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Method used or attempted: HAR and page-complexity review.

A roughly 502-node search entry transfers 715,201 bytes of JavaScript across eight script requests, disproportionate to the initial task surface.

Reference: network-summary, page-probe-summary; retained-private

JavaScript payload is disproportionate to the initial search surface.

  • site-03-finding-05 · medium · JavaScript payload is disproportionate to the initial search surface.
third-party-and-media-budget

Third-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible.

Failed / issue
Source: issues; confidence: medium
Implementation and method

Catalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses.

Method used or attempted: HAR and tracker review.

Eight third-party requests transferred 92,305 bytes across five third-party origins on the initial search page, before the user performs a search.

Reference: network-summary, tracker-summary; retained-private

The initial page has a material third-party and long-lived identifier footprint.

  • site-03-finding-09 · medium · The initial page has a material third-party and long-lived identifier footprint.
Be agent ready · 2 tests
Be agent ready checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
structured-agent-capabilities

Where it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

Method used or attempted: Raw HTML and semantic form inspection.

The primary capability is exposed as a standard server-rendered form with role search, a labelled q combobox, named submit controls, and real links, so agents need not infer a canvas-only interaction.

Reference: discoverability-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

on-device-inference

On-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses.

Method used or attempted: Page purpose and runtime capability review.

The entry is a remote web-search portal; no bounded local summarisation or inference task is exposed.

Reference: page-probe-summary; retained-private

On-device inference does not improve the audited search-entry task.

Be memory-efficient · 3 tests
Be memory-efficient checks for https://www.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-leak-under-repeated-interaction

Repeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses.

Method used or attempted: Compared independent heap summaries before and after ten bounded native scroll down/up cycles.

Post-cycle heap self size decreased by 7,716 bytes and node count decreased by 35 versus baseline, with no retained-growth signal.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

bounded-footprint

Heap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses.

Method used or attempted: Heap summary and DOM-size review.

The page held about [host/path omitted] MB V8 self size and roughly 502 DOM elements, proportionate for this feature-rich search entry.

Reference: memory-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

no-detached-dom-or-unbounded-listeners

No growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses.

Method used or attempted: Compared heap constructor summaries across repeated bounded scroll.

The post-cycle heap did not grow overall and no Detached constructor appeared among the largest retained constructors.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

4. https://www.reddit.com · 58 slots · available
Respect user preferences · 3 tests
Respect user preferences checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
respects-color-scheme

Honours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Desktop and dark/high-contrast screenshots are pixel-identical at 40,303 bytes; DOM reports color-scheme normal and fixed white surfaces.

Reference: No artifact reference; described-only

Dark preference is ignored

  • site-04-finding-01 · medium · Dark preference is ignored
respects-reduced-motion

Honours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Reduced-motion video captured one static frame and the probe found zero active animations.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

respects-contrast

Honours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Text and challenge controls remain visually legible in the high-contrast screenshot; no content disappears.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

Implement natural interactions · 3 tests
Implement natural interactions checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
view-transitions

State and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

Method used or attempted: Permit-bound recon plus strict journey replay result

Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Reference: No artifact reference; described-only

The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

scroll-driven-animations

Scroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The response is exactly one viewport tall and has no scroll-driven content.

Reference: No artifact reference; described-only

The response is exactly one viewport tall and has no scroll-driven content.

physical-gestures

Gesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The static challenge response exposes no gesture-driven UI.

Reference: No artifact reference; described-only

The static challenge response exposes no gesture-driven UI.

Provide guided navigation · 3 tests
Provide guided navigation checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
scroll-state-aware-chrome

Sticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The response has no scroll range or chrome whose state could react to scrolling.

Reference: No artifact reference; described-only

The response has no scroll range or chrome whose state could react to scrolling.

anchored-positioning

Tooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

No page-owned tooltip, popover, or anchored overlay is exposed without performing the prohibited CAPTCHA interaction.

Reference: No artifact reference; described-only

No page-owned tooltip, popover, or anchored overlay is exposed without performing the prohibited CAPTCHA interaction.

directs-attention

Navigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Desktop and mobile screenshots provide a direct H1, explanatory copy, and centered CAPTCHA with clear visual hierarchy.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

Maximize content, reduce noise · 3 tests
Maximize content, reduce noise checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-intrusive-interruptions

No intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

Method used or attempted: Permit-bound recon plus strict journey replay result

Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Reference: No artifact reference; described-only

The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

semantic-dismissible-primitives

Overlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

No page-owned dismissible dialog or popover is present in the observed response.

Reference: No artifact reference; described-only

No page-owned dismissible dialog or popover is present in the observed response.

reduced-chrome

Minimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Screenshots show minimal branding/footer chrome around the single challenge task.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

Adapt to the form factor · 3 tests
Adapt to the form factor checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
responsive-no-horizontal-scroll

Layout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Layout reports 0 horizontal overflow at both 1440x1000 and 360x800.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

component-level-responsiveness

Components adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Desktop/mobile screenshots and DOM CSS show the main surface constrains to 480px and footer switches to a column below 600px.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

input-modality-aware

Touch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Programmatic focus on the first link produced a native auto 1px focus outline.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

Support core task success · 3 tests
Support core task success checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
clear-purpose-and-primary-action

The page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The H1 and copy explicitly state the purpose and point to the reCAPTCHA as the only primary action.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

primary-flow-completion

The representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses.

Method used or attempted: Permit-bound recon plus strict journey replay result

Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Reference: No artifact reference; described-only

The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

clear-system-state-and-recovery

Loading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses.

Method used or attempted: Permit-bound recon plus strict journey replay result

Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Reference: No artifact reference; described-only

The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Be fast and stable · 5 tests
Be fast and stable checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
good-core-web-vitals

Core Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Trace reports LCP/FCP [host/path omitted], 0 long tasks, and 0ms total blocking time; layout reports CLS 0.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

visual-stability

No cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Both layout captures report CLS 0 with no layout shifts.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

efficient-main-thread

The main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Trace reports no long tasks and 0ms total blocking time over [host/path omitted].

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

efficient-resource-delivery

Critical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

HAR summary records 514,675 transferred bytes for the challenge; 346,849 bytes are third-party and a parser-inserted reCAPTCHA script has no async/defer.

Reference: No artifact reference; described-only

The minimal wall transfers half a megabyte, mostly third-party JavaScript

  • site-04-finding-10 · medium · The minimal wall transfers half a megabyte, mostly third-party JavaScript
trim-unused-and-duplicate-code

The page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses.

Method used or attempted: Permit-bound recon plus strict journey replay result

Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Reference: No artifact reference; described-only

The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Be inclusive · 5 tests
Be inclusive checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
names-roles-labels

Interactive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The evaluated control inventory shows the 100x32 Reddit logo link has no text, aria-label, or title; the image audit reports its image has no alt.

Reference: No artifact reference; described-only

The Reddit logo link has no accessible name

  • site-04-finding-02 · medium · The Reddit logo link has no accessible name
sufficient-contrast

Text and essential UI meet WCAG colour-contrast minimums against their background.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Visual review of desktop, mobile, and increased-contrast captures shows readable dark text and visible controls on white surfaces.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

structure-and-focus

Heading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The probe found one H1 and a visible native focus outline, but zero main, nav, header, or footer landmark elements.

Reference: No artifact reference; described-only

The page has no semantic landmarks

  • site-04-finding-04 · medium · The page has no semantic landmarks
legible-text

Text is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Desktop/mobile screenshots show a 24px heading and 16px centered body copy with readable line lengths.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

zoom-reflow-targets-and-media

The experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

At 360px there is no overflow, but the evaluated footer links are only 14 CSS px tall, below a comfortable touch target.

Reference: No artifact reference; described-only

Footer links have undersized touch targets

  • site-04-finding-05 · medium · Footer links have undersized touch targets
Follow best practices · 3 tests
Follow best practices checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-console-errors

The page loads without console errors or uncaught exceptions.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

Method used or attempted: Permit-bound recon plus strict journey replay result

Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Reference: No artifact reference; described-only

The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

sound-document-and-assets

Valid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.)

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The image audit reports the sole image lacks width and height attributes; DOM confirms it renders at 100x28 from a 300x84 intrinsic asset.

Reference: No artifact reference; described-only

The logo image omits intrinsic dimensions

  • site-04-finding-03 · low · The logo image omits intrinsic dimensions
browser-platform-hygiene

The page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

DOM confirms HTML doctype, UTF-8, viewport metadata, and no permission prompt; the page uses HTTPS.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

Be discoverable · 4 tests
Be discoverable checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
title-and-description

The page has a unique, descriptive <title> and a meta description.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The rendered title is present, but the probe found no meta description.

Reference: No artifact reference; described-only

The anti-bot response has no meta description

  • site-04-finding-06 · medium · The anti-bot response has no meta description
crawlable-and-mobile-friendly

Links are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The discoverability primitive found only 4% rendered-word coverage in raw HTML; title and H1 are absent from raw HTML.

Reference: No artifact reference; described-only

The rendered challenge is almost absent without JavaScript

  • site-04-finding-07 · high · The rendered challenge is almost absent without JavaScript
canonical-and-indexing-signals

Public pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The probe found no canonical link and no robots meta on a 200 response that renders an anti-bot interstitial.

Reference: No artifact reference; described-only

Indexing and sharing metadata are missing

  • site-04-finding-08 · medium · Indexing and sharing metadata are missing
structured-and-shareable-metadata

Where the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The probe found zero JSON-LD blocks and zero Open Graph metadata.

Reference: No artifact reference; described-only

Indexing and sharing metadata are missing

  • site-04-finding-08 · medium · Indexing and sharing metadata are missing
Be private and secure · 4 tests
Be private and secure checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
secure-transport-and-headers

Served over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses.

Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence.

Retained evidence records 3 of 6 baseline security-header categories present. Cookie-attribute review records 1 of 1 records with Secure and 0 with HttpOnly. Names, values, routes, and raw headers remain private.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

data-minimisation-and-third-parties

No over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses.

Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence.

Retained evidence records 1 first-party origin categories and 3 third-party origin categories across 8 requests. Destinations, identifiers, routes, headers, and bodies remain private.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

in-context-permissions-and-modern-auth

Permission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses.

Method used or attempted: Permit-bound recon plus strict journey replay result

Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Reference: No artifact reference; described-only

The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

defensive-browser-policies

Browser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses.

Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence.

Retained evidence records 3 of 6 baseline browser-policy header categories present. Values and raw headers remain private.

Reference: No artifact reference; described-only

Defensive response policies are incomplete

  • site-04-finding-09 · high · Defensive response policies are incomplete
Be resilient · 4 tests
Be resilient checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
progressive-enhancement

Core content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The no-JavaScript/raw comparison exposes only 4% of rendered content words and omits the rendered title and H1.

Reference: No artifact reference; described-only

The rendered challenge is almost absent without JavaScript

  • site-04-finding-07 · high · The rendered challenge is almost absent without JavaScript
resilient-runtime-behaviour

The page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The challenge rendered consistently across desktop, mobile, reduced-motion, and crawler captures without runtime-visible breakage.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

offline-and-installable

Where the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.)

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Installability is not applicable to this transient anti-bot response, and the intended Reddit application was blocked.

Reference: No artifact reference; described-only

Installability is not applicable to this transient anti-bot response, and the intended Reddit application was blocked.

network-and-http-failure-states

HTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses.

Method used or attempted: Permit-bound recon plus strict journey replay result

Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Reference: No artifact reference; described-only

The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Be internationalised · 3 tests
Be internationalised checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
lang-dir-and-logical-properties

Correct lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The document declares lang=en; observed text is consistently left-to-right and mobile CSS adapts without clipping.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

locale-aware-data

Dates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The observed response presents no locale-sensitive dates, numbers, currency, or calendars.

Reference: No artifact reference; described-only

The observed response presents no locale-sensitive dates, numbers, currency, or calendars.

time-zone-correctness

Time handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The observed response presents no time-zone-dependent data or scheduling.

Reference: No artifact reference; described-only

The observed response presents no time-zone-dependent data or scheduling.

Be trustworthy · 4 tests
Be trustworthy checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-dark-patterns

No deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive).

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The wall plainly explains why the challenge exists and presents no deceptive choices, confirmshaming, or commercial pressure.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

humane-error-handling

Forms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

Method used or attempted: Permit-bound recon plus strict journey replay result

Attempted permit-bound DOM/visual/journey inspection. The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

Reference: No artifact reference; described-only

The exact URL rendered a reCAPTCHA anti-bot wall instead of Reddit content, and the strict pre-replayed journey aborted after its baseline navigation; direct evidence for the intended application was unavailable.

trustworthy-input-assistance

Input is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The only page form control is reCAPTCHA implementation plumbing; there is no address, payment, sign-in, or signup field to autofill.

Reference: No artifact reference; described-only

The only page form control is reCAPTCHA implementation plumbing; there is no address, payment, sign-in, or signup field to autofill.

safe-commercial-and-account-flows

Checkout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

No commercial, subscription, checkout, or account-management flow is exposed on the observed response.

Reference: No artifact reference; described-only

No commercial, subscription, checkout, or account-management flow is exposed on the observed response.

Be sustainable · 3 tests
Be sustainable checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
optimised-assets

Images and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The one page-owned visual is an inline SVG logo; no raster transfer is used for it.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

no-wasteful-work

Background work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Trace shows no long tasks and the reduced-motion capture produced no continuing animation activity.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

third-party-and-media-budget

Third-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

HAR summary records 67% of transfer bytes from third parties, dominated by 340,589 bytes of reCAPTCHA JavaScript for a minimal wall.

Reference: No artifact reference; described-only

The minimal wall transfers half a megabyte, mostly third-party JavaScript

  • site-04-finding-10 · medium · The minimal wall transfers half a megabyte, mostly third-party JavaScript
Be agent ready · 2 tests
Be agent ready checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
structured-agent-capabilities

Where it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The anti-bot response is not an agent-capability surface, and the intended Reddit content could not be reached.

Reference: No artifact reference; described-only

The anti-bot response is not an agent-capability surface, and the intended Reddit content could not be reached.

on-device-inference

On-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The observed challenge has no inference use case or built-in AI surface.

Reference: No artifact reference; described-only

The observed challenge has no inference use case or built-in AI surface.

Be memory-efficient · 3 tests
Be memory-efficient checks for https://www.reddit.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-leak-under-repeated-interaction

Repeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

The document has no scroll range and the strict journey did not permit a repeatable interaction; fabricating one would not test a real user action.

Reference: No artifact reference; described-only

The document has no scroll range and the strict journey did not permit a repeatable interaction; fabricating one would not test a real user action.

bounded-footprint

Heap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

Baseline heap summary is 6,490,062 bytes for 105,620 nodes, proportionate to the challenge plus reCAPTCHA.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

no-detached-dom-or-unbounded-listeners

No growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, trace, HAR, security, discoverability, or heap evidence as applicable

After ten bounded scroll attempts, heap growth was only 11,487 bytes and 26 nodes; no Detached* constructor appears in either top-constructor summary.

Reference: No artifact reference; described-only

The retained evidence directly supported this check under the captured conditions.

5. https://www.amazon.com · 58 slots · available
Respect user preferences · 3 tests
Respect user preferences checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
respects-color-scheme

Honours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The dark-mode probe reports body rgb(255, 255, 255), text rgb(15, 17, 17), and color-scheme normal while the dark media query matches.

Reference: page-probe-summary, screenshot; retained-private

The homepage remains light when the user requests dark mode.

  • site-05-finding-01 · high · The homepage remains light when the user requests dark mode.
respects-reduced-motion

Honours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

Reduced-motion emulation matched and the page exposed zero active animations with automatic scroll behavior.

Reference: layout-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

respects-contrast

Honours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

The high-contrast screenshot retained readable navigation, promotional text, and controls without disappearing content.

Reference: other-private-evidence; retained-private

The retained evidence directly supported this check under the captured conditions.

Implement natural interactions · 3 tests
Implement natural interactions checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
view-transitions

State and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

The only replayed baseline triggered an automatic POST and the strict journey aborted; no safe state transition was available to inspect.

Reference: No artifact reference; described-only

The only replayed baseline triggered an automatic POST and the strict journey aborted; no safe state transition was available to inspect.

scroll-driven-animations

Scroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

The bounded-scroll journey step was skipped after the baseline mutation blocker, so scroll-linked behavior could not be exercised safely.

Reference: No artifact reference; described-only

The bounded-scroll journey step was skipped after the baseline mutation blocker, so scroll-linked behavior could not be exercised safely.

physical-gestures

Gesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

No permitted reversible gesture surface could be exercised after the strict journey aborted.

Reference: No artifact reference; described-only

No permitted reversible gesture surface could be exercised after the strict journey aborted.

Provide guided navigation · 3 tests
Provide guided navigation checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
scroll-state-aware-chrome

Sticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

The planned bounded scroll was skipped after the automatic POST blocker.

Reference: No artifact reference; described-only

The planned bounded scroll was skipped after the automatic POST blocker.

anchored-positioning

Tooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Opening menus or overlays was outside the already replayed journey and no additional control interaction was permitted.

Reference: No artifact reference; described-only

Opening menus or overlays was outside the already replayed journey and no additional control interaction was permitted.

directs-attention

Navigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

The strict journey could not advance beyond baseline, so post-navigation attention movement was not observable.

Reference: No artifact reference; described-only

The strict journey could not advance beyond baseline, so post-navigation attention movement was not observable.

Maximize content, reduce noise · 3 tests
Maximize content, reduce noise checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-intrusive-interruptions

No intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

Independent desktop and mobile load screenshots showed the page content directly with no modal interstitial obscuring it.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

semantic-dismissible-primitives

Overlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The DOM probe found many div role=dialog nodes but zero dialog, popover, or details elements.

Reference: page-probe-summary; retained-private

Overlay semantics rely on custom role-based containers rather than native primitives.

  • site-05-finding-02 · medium · Overlay semantics rely on custom role-based containers rather than native primitives.
reduced-chrome

Minimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

Desktop and mobile screenshots devote the first viewport to two navigation rows and large promotional panels; the mobile capture reaches footer chrome with little product-detail content.

Reference: screenshot; retained-private

Navigation, promotional chrome, location messaging, and dense merchandising compete with the shopping content.

  • site-05-finding-03 · medium · Navigation, promotional chrome, location messaging, and dense merchandising compete with the shopping content.
Adapt to the form factor · 3 tests
Adapt to the form factor checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
responsive-no-horizontal-scroll

Layout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The mobile layout reports no viewport meta and 620 px horizontal overflow at 360x800; desktop also reports 220 px overflow. The mobile screenshot is visibly clipped to a desktop-width surface.

Reference: layout-summary, screenshot; retained-private

The entry page does not establish a mobile viewport and overflows horizontally.

  • site-05-finding-04 · critical · The entry page does not establish a mobile viewport and overflows horizontally.
component-level-responsiveness

Components adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Rendered evidence proves page-level overflow, but authored CSS and safe component resizing were unavailable, so container-level behavior could not be conclusively judged.

Reference: No artifact reference; described-only

Rendered evidence proves page-level overflow, but authored CSS and safe component resizing were unavailable, so container-level behavior could not be conclusively judged.

input-modality-aware

Touch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The focus probe reports 0 px outline on the Amazon home link, delivery control, search field, search submit, language, account, orders, and cart controls.

Reference: page-probe-summary; retained-private

Keyboard focus is not visibly exposed on several primary controls.

  • site-05-finding-05 · high · Keyboard focus is not visibly exposed on several primary controls.
Support core task success · 3 tests
Support core task success checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
clear-purpose-and-primary-action

The page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

The Amazon identity, prominent search control, category navigation, and shopping promotions make the commerce purpose and next action clear.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

primary-flow-completion

The representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Search, account, product, cart, and checkout routes were outside the exact-URL permit, and the strict baseline aborted on an automatic POST.

Reference: No artifact reference; described-only

Search, account, product, cart, and checkout routes were outside the exact-URL permit, and the strict baseline aborted on an automatic POST.

clear-system-state-and-recovery

Loading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Failure, empty, success, and retry states require prohibited route or mutation testing and were not safely reachable.

Reference: No artifact reference; described-only

Failure, empty, success, and retry states require prohibited route or mutation testing and were not safely reachable.

Be fast and stable · 5 tests
Be fast and stable checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
good-core-web-vitals

Core Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

The trace did not yield FCP or LCP and no permitted interaction was available for an INP-like measurement.

Reference: No artifact reference; described-only

The trace did not yield FCP or LCP and no permitted interaction was available for an INP-like measurement.

visual-stability

No cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

Desktop and mobile layout observers recorded CLS 0 with no shift entries during their capture windows.

Reference: layout-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

efficient-main-thread

The main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

The performance trace recorded zero long tasks and 0 ms total blocking time over [host/path omitted] seconds.

Reference: performance-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

efficient-resource-delivery

Critical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The HAR records 43 requests and 1,101,460 transferred bytes, including ten parser-inserted render-blocking candidates and a 286,614-byte WAF script.

Reference: network-summary; retained-private

The landing load has a long render-blocking resource chain and a large anti-bot script cost.

  • site-05-finding-06 · high · The landing load has a long render-blocking resource chain and a large anti-bot script cost.
trim-unused-and-duplicate-code

The page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

No permit-aware code-coverage evidence was gathered, so shipped bytes cannot be classified as used or unused.

Reference: No artifact reference; described-only

No permit-aware code-coverage evidence was gathered, so shipped bytes cannot be classified as used or unused.

Be inclusive · 5 tests
Be inclusive checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
names-roles-labels

Interactive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The accessibility probe found 56 unnamed controls among 268 controls, while the image audit found three images without alt in its captured variant.

Reference: image-summary, page-probe-summary; retained-private

A material number of interactive controls lack a detectable accessible name.

  • site-05-finding-07 · high · A material number of interactive controls lack a detectable accessible name.
sufficient-contrast

Text and essential UI meet WCAG colour-contrast minimums against their background.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

Desktop, mobile, and prefers-contrast screenshots show legible text and controls against their immediate surfaces; no visual contrast failure was observed.

Reference: other-private-evidence, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

structure-and-focus

Heading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The probe reports zero H1 elements and 0 px outlines on primary navigation and search controls.

Reference: page-probe-summary; retained-private

The page has no H1 and several major controls suppress visible focus.

  • site-05-finding-08 · high · The page has no H1 and several major controls suppress visible focus.
legible-text

Text is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

Captured headings, navigation labels, and body text are readable and not clipped vertically in the inspected states.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

zoom-reflow-targets-and-media

The experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

At 360x800 the layout viewport remains 980 px wide and horizontal overflow is 620 px, making content clipped and zoomed out.

Reference: layout-summary, screenshot; retained-private

Narrow-screen reflow fails because the page omits a viewport declaration.

  • site-05-finding-09 · critical · Narrow-screen reflow fails because the page omits a viewport declaration.
Follow best practices · 3 tests
Follow best practices checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-console-errors

The page loads without console errors or uncaught exceptions.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

The journey console dimension was blocked by the unavailable Stage 2 collector and no independent console-event capture was retained.

Reference: No artifact reference; described-only

The journey console dimension was blocked by the unavailable Stage 2 collector and no independent console-event capture was retained.

sound-document-and-assets

Valid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.)

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The image audit found 11 of 12 captured images without width/height, two oversized images, seven without srcset, and eight legacy-format images.

Reference: image-summary; retained-private

Image sizing and delivery are structurally fragile.

  • site-05-finding-10 · medium · Image sizing and delivery are structurally fragile.
browser-platform-hygiene

The page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Deprecation, BFCache, source-map, prompt timing, and vulnerable-library checks require evidence not available within the bounded entry-only run.

Reference: No artifact reference; described-only

Deprecation, BFCache, source-map, prompt timing, and vulnerable-library checks require evidence not available within the bounded entry-only run.

Be discoverable · 4 tests
Be discoverable checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
title-and-description

The page has a unique, descriptive <title> and a meta description.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

The rendered probe found a descriptive title and a detailed meta description.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

crawlable-and-mobile-friendly

Links are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The discoverability probe reports 1% raw-to-rendered content coverage and isJsShell=true; the page also lacks viewport meta.

Reference: discoverability-summary, layout-summary; retained-private

The rendered landing page is not mobile-friendly and most content is absent from the raw HTML response.

  • site-05-finding-11 · high · The rendered landing page is not mobile-friendly and most content is absent from the raw HTML response.
canonical-and-indexing-signals

Public pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

Discoverability evidence records fetchedStatus 202, 1% coverage, no title in raw HTML, and no raw meta description, despite a rendered canonical URL.

Reference: discoverability-summary; retained-private

The non-JavaScript fetch returns an HTTP 202 challenge shell rather than indexable page content.

  • site-05-finding-12 · high · The non-JavaScript fetch returns an HTTP 202 challenge shell rather than indexable page content.
structured-and-shareable-metadata

Where the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The probe found two Open Graph tags but no JSON-LD blocks for the organization or commerce destination.

Reference: page-probe-summary; retained-private

The homepage exposes limited entity metadata.

  • site-05-finding-13 · low · The homepage exposes limited entity metadata.
Be private and secure · 4 tests
Be private and secure checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
secure-transport-and-headers

Served over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses.

Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence.

Retained evidence records 0 of 6 baseline security-header categories present. Cookie-attribute review records 3 of 7 records with Secure and 0 with HttpOnly. Names, values, routes, and raw headers remain private.

Reference: cookie-attribute-summary, security-header-summary; retained-private

The captured main response lacks baseline defensive headers and several cookies have weak attributes.

  • site-05-finding-14 · high · The captured main response lacks baseline defensive headers and several cookies have weak attributes.
data-minimisation-and-third-parties

No over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses.

Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence.

Retained evidence records 1 first-party origin categories and 3 third-party origin categories across 43 requests. Destinations, identifiers, routes, headers, and bodies remain private.

Reference: network-summary, tracker-summary; retained-private

Most landing-load requests and bytes are delivered from third-party-classified origins.

  • site-05-finding-15 · medium · Most landing-load requests and bytes are delivered from third-party-classified origins.
in-context-permissions-and-modern-auth

Permission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Authentication is on an unpermitted path and no permission prompt fired on the captured landing state; modern-auth support could not be verified.

Reference: No artifact reference; described-only

Authentication is on an unpermitted path and no permission prompt fired on the captured landing state; modern-auth support could not be verified.

defensive-browser-policies

Browser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses.

Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence.

Retained evidence records 0 of 6 baseline browser-policy header categories present. Values and raw headers remain private.

Reference: security-header-summary; retained-private

Browser-enforced defense headers are absent in the captured response.

  • site-05-finding-16 · high · Browser-enforced defense headers are absent in the captured response.
Be resilient · 4 tests
Be resilient checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
progressive-enhancement

Core content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The raw response has 14 content tokens versus 304 rendered tokens, 1% coverage, no raw title, and no raw description; the crawler screenshot is a challenge/blank shell.

Reference: discoverability-summary; retained-private

Core landing content is effectively unavailable without JavaScript.

  • site-05-finding-17 · high · Core landing content is effectively unavailable without JavaScript.
resilient-runtime-behaviour

The page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Menus and async state could not be exercised because the strict replay aborted after baseline network mutation.

Reference: No artifact reference; described-only

Menus and async state could not be exercised because the strict replay aborted after baseline network mutation.

offline-and-installable

Where the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.)

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

Method used or attempted: Applicability judgement from the audited commerce entry state.

Offline/installability is contextual and not essential to this public commerce landing page; progressive enhancement is judged separately.

Reference: No artifact reference; described-only

Offline/installability is contextual and not essential to this public commerce landing page; progressive enhancement is judged separately.

network-and-http-failure-states

HTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Network failure simulation and error routes were outside the bounded execution plan.

Reference: No artifact reference; described-only

Network failure simulation and error routes were outside the bounded execution plan.

Be internationalised · 3 tests
Be internationalised checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
lang-dir-and-logical-properties

Correct lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

The rendered document declares lang=en-us and computes ltr direction; the captured English content follows that direction.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

locale-aware-data

Dates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

Method used or attempted: Applicability judgement from the audited commerce entry state.

No date, currency, duration, or other locale-sensitive value was present in the audited entry state.

Reference: No artifact reference; described-only

No date, currency, duration, or other locale-sensitive value was present in the audited entry state.

time-zone-correctness

Time handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

Method used or attempted: Applicability judgement from the audited commerce entry state.

No time or event data was present in the audited entry state.

Reference: No artifact reference; described-only

No time or event data was present in the audited entry state.

Be trustworthy · 4 tests
Be trustworthy checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-dark-patterns

No deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive).

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

No consent wall, confirmshaming, forced continuity, disguised ad control, or purchase prompt appeared on the audited entry state.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

humane-error-handling

Forms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Form mutation and submission are hard-denied by the journey protocol.

Reference: No artifact reference; described-only

Form mutation and submission are hard-denied by the journey protocol.

trustworthy-input-assistance

Input is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Account, address, and payment forms were outside the exact-URL boundary, so autofill assistance could not be assessed.

Reference: No artifact reference; described-only

Account, address, and payment forms were outside the exact-URL boundary, so autofill assistance could not be assessed.

safe-commercial-and-account-flows

Checkout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses.

Method used or attempted: Attempted bounded entry-page inspection and/or strict pre-replayed journey.

Cart, checkout, purchase, authentication, account management, and cancellation are hard-denied or outside the exact-URL permit.

Reference: No artifact reference; described-only

Cart, checkout, purchase, authentication, account management, and cancellation are hard-denied or outside the exact-URL permit.

Be sustainable · 3 tests
Be sustainable checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
optimised-assets

Images and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The image audit reports two oversized images, seven without srcset, eight legacy-format images, and no modern-format image in the captured set.

Reference: image-summary; retained-private

Captured image delivery omits responsive sizing and modern formats.

  • site-05-finding-18 · medium · Captured image delivery omits responsive sizing and modern formats.
no-wasteful-work

Background work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

The flow baseline was aborted after an automatic WAF verification POST attempt; the independent HAR still recorded 43 requests and [host/path omitted] MB transferred.

Reference: journey-summary, network-summary; retained-private

The baseline immediately starts substantial challenge and telemetry work.

  • site-05-finding-19 · medium · The baseline immediately starts substantial challenge and telemetry work.
third-party-and-media-budget

Third-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM/evaluate, layout, HAR, discoverability, security, image, or journey evidence as cited.

Third-party-classified traffic accounts for 794,498 bytes and 37 requests; scripts alone transfer 449,270 bytes.

Reference: network-summary; retained-private

Third-party-classified delivery dominates the landing-page transfer budget.

  • site-05-finding-20 · high · Third-party-classified delivery dominates the landing-page transfer budget.
Be agent ready · 2 tests
Be agent ready checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
structured-agent-capabilities

Where it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

Method used or attempted: Applicability judgement from the audited commerce entry state.

Agentic commerce is an emerging opportunity, but no declared agent-facing intent or surface was available; absence is not penalized.

Reference: No artifact reference; described-only

Agentic commerce is an emerging opportunity, but no declared agent-facing intent or surface was available; absence is not penalized.

on-device-inference

On-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses.

Method used or attempted: Applicability judgement from the audited commerce entry state.

No on-device inference use case or declared intent was present on the commerce landing state.

Reference: No artifact reference; described-only

No on-device inference use case or declared intent was present on the commerce landing state.

Be memory-efficient · 3 tests
Be memory-efficient checks for https://www.amazon.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-leak-under-repeated-interaction

Repeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

After ten bounded scroll-to-500-and-back cycles, summarized heap size rose only 200,845 bytes ([host/path omitted]%) and node count 1,851 ([host/path omitted]%) across independent captures, with no unbounded-growth signal.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

bounded-footprint

Heap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

The rendered commerce landing state used about [host/path omitted] MB summarized self-size; this is substantial but proportionate to a dense, script-heavy homepage and did not grow materially in the bounded repeat.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

no-detached-dom-or-unbounded-listeners

No growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses.

Method used or attempted: Permit-bound evidence review for the entry state.

No Detached* constructor appeared among top heap constructors, and closure count changed by only 74 across the repeated-scroll pair; no accumulating pattern was observed.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

6. https://en.wikipedia.org · 58 slots · available

Method warning: this site has a report, but its no-console-errors pass is method-invalid and is labelled in the table.

Respect user preferences · 3 tests
Respect user preferences checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
respects-color-scheme

Honours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

Desktop and prefers-color-scheme: dark screenshots are byte-identical; computed color-scheme is normal and the rendered body remains rgb(248,249,250).

Reference: page-probe-summary, screenshot; retained-private

The page does not follow the system dark-color preference by default.

  • site-06-finding-01 · medium · The page does not follow the system dark-color preference by default.
respects-reduced-motion

Honours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Reduced-motion probe matched the preference, reported scroll-behavior auto, and found zero active animations; CSS also contains a reduced-motion rule.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

respects-contrast

Honours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

The prefers-contrast: more screenshot retains visible dark text, blue links, borders, and controls across the full page.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Implement natural interactions · 3 tests
Implement natural interactions checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
view-transitions

State and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

CSS probe found no view-transition rules or declarations; the MPA exposes hundreds of same-origin article links.

Reference: journey-summary, page-probe-summary; retained-private

Same-origin document navigation has no View Transition enhancement.

  • site-06-finding-02 · low · Same-origin document navigation has no View Transition enhancement.
scroll-driven-animations

Scroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

The audited main page and bounded-scroll journey contain no parallax, scrollytelling, reveal, or other scroll-linked animation to implement.

physical-gestures

Gesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

The recorded bounded wheel scroll moved 246 CSS px using native scrolling, with no pointer interception, unexpected state change, or network request.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Provide guided navigation · 3 tests
Provide guided navigation checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
scroll-state-aware-chrome

Sticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

The bounded-scroll screenshot shows content moving predictably without a static header covering the reading area; no JS scroll toggling was detected.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

anchored-positioning

Tooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

No tooltip, popover, or menu overlay was opened or present in the bounded journey, so there is no positioned overlay to assess.

directs-attention

Navigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

The baseline has a clear selected Main Page tab, section headings, blue link affordances, and a visible skip link focus treatment.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Maximize content, reduce noise · 3 tests
Maximize content, reduce noise checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-intrusive-interruptions

No intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Baseline and full-page screenshots contain no load-time modal, interstitial, consent wall, or banner obscuring content.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

semantic-dismissible-primitives

Overlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

DOM inspection found no active overlay in the audited state; existing controls use native form and button elements rather than a visible ad-hoc modal.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

reduced-chrome

Minimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Screenshots show the viewport dominated by encyclopedia content with compact navigation and no decorative application frame.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Adapt to the form factor · 3 tests
Adapt to the form factor checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
responsive-no-horizontal-scroll

Layout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

At requested 360x800, layout reports a 1120 CSS-pixel client width and the DOM declares meta viewport width=1120; the full-page mobile capture shows desktop columns and tiny text.

Reference: layout-summary, page-probe-summary, screenshot; retained-private

The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.

  • site-06-finding-03 · high · The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.
component-level-responsiveness

Components adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

At requested 360x800, layout reports a 1120 CSS-pixel client width and the DOM declares meta viewport width=1120; the full-page mobile capture shows desktop columns and tiny text.

Reference: layout-summary, page-probe-summary, screenshot; retained-private

The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.

  • site-06-finding-03 · high · The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.
input-modality-aware

Touch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

The focus/target probe measured Donate, Create account, Log in and many content links at 15 to 17 px high, below a comfortable touch-target size.

Reference: page-probe-summary, screenshot; retained-private

Many important navigation links have touch boxes only 15 to 17 CSS pixels high.

  • site-06-finding-04 · medium · Many important navigation links have touch boxes only 15 to 17 CSS pixels high.
Support core task success · 3 tests
Support core task success checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
clear-purpose-and-primary-action

The page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

The first viewport clearly says “Welcome to Wikipedia”, exposes search/navigation, and immediately presents the featured article.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

primary-flow-completion

The representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Both supplied journey actions completed: same-origin baseline load and bounded scroll, with 2 planned and 2 performed actions and no blocked action.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

clear-system-state-and-recovery

Loading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

The supplied non-mutating baseline and bounded-scroll journey has no loading, empty, error, success, or partial-completion state.

Be fast and stable · 5 tests
Be fast and stable checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
good-core-web-vitals

Core Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Permit-bound trace measured LCP/FCP [host/path omitted] ms and TBT [host/path omitted] ms; layout measured CLS [host/path omitted], all within good lab ranges.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

visual-stability

No cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Layout observers measured CLS [host/path omitted] desktop and 0 mobile, with only one very small desktop shift.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

efficient-main-thread

The main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Trace found one [host/path omitted] ms long task and [host/path omitted] ms total blocking time, while the page reached LCP in [host/path omitted] s.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

efficient-resource-delivery

Critical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

Image audit found 16 below-fold images without loading=lazy, 17 legacy-format images, 3 oversized images, and one missing dimension; HAR lists 10 images without cache headers, including a 129 KB portrait.

Reference: image-summary, layout-summary, network-summary; retained-private

Below-fold imagery is eagerly loaded and several image delivery details waste bytes or weaken stability.

  • site-06-finding-07 · medium · Below-fold imagery is eagerly loaded and several image delivery details waste bytes or weaken stability.
trim-unused-and-duplicate-code

The page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

HAR transferred 682,131 bytes total with 281,579 script bytes and no third-party tracker package; this is proportionate for the feature-rich reference landing page.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Be inclusive · 5 tests
Be inclusive checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
names-roles-labels

Interactive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Probe found zero unlabeled visible controls; all images carry an alt attribute and the page exposes 13 landmarks.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

sufficient-contrast

Text and essential UI meet WCAG colour-contrast minimums against their background.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Desktop and high-contrast screenshots retain legible dark text, blue links, controls, and section boundaries; no content disappears in the preference condition.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

structure-and-focus

Heading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Visible heading sequence is one H1 followed by eight H2s with no skipped levels; first keyboard target has a 2 px solid focus outline.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

legible-text

Text is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

At requested 360x800, layout reports a 1120 CSS-pixel client width and the DOM declares meta viewport width=1120; the full-page mobile capture shows desktop columns and tiny text.

Reference: layout-summary, page-probe-summary, screenshot; retained-private

The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.

  • site-06-finding-03 · high · The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.
zoom-reflow-targets-and-media

The experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

At requested 360x800, layout reports a 1120 CSS-pixel client width and the DOM declares meta viewport width=1120; the full-page mobile capture shows desktop columns and tiny text. The focus/target probe measured Donate, Create account, Log in and many content links at 15 to 17 px high, below a comfortable touch-target size.

Reference: layout-summary, page-probe-summary, screenshot; retained-private

The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down. Many important navigation links have touch boxes only 15 to 17 CSS pixels high.

  • site-06-finding-03 · high · The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.
  • site-06-finding-04 · medium · Many important navigation links have touch boxes only 15 to 17 CSS pixels high.
Follow best practices · 3 tests
Follow best practices checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-console-errors

The page loads without console errors or uncaught exceptions.

Pass
Source: pass; confidence: medium

Method invalid: this console pass used absence of a surfaced error even though the authoritative console collector was unavailable. Treat it as unreliable evidence, not a valid pass.

Implementation and method

Catalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

All permit-bound DOM, layout, trace, HAR, evaluate, and journey navigations completed without an uncaught exception surfaced by the harness.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

sound-document-and-assets

Valid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.)

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

Image audit found 16 below-fold images without loading=lazy, 17 legacy-format images, 3 oversized images, and one missing dimension; HAR lists 10 images without cache headers, including a 129 KB portrait.

Reference: image-summary, layout-summary, network-summary; retained-private

Below-fold imagery is eagerly loaded and several image delivery details waste bytes or weaken stability.

  • site-06-finding-07 · medium · Below-fold imagery is eagerly loaded and several image delivery details waste bytes or weaken stability.
browser-platform-hygiene

The page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

The document uses HTML5, no inline event-handler attributes were found, and no permission prompt appeared during repeated permit-bound loads.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Be discoverable · 4 tests
Be discoverable checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
title-and-description

The page has a unique, descriptive <title> and a meta description.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

DOM and discoverability probes both report metaDescription null, although the title survives without JavaScript.

Reference: discoverability-summary, page-probe-summary; retained-private

The Main Page has no meta description.

  • site-06-finding-05 · low · The Main Page has no meta description.
crawlable-and-mobile-friendly

Links are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

At requested 360x800, layout reports a 1120 CSS-pixel client width and the DOM declares meta viewport width=1120; the full-page mobile capture shows desktop columns and tiny text.

Reference: layout-summary, page-probe-summary, screenshot; retained-private

The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.

  • site-06-finding-03 · high · The mobile viewport is forced to a 1120 CSS-pixel desktop canvas and scaled down.
canonical-and-indexing-signals

Public pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Fetch returned 200 after the canonical root redirect; rendered DOM declares canonical [origin omitted] and robots allows previews.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

structured-and-shareable-metadata

Where the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

The page contains one JSON-LD block plus matching og:title and og:image metadata for the visible Main Page content.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Be private and secure · 4 tests
Be private and secure checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
secure-transport-and-headers

Served over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses.

Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence.

Retained evidence records 3 of 6 baseline security-header categories present. Cookie-attribute review records 6 of 7 records with Secure and 4 with HttpOnly. Names, values, routes, and raw headers remain private.

Reference: cookie-attribute-summary, security-header-summary; retained-private

Browser defenses and one session-like cookie are weaker than expected.

  • site-06-finding-06 · high · Browser defenses and one session-like cookie are weaker than expected.
data-minimisation-and-third-parties

No over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses.

Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence.

Retained evidence records 1 first-party origin categories and 2 third-party origin categories across 40 requests. Destinations, identifiers, routes, headers, and bodies remain private.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

in-context-permissions-and-modern-auth

Permission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

The public anonymous main-page journey requests no browser permission and contains no authentication step.

defensive-browser-policies

Browser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses.

Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence.

Retained evidence records 3 of 6 baseline browser-policy header categories present. Values and raw headers remain private.

Reference: cookie-attribute-summary, security-header-summary; retained-private

Browser defenses and one session-like cookie are weaker than expected.

  • site-06-finding-06 · high · Browser defenses and one session-like cookie are weaker than expected.
Be resilient · 4 tests
Be resilient checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
progressive-enhancement

Core content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Discoverability measured 98% content coverage without JavaScript, isJsShell false, and matching browser/crawler views.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

resilient-runtime-behaviour

The page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

The recorded load and scroll retained the same URL, 2,375-node state and 142 listeners, with no overlay clipping or state loss.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

offline-and-installable

Where the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.)

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

Wikipedia Main Page is a server-rendered reference document rather than an installable application; PWA installability is not needed for this path.

network-and-http-failure-states

HTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

The reviewed exact-URL journey contains no client-side mutation or async task with an application-owned retry/error state; testing a different error URL was explicitly prohibited.

Be internationalised · 3 tests
Be internationalised checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
lang-dir-and-logical-properties

Correct lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

DOM reports html lang=en and dir=ltr; CSS inspection found logical inline/block properties and the page links to 347 language editions.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

locale-aware-data

Dates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

This English reference landing page does not perform locale-sensitive number, currency, duration, or calendar calculations in the supplied journey.

time-zone-correctness

Time handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

The supplied journey exposes no scheduled event or time-zone-sensitive transaction.

Be trustworthy · 4 tests
Be trustworthy checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-dark-patterns

No deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive).

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Screenshots show no consent wall, disguised ad, confirmshaming, forced continuity, or hidden commitment; controls and links use direct labels.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

humane-error-handling

Forms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

The bounded non-mutating journey does not submit a form or create an invalid-input state.

trustworthy-input-assistance

Input is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

The only relevant visible input is encyclopedia search, for which address, payment, sign-in, and sign-up autocomplete tokens do not apply.

safe-commercial-and-account-flows

Checkout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

No checkout, subscription, consent, authentication, or account-management flow appears in the supplied anonymous journey.

Be sustainable · 3 tests
Be sustainable checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
optimised-assets

Images and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Method used or attempted: Permit-bound visual, DOM/evaluate, network, layout, trace, or journey evidence as cited.

Image audit found 16 below-fold images without loading=lazy, 17 legacy-format images, 3 oversized images, and one missing dimension; HAR lists 10 images without cache headers, including a 129 KB portrait.

Reference: image-summary, layout-summary, network-summary; retained-private

Below-fold imagery is eagerly loaded and several image delivery details waste bytes or weaken stability.

  • site-06-finding-07 · medium · Below-fold imagery is eagerly loaded and several image delivery details waste bytes or weaken stability.
no-wasteful-work

Background work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

After the bounded scroll, no network requests occurred; trace and journey metrics show no ongoing animation or increasing task work.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

third-party-and-media-budget

Third-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

HAR totals 682 KB, has no audio/video/autoplay media and no known trackers; third-party bytes are 253 KB of Wikimedia assets/auth.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Be agent ready · 2 tests
Be agent ready checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
structured-agent-capabilities

Where it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

No developer intent to expose action tools to agents was declared; high server-rendered content coverage already supports read-only agents without making emerging WebMCP mandatory.

on-device-inference

On-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses.

Method used or attempted: Applicability judgement against the supplied exact-URL non-mutating journey.

No applicable surface exists in the audited path.

Reference: No artifact reference; described-only

The static encyclopedia landing page has no summarisation or generation interaction for which built-in on-device inference is necessary.

Be memory-efficient · 3 tests
Be memory-efficient checks for https://en.wikipedia.org
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-leak-under-repeated-interaction

Repeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

After ten scroll-down/up repetitions, heap self size fell from 13,165,795 to 12,918,423 bytes and heap node count fell from 190,046 to 167,942.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

bounded-footprint

Heap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Baseline heap was [host/path omitted] MB self size; journey JSHeapUsedSize was about [host/path omitted] MB for a content-rich 2,375-element page.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

no-detached-dom-or-unbounded-listeners

No growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses.

Method used or attempted: Permit-bound screenshot, journey, DOM/evaluate, layout, trace, HAR, discoverability, or heap evidence.

Heap constructor summaries show no growing Detached* population; flow listeners stayed at 142 and DOM node count at 2,375 before/after scroll.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

7. https://www.nytimes.com · 58 slots · available
Respect user preferences · 3 tests
Respect user preferences checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
respects-color-scheme

Honours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method.

Method used or attempted: multi-modal evidence review

Dark emulation retained a white surface.

Reference: page-probe-summary, screenshot; retained-private

The homepage stays light when the user requests dark mode.

  • site-07-finding-01 · medium · The homepage stays light when the user requests dark mode.
respects-reduced-motion

Honours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses.

Method used or attempted: permit-bound evidence review

The reduced-motion query matched; only three finished 300ms animations remained and the CSS capture contains reduced-motion handling.

Reference: page-probe-summary, video; retained-private

The retained evidence directly supported this check under the captured conditions.

respects-contrast

Honours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Method used or attempted: permit-bound evidence review

Forced-colors/high-contrast capture preserved visible text, controls, borders, and links.

Reference: other-private-evidence; retained-private

The retained evidence directly supported this check under the captured conditions.

Implement natural interactions · 3 tests
Implement natural interactions checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
view-transitions

State and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps.

Blocked
Source: blocked; confidence: low
Implementation and method

Catalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

Method used or attempted: strict journey replay review

The strict supplied journey aborted before any state or route transition could be exercised.

Reference: journey-summary; retained-private

A blocked POST caused the journey to stop before a transition state.

scroll-driven-animations

Scroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

Method used or attempted: permit-bound evidence review

No scroll-linked animation or scrollytelling surface is exposed in the exact homepage state.

Reference: page-probe-summary; retained-private

The audited state has no scroll-linked visual effect requiring a declarative timeline.

physical-gestures

Gesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

Method used or attempted: permit-bound evidence review

No gesture-driven removal, pull, swipe, or comparable physical interaction is exposed on the exact homepage state.

Reference: screenshot; retained-private

The audited homepage state has no representative gesture surface.

Provide guided navigation · 3 tests
Provide guided navigation checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
scroll-state-aware-chrome

Sticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Method used or attempted: strict journey replay review

The supplied bounded-scroll state was skipped, so responsive scroll chrome could not be observed in that reviewed path.

Reference: journey-summary; retained-private

The strict journey stopped before its bounded-scroll action.

anchored-positioning

Tooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned.

Blocked
Source: blocked; confidence: low
Implementation and method

Catalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Method used or attempted: permit-bound evidence review

No tooltip or menu could be safely opened in the supplied strict journey after it aborted.

Reference: journey-summary; retained-private

The strict journey stopped before overlay interactions.

directs-attention

Navigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved.

Blocked
Source: blocked; confidence: low
Implementation and method

Catalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

Method used or attempted: permit-bound evidence review

No in-page jump or disclosure state was reached after the journey aborted.

Reference: journey-summary; retained-private

The strict journey stopped before a focus-moving navigation action.

Maximize content, reduce noise · 3 tests
Maximize content, reduce noise checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-intrusive-interruptions

No intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

Method used or attempted: multi-modal evidence review

The consent dialog obscures nearly all primary content.

Reference: screenshot; retained-private

A privacy alert dialog obscures the news on every fresh load.

  • site-07-finding-02 · high · A privacy alert dialog obscures the news on every fresh load.
semantic-dismissible-primitives

Overlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

Method used or attempted: permit-bound evidence review

The privacy surface uses a real dialog plus role=alertdialog and aria-modal=true, with explicit actions.

Reference: page-probe-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

reduced-chrome

Minimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

Method used or attempted: permit-bound evidence review

Below the modal, the publication layout is content-dense and uses restrained separators rather than a heavy app frame.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

Adapt to the form factor · 3 tests
Adapt to the form factor checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
responsive-no-horizontal-scroll

Layout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

Method used or attempted: permit-bound evidence review

Desktop and 360px layouts reported zero horizontal overflow and a valid viewport meta tag.

Reference: layout-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

component-level-responsiveness

Components adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

Method used or attempted: multi-modal evidence review

Captured CSS contains no container-query declarations.

Reference: page-probe-summary; retained-private

No container-query implementation was found in the captured homepage CSS.

  • site-07-finding-03 · medium · No container-query implementation was found in the captured homepage CSS.
input-modality-aware

Touch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

Method used or attempted: multi-modal evidence review

Visible focus and target-size probes found undersized and focus-invisible controls.

Reference: page-probe-summary; retained-private

Many visible controls are undersized and lose a visible keyboard focus indicator.

  • site-07-finding-04 · high · Many visible controls are undersized and lose a visible keyboard focus indicator.
Support core task success · 3 tests
Support core task success checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
clear-purpose-and-primary-action

The page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses.

Method used or attempted: multi-modal evidence review

Consent management displaces the news purpose in the first viewport.

Reference: screenshot; retained-private

The first viewport prioritises consent management over the publication and its current news.

  • site-07-finding-05 · high · The first viewport prioritises consent management over the publication and its current news.
primary-flow-completion

The representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses.

Method used or attempted: strict journey replay review

The supplied baseline action encountered a blocked POST and the bounded scroll was skipped, so the reviewed primary journey did not complete.

Reference: journey-summary; retained-private

The strict journey result is partial and action 2 is skipped-after-blocker.

clear-system-state-and-recovery

Loading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses.

Method used or attempted: permit-bound evidence review

The privacy state is explicit and offers Accept all, Reject all, and Manage preferences without hiding the rejection path.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

Be fast and stable · 5 tests
Be fast and stable checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
good-core-web-vitals

Core Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses.

Method used or attempted: multi-modal evidence review

LCP was [host/path omitted] and TBT 994ms.

Reference: performance-summary; retained-private

Observed load performance misses the good LCP range and has substantial blocking work.

  • site-07-finding-06 · high · Observed load performance misses the good LCP range and has substantial blocking work.
visual-stability

No cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Method used or attempted: multi-modal evidence review

Mobile CLS was [host/path omitted] and most images omitted dimensions.

Reference: image-summary, layout-summary; retained-private

The mobile homepage has poor observed layout stability.

  • site-07-finding-07 · high · The mobile homepage has poor observed layout stability.
efficient-main-thread

The main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

Method used or attempted: multi-modal evidence review

Trace captured 13 long tasks and 994ms TBT.

Reference: performance-summary; retained-private

Heavy scripting blocks the main thread during load.

  • site-07-finding-08 · high · Heavy scripting blocks the main thread during load.
efficient-resource-delivery

Critical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses.

Method used or attempted: multi-modal evidence review

HAR captured 368 requests and [host/path omitted].

Reference: network-summary; retained-private

The homepage load is exceptionally request-heavy for a content entry page.

  • site-07-finding-09 · high · The homepage load is exceptionally request-heavy for a content entry page.
trim-unused-and-duplicate-code

The page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses.

Method used or attempted: multi-modal evidence review

259 script requests transferred [host/path omitted].

Reference: network-summary; retained-private

JavaScript and third-party code dominate transfer cost.

  • site-07-finding-10 · high · JavaScript and third-party code dominate transfer cost.
Be inclusive · 5 tests
Be inclusive checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
names-roles-labels

Interactive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses.

Method used or attempted: multi-modal evidence review

Empty interactive anchors and null image alternatives were observed.

Reference: image-summary, page-probe-summary; retained-private

The homepage includes interactive links with no accessible text and editorial images with absent descriptions.

  • site-07-finding-11 · high · The homepage includes interactive links with no accessible text and editorial images with absent descriptions.
sufficient-contrast

Text and essential UI meet WCAG colour-contrast minimums against their background.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Method used or attempted: permit-bound evidence review

The default consent surface is high contrast and forced-colors preserves all essential text and controls.

Reference: other-private-evidence, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

structure-and-focus

Heading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Method used or attempted: multi-modal evidence review

Many focused interactive elements expose transparent or none outlines.

Reference: page-probe-summary; retained-private

Keyboard focus is not consistently visible.

  • site-07-finding-12 · high · Keyboard focus is not consistently visible.
legible-text

Text is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Method used or attempted: permit-bound evidence review

Desktop and mobile captures show readable type, comfortable line lengths, and no clipping in the privacy content.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

zoom-reflow-targets-and-media

The experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses.

Method used or attempted: multi-modal evidence review

The mobile layout reflows, but several controls are 12px to 32px high and carousel buttons are 24px square.

Reference: layout-summary, page-probe-summary; retained-private

Several primary and repeated controls miss an inclusive target-size floor.

  • site-07-finding-13 · high · Several primary and repeated controls miss an inclusive target-size floor.
Follow best practices · 3 tests
Follow best practices checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-console-errors

The page loads without console errors or uncaught exceptions.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

Method used or attempted: strict journey capture review

The supplied journey explicitly reports the console collector as blocked, and no pre-navigation console capture was available.

Reference: journey-summary; retained-private

Required Stage 2 console collector was unavailable.

sound-document-and-assets

Valid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.)

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Method used or attempted: multi-modal evidence review

Document basics pass, but 20 of 25 images omit intrinsic dimensions.

Reference: image-summary, page-probe-summary; retained-private

Most rendered images omit intrinsic dimensions.

  • site-07-finding-14 · medium · Most rendered images omit intrinsic dimensions.
browser-platform-hygiene

The page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

Method used or attempted: multi-modal evidence review

Three requests for the same VTT resource returned 404.

Reference: network-summary; retained-private

The load includes repeated failed media-caption requests.

  • site-07-finding-15 · medium · The load includes repeated failed media-caption requests.
Be discoverable · 4 tests
Be discoverable checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
title-and-description

The page has a unique, descriptive <title> and a meta description.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

Method used or attempted: permit-bound evidence review

A descriptive title and meta description are present in raw and rendered HTML.

Reference: discoverability-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

crawlable-and-mobile-friendly

Links are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

Method used or attempted: permit-bound evidence review

The page has a viewport meta tag, 95% no-JS content coverage, and 392 of 412 anchors expose hrefs.

Reference: discoverability-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

canonical-and-indexing-signals

Public pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses.

Method used or attempted: permit-bound evidence review

HTTP status is 200; canonical, nine hreflang links, and no accidental noindex were observed.

Reference: discoverability-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

structured-and-shareable-metadata

Where the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses.

Method used or attempted: permit-bound evidence review

The rendered homepage exposes two JSON-LD blocks and five Open Graph tags.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Be private and secure · 4 tests
Be private and secure checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
secure-transport-and-headers

Served over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses.

Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence.

Retained evidence records 5 of 6 baseline security-header categories present. Cookie-attribute review records 9 of 14 records with Secure and 3 with HttpOnly. Names, values, routes, and raw headers remain private.

Reference: cookie-attribute-summary, security-header-summary; retained-private

Transport is strong, but CSP and cookie posture remain permissive.

  • site-07-finding-16 · high · Transport is strong, but CSP and cookie posture remain permissive.
data-minimisation-and-third-parties

No over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses.

Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence.

Retained evidence records 1 first-party origin categories and 27 third-party origin categories across 368 requests. Destinations, identifiers, routes, headers, and bodies remain private.

Reference: network-summary, screenshot, tracker-summary; retained-private

Large advertising and measurement traffic begins before a privacy choice is made.

  • site-07-finding-17 · high · Large advertising and measurement traffic begins before a privacy choice is made.
in-context-permissions-and-modern-auth

Permission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses.

Method used or attempted: permit-bound evidence review

No browser permission prompt appeared on load; an authentication form was not part of the exact permitted page state.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

defensive-browser-policies

Browser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses.

Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence.

Retained evidence records 5 of 6 baseline browser-policy header categories present. Values and raw headers remain private.

Reference: security-header-summary; retained-private

Defensive policy coverage is incomplete.

  • site-07-finding-18 · medium · Defensive policy coverage is incomplete.
Be resilient · 4 tests
Be resilient checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
progressive-enhancement

Core content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Method used or attempted: permit-bound evidence review

Raw HTML returned 200 with 95% rendered-content coverage, including title, H1, and description.

Reference: discoverability-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

resilient-runtime-behaviour

The page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

Method used or attempted: permit-bound evidence review

The privacy dialog remains contained and usable at desktop and 360px without overflow.

Reference: layout-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

offline-and-installable

Where the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.)

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

Method used or attempted: permit-bound evidence review

The audited surface is a public news homepage, not an app-like offline workflow; progressive enhancement is judged separately and passes.

Reference: discoverability-summary; retained-private

Offline installability is not required for this publication homepage archetype.

network-and-http-failure-states

HTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states.

Not run
Source: not-run; confidence: high
Implementation and method

Catalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses.

Method used or attempted: permit-bound evidence review

No alternate URL or failure route was permitted, and injecting a network failure was outside the supplied bounded journey.

Reference: journey-summary; retained-private

The exact-URL and strict-journey boundary did not permit a representative error-state path.

Be internationalised · 3 tests
Be internationalised checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
lang-dir-and-logical-properties

Correct lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

Method used or attempted: permit-bound evidence review

The page declares lang=en, offers language choices, exposes nine hreflang links, and captured CSS uses inline logical properties.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

locale-aware-data

Dates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

Method used or attempted: permit-bound evidence review

The UK-observed page localises subscription currency to GBP and provides language selection.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

time-zone-correctness

Time handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

Method used or attempted: permit-bound evidence review

No event scheduling or time-zone-sensitive task is present in the exact homepage state.

Reference: screenshot; retained-private

The audited homepage state exposes no time-zone-sensitive flow.

Be trustworthy · 4 tests
Be trustworthy checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-dark-patterns

No deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive).

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

Method used or attempted: permit-bound evidence review

Accept all and Reject all have equal visual prominence, Manage preferences is explicit, and policy links are visible.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

humane-error-handling

Forms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

Method used or attempted: permit-bound evidence review

No user-submitted form or invalid-input state is present in the exact homepage state.

Reference: page-probe-summary; retained-private

The only visible choices are consent actions, not data-entry validation.

trustworthy-input-assistance

Input is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

Method used or attempted: permit-bound evidence review

No sign-in, address, payment, or comparable text input is present in the audited state.

Reference: page-probe-summary; retained-private

No user text-entry flow is present on the exact homepage state.

safe-commercial-and-account-flows

Checkout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity.

Not run
Source: not-run; confidence: high
Implementation and method

Catalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses.

Method used or attempted: permit-bound evidence review

A subscription link is present, but cross-origin or different-path commercial navigation was expressly outside the exact-URL boundary.

Reference: screenshot; retained-private

The commercial flow could not be opened without leaving the only permitted URL.

Be sustainable · 3 tests
Be sustainable checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
optimised-assets

Images and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Method used or attempted: multi-modal evidence review

All 25 observed images use legacy JPG or PNG formats.

Reference: image-summary; retained-private

Image delivery relies entirely on legacy formats in the observed state.

  • site-07-finding-19 · medium · Image delivery relies entirely on legacy formats in the observed state.
no-wasteful-work

Background work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Method used or attempted: multi-modal evidence review

The initial state transfers [host/path omitted] and executes 259 script requests before content access.

Reference: network-summary, performance-summary; retained-private

The initial page performs disproportionate background and third-party work.

  • site-07-finding-20 · high · The initial page performs disproportionate background and third-party work.
third-party-and-media-budget

Third-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses.

Method used or attempted: multi-modal evidence review

Third parties account for [host/path omitted] and 114 requests.

Reference: network-summary, tracker-summary; retained-private

Third-party scripts consume more than half of transferred bytes.

  • site-07-finding-21 · high · Third-party scripts consume more than half of transferred bytes.
Be agent ready · 2 tests
Be agent ready checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
structured-agent-capabilities

Where it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

Method used or attempted: permit-bound evidence review

The publication exposes crawlable structured content but no declared agent-action surface; emerging WebMCP capability is treated as an opportunity, not a failure.

Reference: discoverability-summary, page-probe-summary; retained-private

No agent-facing transactional capability is declared for this public homepage.

on-device-inference

On-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses.

Method used or attempted: permit-bound evidence review

No on-device inference feature is part of the homepage purpose.

Reference: page-probe-summary; retained-private

Built-in AI is not necessary for the audited homepage task.

Be memory-efficient · 3 tests
Be memory-efficient checks for https://www.nytimes.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-leak-under-repeated-interaction

Repeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses.

Method used or attempted: permit-bound evidence review

After ten bounded down/up scroll cycles, heap self-size fell from [host/path omitted] to [host/path omitted] and node count fell from [host/path omitted] to [host/path omitted].

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

bounded-footprint

Heap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses.

Method used or attempted: multi-modal evidence review

Both snapshots retain roughly 116MB and [host/path omitted] heap nodes.

Reference: memory-summary; retained-private

The baseline memory footprint is large for a news homepage.

  • site-07-finding-22 · medium · The baseline memory footprint is large for a news homepage.
no-detached-dom-or-unbounded-listeners

No growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses.

Method used or attempted: permit-bound evidence review

The post-scroll snapshot did not grow: nodes, closures, objects, arrays, and total self-size all decreased; Detached constructors were not among the retained leaders.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

8. https://gemini.google.com · 58 slots · available
Respect user preferences · 3 tests
Respect user preferences checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
respects-color-scheme

Honours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

respects-reduced-motion

Honours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The permit-bound reduced-motion probe reported prefers-reduced-motion: reduce=true but 9 running animations, including two with 1,498,500 ms durations.

Reference: other-private-evidence; retained-private

Long-running motion remains active when reduced motion is requested.

  • site-08-finding-01 · high · Long-running motion remains active when reduced motion is requested.
respects-contrast

Honours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

Implement natural interactions · 3 tests
Implement natural interactions checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
view-transitions

State and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

scroll-driven-animations

Scroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

physical-gestures

Gesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

Provide guided navigation · 3 tests
Provide guided navigation checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
scroll-state-aware-chrome

Sticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

anchored-positioning

Tooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

directs-attention

Navigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

Maximize content, reduce noise · 3 tests
Maximize content, reduce noise checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-intrusive-interruptions

No intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

Desktop, mobile, and pre-replayed journey screenshots show a modal covering the prompt and main content. On 360x800, the dialog fills most of the viewport and initially exposes only Read more while the decisive controls are below the fold.

Reference: screenshot; retained-private

The consent dialog obscures the primary assistant on first load.

  • site-08-finding-02 · high · The consent dialog obscures the primary assistant on first load.
semantic-dismissible-primitives

Overlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: page-probe-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

reduced-chrome

Minimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: page-probe-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

Adapt to the form factor · 3 tests
Adapt to the form factor checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
responsive-no-horizontal-scroll

Layout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: layout-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

component-level-responsiveness

Components adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

input-modality-aware

Touch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: journey-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Support core task success · 3 tests
Support core task success checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
clear-purpose-and-primary-action

The page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: page-probe-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

primary-flow-completion

The representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

clear-system-state-and-recovery

Loading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

Be fast and stable · 5 tests
Be fast and stable checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
good-core-web-vitals

Core Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The permit-bound trace measured FCP 6,623 ms and LCP 6,862 ms, well outside the good LCP range; it also recorded 108 ms total blocking time.

Reference: performance-summary; retained-private

Cold-load rendering is too slow.

  • site-08-finding-03 · high · Cold-load rendering is too slow.
visual-stability

No cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: layout-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

efficient-main-thread

The main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The permit-bound trace measured FCP 6,623 ms and LCP 6,862 ms, well outside the good LCP range; it also recorded 108 ms total blocking time.

Reference: performance-summary; retained-private

Cold-load rendering is too slow.

  • site-08-finding-03 · high · Cold-load rendering is too slow.
efficient-resource-delivery

Critical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The HAR summary recorded 100 requests and 5,424,867 transferred bytes: [host/path omitted] MB of scripts and [host/path omitted] MB of fonts. Two parser-inserted high-priority scripts were render-blocking candidates.

Reference: network-summary, other-private-evidence; retained-private

The initial app shell transfers a disproportionate payload.

  • site-08-finding-04 · high · The initial app shell transfers a disproportionate payload.
trim-unused-and-duplicate-code

The page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The HAR captured 51 script requests totalling 3,787,375 transferred bytes before any assistant task was performed, including individual bundles over 1 MB and 700 KB.

Reference: other-private-evidence; retained-private

The signed-out shell eagerly loads a large script surface.

  • site-08-finding-05 · medium · The signed-out shell eagerly loads a large script surface.
Be inclusive · 5 tests
Be inclusive checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
names-roles-labels

Interactive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The images primitive found all 5 images without width/height attributes and 4 without alt attributes; the narrow layout still measured CLS [host/path omitted].

Reference: image-summary, layout-summary; retained-private

Rendered images omit intrinsic dimensions and several omit alt text.

  • site-08-finding-06 · medium · Rendered images omit intrinsic dimensions and several omit alt text.
sufficient-contrast

Text and essential UI meet WCAG colour-contrast minimums against their background.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

structure-and-focus

Heading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: page-probe-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

legible-text

Text is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

zoom-reflow-targets-and-media

The experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

Follow best practices · 3 tests
Follow best practices checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-console-errors

The page loads without console errors or uncaught exceptions.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

sound-document-and-assets

Valid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.)

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The images primitive found all 5 images without width/height attributes and 4 without alt attributes; the narrow layout still measured CLS [host/path omitted].

Reference: image-summary, layout-summary; retained-private

Rendered images omit intrinsic dimensions and several omit alt text.

  • site-08-finding-06 · medium · Rendered images omit intrinsic dimensions and several omit alt text.
browser-platform-hygiene

The page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

Be discoverable · 4 tests
Be discoverable checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
title-and-description

The page has a unique, descriptive <title> and a meta description.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

crawlable-and-mobile-friendly

Links are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

Discoverability evidence measured only 2% rendered-word coverage in raw HTML and classified the page as isJsShell=true; the crawler screenshot contains no useful assistant content.

Reference: discoverability-summary; retained-private

The public entry page is effectively a JavaScript shell for non-JS crawlers.

  • site-08-finding-07 · high · The public entry page is effectively a JavaScript shell for non-JS crawlers.
canonical-and-indexing-signals

Public pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

structured-and-shareable-metadata

Where the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content.

Not applicable
Source: not-applicable; confidence: medium
Implementation and method

Catalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses.

Method used or attempted: Applicability judgement from the signed-out Gemini app archetype

The audited entry state does not present this contextual surface.

Reference: No artifact reference; described-only

This contextual check is not applicable to the exact signed-out entry state and bounded non-mutating journey.

Be private and secure · 4 tests
Be private and secure checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
secure-transport-and-headers

Served over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses.

Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence.

Retained evidence records 4 of 6 baseline security-header categories present. Cookie-attribute review records 1 of 1 records with Secure and 1 with HttpOnly. Names, values, routes, and raw headers remain private.

Reference: security-header-summary; retained-private

The document security policy leaves avoidable browser-defense gaps.

  • site-08-finding-08 · medium · The document security policy leaves avoidable browser-defense gaps.
data-minimisation-and-third-parties

No over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses.

Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence.

Retained evidence records 1 first-party origin categories and 8 third-party origin categories across 100 requests. Destinations, identifiers, routes, headers, and bodies remain private.

Reference: other-private-evidence, screenshot, tracker-summary; retained-private

The signed-out baseline contacts a broad third-party-origin surface.

  • site-08-finding-09 · medium · The signed-out baseline contacts a broad third-party-origin surface.
in-context-permissions-and-modern-auth

Permission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

defensive-browser-policies

Browser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses.

Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence.

Retained evidence records 4 of 6 baseline browser-policy header categories present. Values and raw headers remain private.

Reference: security-header-summary; retained-private

The document security policy leaves avoidable browser-defense gaps.

  • site-08-finding-08 · medium · The document security policy leaves avoidable browser-defense gaps.
Be resilient · 4 tests
Be resilient checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
progressive-enhancement

Core content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The raw document returned HTTP 200 but exposed only 2% of rendered content words without JavaScript and was classified as a JS shell.

Reference: discoverability-summary; retained-private

Core public content does not progressively enhance without JavaScript.

  • site-08-finding-10 · high · Core public content does not progressively enhance without JavaScript.
resilient-runtime-behaviour

The page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

offline-and-installable

Where the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.)

Not applicable
Source: not-applicable; confidence: medium
Implementation and method

Catalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

Method used or attempted: Applicability judgement from the signed-out Gemini app archetype

The audited entry state does not present this contextual surface.

Reference: No artifact reference; described-only

This contextual check is not applicable to the exact signed-out entry state and bounded non-mutating journey.

network-and-http-failure-states

HTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

Be internationalised · 3 tests
Be internationalised checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
lang-dir-and-logical-properties

Correct lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

locale-aware-data

Dates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly.

Not applicable
Source: not-applicable; confidence: medium
Implementation and method

Catalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

Method used or attempted: Applicability judgement from the signed-out Gemini app archetype

The audited entry state does not present this contextual surface.

Reference: No artifact reference; described-only

This contextual check is not applicable to the exact signed-out entry state and bounded non-mutating journey.

time-zone-correctness

Time handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous.

Not applicable
Source: not-applicable; confidence: medium
Implementation and method

Catalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

Method used or attempted: Applicability judgement from the signed-out Gemini app archetype

The audited entry state does not present this contextual surface.

Reference: No artifact reference; described-only

This contextual check is not applicable to the exact signed-out entry state and bounded non-mutating journey.

Be trustworthy · 4 tests
Be trustworthy checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-dark-patterns

No deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive).

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

Method used or attempted: Permit-bound screenshot, DOM, layout, or evaluate probe

Direct evidence on the signed-out entry state supported this outcome under the audited desktop/mobile condition.

Reference: page-probe-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

humane-error-handling

Forms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

trustworthy-input-assistance

Input is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

safe-commercial-and-account-flows

Checkout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

Be sustainable · 3 tests
Be sustainable checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
optimised-assets

Images and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The images primitive found all 5 images without width/height attributes and 4 without alt attributes; the narrow layout still measured CLS [host/path omitted].

Reference: image-summary, layout-summary; retained-private

Rendered images omit intrinsic dimensions and several omit alt text.

  • site-08-finding-06 · medium · Rendered images omit intrinsic dimensions and several omit alt text.
no-wasteful-work

Background work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

third-party-and-media-budget

Third-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The baseline transferred [host/path omitted] MB, including [host/path omitted] MB from non-entry origins, to render a consent-covered assistant shell; scripts and fonts dominate.

Reference: other-private-evidence, screenshot; retained-private

Initial resource use is disproportionate to the visible signed-out state.

  • site-08-finding-11 · medium · Initial resource use is disproportionate to the visible signed-out state.
Be agent ready · 2 tests
Be agent ready checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
structured-agent-capabilities

Where it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

The discoverability probe found only 2% raw-HTML word coverage and no useful crawler-visible task surface, leaving non-JS agents to infer the product from metadata.

Reference: discoverability-summary; retained-private

The AI assistant entry surface is not robustly machine-readable without executing the application.

  • site-08-finding-12 · medium · The AI assistant entry surface is not robustly machine-readable without executing the application.
on-device-inference

On-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server.

Not applicable
Source: not-applicable; confidence: medium
Implementation and method

Catalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses.

Method used or attempted: Applicability judgement from the signed-out Gemini app archetype

The audited entry state does not present this contextual surface.

Reference: No artifact reference; described-only

This contextual check is not applicable to the exact signed-out entry state and bounded non-mutating journey.

Be memory-efficient · 3 tests
Be memory-efficient checks for https://gemini.google.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-leak-under-repeated-interaction

Repeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

bounded-footprint

Heap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses.

Method used or attempted: Permit-bound raw-CDP evidence and model judgement

A single permit-bound heap summary contained 763,142 nodes, 3,168,144 edges, and 36,441,757 self-size bytes before any assistant interaction.

Reference: memory-summary; retained-private

The signed-out, consent-covered shell has a large baseline heap footprint.

  • site-08-finding-13 · medium · The signed-out, consent-covered shell has a large baseline heap footprint.
no-detached-dom-or-unbounded-listeners

No growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses.

Method used or attempted: Attempted strict V1 journey or required a prohibited/unavailable interaction state

The pre-replayed strict journey aborted after mutation-method-OPTIONS on baseline load; the bounded scroll was skipped, and login, prompt submission, account, permission, error, and repeated-interaction states are prohibited by the supplied boundary.

Reference: journey-summary; retained-private

Exact execution policy and the partial journey prevented direct evidence for this check.

9. https://www.netflix.com · 58 slots · available
Respect user preferences · 3 tests
Respect user preferences checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
respects-color-scheme

Honours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Dark and default screenshots are byte-identical, while computed color-scheme is normal. The dark visual design is hard-coded rather than preference-driven.

Reference: page-probe-summary, screenshot; retained-private

The page does not react to the user color-scheme preference

  • site-09-finding-01 · low · The page does not react to the user color-scheme preference
respects-reduced-motion

Honours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Reduced-motion evaluate probe reports zero animations, while the default probe reports ten running 4,000ms animations.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

respects-contrast

Honours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Forced-colors/prefers-contrast screenshot keeps text, form borders, buttons, and dismiss control visible.

Reference: other-private-evidence; retained-private

The retained evidence directly supported this check under the captured conditions.

Implement natural interactions · 3 tests
Implement natural interactions checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
view-transitions

State and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The retained DOM/CSS contains no view-transition declarations despite interactive FAQ, carousel, signup, and navigation surfaces.

Reference: page-probe-summary; retained-private

State and route changes do not expose View Transition support

  • site-09-finding-02 · low · State and route changes do not expose View Transition support
scroll-driven-animations

Scroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

Method used or attempted: Contextual applicability judgement against the audited landing state.

No parallax, scrollytelling, or entry/exit reveal motion was observed; the horizontal carousel uses native scroll snap without scroll-linked animation.

Reference: No artifact reference; described-only

No parallax, scrollytelling, or entry/exit reveal motion was observed; the horizontal carousel uses native scroll snap without scroll-linked animation.

physical-gestures

Gesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Retained CSS uses mandatory horizontal scroll-snap and logical scroll margins for the trending carousel.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Provide guided navigation · 3 tests
Provide guided navigation checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
scroll-state-aware-chrome

Sticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The document is 3,266px tall at desktop and 3,970px at mobile, but the retained CSS has no animation-timeline or scroll-state query markers and the planned scroll state was skipped.

Reference: journey-summary, layout-summary; retained-private

The long landing page has no scroll-state-aware guidance

  • site-09-finding-03 · low · The long landing page has no scroll-state-aware guidance
anchored-positioning

Tooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Method used or attempted: Contextual applicability judgement against the audited landing state.

No tooltip, popover, or menu overlay requiring anchored positioning exists in the audited landing state.

Reference: No artifact reference; described-only

No tooltip, popover, or menu overlay requiring anchored positioning exists in the audited landing state.

directs-attention

Navigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The carousel uses scroll-snap alignment and the visual hierarchy clearly highlights the current hero action.

Reference: other-private-evidence, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Maximize content, reduce noise · 3 tests
Maximize content, reduce noise checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-intrusive-interruptions

No intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Desktop and mobile screenshots show the cookie banner over the hero; on 360x800 it occupies roughly the top quarter of the viewport before the core task.

Reference: other-private-evidence, screenshot; retained-private

The cookie banner obscures primary content on load

  • site-09-finding-04 · medium · The cookie banner obscures primary content on load
semantic-dismissible-primitives

Overlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The live DOM reports zero dialog, popover, and details elements while the cookie preference overlay and six FAQ disclosures are present.

Reference: other-private-evidence, page-probe-summary; retained-private

Dismissible UI is implemented without native overlay primitives

  • site-09-finding-05 · low · Dismissible UI is implemented without native overlay primitives
reduced-chrome

Minimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The first viewport is dominated by content and the core signup action, with minimal persistent chrome.

Reference: other-private-evidence; retained-private

The retained evidence directly supported this check under the captured conditions.

Adapt to the form factor · 3 tests
Adapt to the form factor checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
responsive-no-horizontal-scroll

Layout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Layout evidence reports zero horizontal overflow at both 780x493 and 360x800, with viewport meta present.

Reference: layout-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

component-level-responsiveness

Components adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The retained authored CSS contains no @container rule and computed container-type is normal on inspected components.

Reference: page-probe-summary; retained-private

Components rely on viewport styling rather than container queries

  • site-09-finding-06 · low · Components rely on viewport styling rather than container queries
input-modality-aware

Touch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The focused Sign In link has a 2px solid outline; primary input and button are 56px tall and work without hover.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Support core task success · 3 tests
Support core task success checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
clear-purpose-and-primary-action

The page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The headline, £[host/path omitted] starting price, cancellation promise, email field, and Get Started action are prominent in the first viewport.

Reference: other-private-evidence; retained-private

The retained evidence directly supported this check under the captured conditions.

primary-flow-completion

The representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses.

Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy.

Signup requires entering and submitting an email, a hard-denied mutating journey action. It was not executed.

Reference: No artifact reference; described-only

Signup requires entering and submitting an email, a hard-denied mutating journey action. It was not executed.

clear-system-state-and-recovery

Loading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses.

Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy.

Invalid submission and network-failure states require prohibited mutation/failure injection and were not executed.

Reference: No artifact reference; described-only

Invalid submission and network-failure states require prohibited mutation/failure injection and were not executed.

Be fast and stable · 5 tests
Be fast and stable checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
good-core-web-vitals

Core Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Trace reports LCP 2,[host/path omitted] and TBT [host/path omitted]; layout reports CLS [host/path omitted] desktop and 0 mobile. No field INP was claimed.

Reference: layout-summary, performance-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

visual-stability

No cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Observed CLS is [host/path omitted] desktop and 0 mobile, both in the good range.

Reference: layout-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

efficient-main-thread

The main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Trace reports two long tasks, longest [host/path omitted], with [host/path omitted] total blocking time.

Reference: performance-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

efficient-resource-delivery

Critical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

HAR summary reports 53 requests and [host/path omitted] MB transferred, including [host/path omitted] MB scripts and 681 KB fonts. A 753 KB framework script and parser-inserted OneTrust script are high-priority candidates.

Reference: network-summary, performance-summary; retained-private

Initial delivery is heavier than the acquisition task requires

  • site-09-finding-07 · medium · Initial delivery is heavier than the acquisition task requires
trim-unused-and-duplicate-code

The page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The HAR attributes [host/path omitted] MB of [host/path omitted] MB to non-main origins, including 341 KB reCAPTCHA and 352 KB OneTrust resources before signup interaction.

Reference: network-summary; retained-private

Third-party and support code dominate initial transfer

  • site-09-finding-08 · medium · Third-party and support code dominate initial transfer
Be inclusive · 5 tests
Be inclusive checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
names-roles-labels

Interactive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Both email inputs have programmatic labels and autocomplete=email; visible controls have names. Missing alt values apply to decorative hero/logo images.

Reference: image-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

sufficient-contrast

Text and essential UI meet WCAG colour-contrast minimums against their background.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Default, mobile, and forced-colors screenshots keep essential text and controls visibly distinct from their backgrounds.

Reference: other-private-evidence, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

structure-and-focus

Heading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The probe finds an H1 followed directly by H3 and no main or nav landmark, although keyboard focus is visibly outlined.

Reference: page-probe-summary; retained-private

The document hierarchy and landmarks are incomplete

  • site-09-finding-09 · medium · The document hierarchy and landmarks are incomplete
legible-text

Text is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Desktop and mobile screenshots show readable type, sensible wrapping, and no clipped primary copy.

Reference: other-private-evidence, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

zoom-reflow-targets-and-media

The experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The viewport meta is width=device-width, initial-scale=[host/path omitted], minimum-scale=[host/path omitted], maximum-scale=[host/path omitted]. The maximum-scale restriction prevents user scaling in affected browsers.

Reference: page-probe-summary, screenshot; retained-private

The viewport metadata disables pinch zoom

  • site-09-finding-10 · high · The viewport metadata disables pinch zoom
Follow best practices · 3 tests
Follow best practices checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-console-errors

The page loads without console errors or uncaught exceptions.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy.

The imported V1 journey explicitly records console capture as blocked, and no permit-bound console collector primitive was available.

Reference: No artifact reference; described-only

The imported V1 journey explicitly records console capture as blocked, and no permit-bound console collector primitive was available.

sound-document-and-assets

Valid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.)

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The image audit reports 4 of 4 images without width/height and 2 oversized images; the logo natural width is 370px for an 89px display width.

Reference: image-summary; retained-private

Images omit intrinsic dimensions and some are oversized

  • site-09-finding-11 · medium · Images omit intrinsic dimensions and some are oversized
browser-platform-hygiene

The page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy.

BFCache, notification prompt, vulnerable library, and paste-prevention coverage was not available from the retained bounded paths.

Reference: No artifact reference; described-only

BFCache, notification prompt, vulnerable library, and paste-prevention coverage was not available from the retained bounded paths.

Be discoverable · 4 tests
Be discoverable checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
title-and-description

The page has a unique, descriptive <title> and a meta description.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The probe records a descriptive localized title and meta description.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

crawlable-and-mobile-friendly

Links are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

All 30 links have hrefs and non-empty text; viewport meta is present; raw HTML returns 200.

Reference: discoverability-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

canonical-and-indexing-signals

Public pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The page succeeds with HTTP 200 at [route omitted] but the DOM probe finds no canonical URL and source inspection finds no hreflang markers.

Reference: discoverability-summary, page-probe-summary; retained-private

The localized landing page has no canonical or hreflang signal

  • site-09-finding-12 · low · The localized landing page has no canonical or hreflang signal
structured-and-shareable-metadata

Where the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The DOM includes Open Graph and Twitter metadata, but contains no JSON-LD for the organization/service represented by the page.

Reference: page-probe-summary; retained-private

Share metadata exists but structured entity metadata is absent

  • site-09-finding-13 · low · Share metadata exists but structured entity metadata is absent
Be private and secure · 4 tests
Be private and secure checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
secure-transport-and-headers

Served over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses.

Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence.

Retained evidence records 3 of 6 baseline security-header categories present. Cookie-attribute review records 3 of 9 records with Secure and 3 with HttpOnly. Names, values, routes, and raw headers remain private.

Reference: cookie-attribute-summary, security-header-summary; retained-private

Security headers and cookie flags have material gaps

  • site-09-finding-14 · high · Security headers and cookie flags have material gaps
data-minimisation-and-third-parties

No over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses.

Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence.

Retained evidence records 1 first-party origin categories and 11 third-party origin categories across 53 requests. Destinations, identifiers, routes, headers, and bodies remain private.

Reference: network-summary, tracker-summary; retained-private

Extensive cross-origin code and telemetry load before task interaction

  • site-09-finding-15 · medium · Extensive cross-origin code and telemetry load before task interaction
in-context-permissions-and-modern-auth

Permission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses.

Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy.

Authentication was outside the permitted path; passkey support and prompt timing could not be exercised.

Reference: No artifact reference; described-only

Authentication was outside the permitted path; passkey support and prompt timing could not be exercised.

defensive-browser-policies

Browser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses.

Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence.

Retained evidence records 3 of 6 baseline browser-policy header categories present. Values and raw headers remain private.

Reference: security-header-summary; retained-private

Several browser-enforced policies are absent

  • site-09-finding-16 · medium · Several browser-enforced policies are absent
Be resilient · 4 tests
Be resilient checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
progressive-enhancement

Core content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Discoverability evidence shows 90% rendered-word coverage in raw HTML, with title, H1, and description present and no JS shell.

Reference: discoverability-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

resilient-runtime-behaviour

The page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The cookie banner and hero remain contained at desktop and mobile widths, and the carousel uses native scroll snap.

Reference: other-private-evidence, page-probe-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

offline-and-installable

Where the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.)

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

Method used or attempted: Contextual applicability judgement against the audited landing state.

The public acquisition page is not an installed app surface; offline signup cannot complete meaningfully.

Reference: No artifact reference; described-only

The public acquisition page is not an installed app surface; offline signup cannot complete meaningfully.

network-and-http-failure-states

HTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses.

Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy.

Network failure and error-route injection were outside the exact-URL bounded run.

Reference: No artifact reference; described-only

Network failure and error-route injection were outside the exact-URL bounded run.

Be internationalised · 3 tests
Be internationalised checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
lang-dir-and-logical-properties

Correct lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

html lang is en but dir is empty; retained CSS contains many physical left/right declarations, although some logical scroll-margin-inline is used.

Reference: page-probe-summary; retained-private

The localized page does not declare direction and still ships substantial physical-direction CSS

  • site-09-finding-17 · low · The localized page does not declare direction and still ships substantial physical-direction CSS
locale-aware-data

Dates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The GB page renders pound pricing and retained source includes Intl usage rather than only hand-written locale formatting.

Reference: other-private-evidence, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

time-zone-correctness

Time handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

Method used or attempted: Contextual applicability judgement against the audited landing state.

The audited landing page presents no dates, events, recurrence, or time-zone-sensitive data.

Reference: No artifact reference; described-only

The audited landing page presents no dates, events, recurrence, or time-zone-sensitive data.

Be trustworthy · 4 tests
Be trustworthy checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-dark-patterns

No deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive).

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Visible pricing and cancellation terms are adjacent to the CTA; Reject and Accept consent actions have equal prominence.

Reference: other-private-evidence, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

humane-error-handling

Forms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy.

Invalid form submission is a prohibited mutating action and was not executed.

Reference: No artifact reference; described-only

Invalid form submission is a prohibited mutating action and was not executed.

trustworthy-input-assistance

Input is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Both email fields use type=email, programmatic labels, and autocomplete=email.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

safe-commercial-and-account-flows

Checkout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses.

Method used or attempted: Attempted coverage planning under the strict exact-URL journey policy.

Subscription and account flows are hard-denied by the strict journey policy and were not executed.

Reference: No artifact reference; described-only

Subscription and account flows are hard-denied by the strict journey policy and were not executed.

Be sustainable · 3 tests
Be sustainable checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
optimised-assets

Images and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Image evidence reports 2 oversized images, 2 legacy-format images, and missing intrinsic dimensions on all four img elements.

Reference: image-summary; retained-private

Several first-view assets are not optimally sized or encoded

  • site-09-finding-18 · medium · Several first-view assets are not optimally sized or encoded
no-wasteful-work

Background work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

The initial load transfers a 341 KB reCAPTCHA script, 352 KB of OneTrust resources, and six logging requests before the signup action is used.

Reference: network-summary, tracker-summary; retained-private

Task-specific work is eagerly loaded before user intent

  • site-09-finding-19 · medium · Task-specific work is eagerly loaded before user intent
third-party-and-media-budget

Third-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

HAR summary reports [host/path omitted] MB cross-origin transfer, 681 KB fonts, and [host/path omitted] MB scripts for a static acquisition first view.

Reference: network-summary; retained-private

The initial third-party and font budget is disproportionate

  • site-09-finding-20 · medium · The initial third-party and font budget is disproportionate
Be agent ready · 2 tests
Be agent ready checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
structured-agent-capabilities

Where it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

Method used or attempted: Contextual applicability judgement against the audited landing state.

Agent capabilities are emerging and no agent-facing surface is intended for this consumer acquisition state.

Reference: No artifact reference; described-only

Agent capabilities are emerging and no agent-facing surface is intended for this consumer acquisition state.

on-device-inference

On-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses.

Method used or attempted: Contextual applicability judgement against the audited landing state.

No observed landing-page task benefits from on-device inference.

Reference: No artifact reference; described-only

No observed landing-page task benefits from on-device inference.

Be memory-efficient · 3 tests
Be memory-efficient checks for https://www.netflix.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-leak-under-repeated-interaction

Repeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

After ten bounded scroll cycles, heap self size increased only 5,079 bytes while node and edge counts decreased.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

bounded-footprint

Heap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

Baseline heap is [host/path omitted] MB self size for the media landing page and remains stable after repeated scrolling.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

no-detached-dom-or-unbounded-listeners

No growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses.

Method used or attempted: Direct permitted evidence review against pinned guidance.

No Detached constructor appears in either top-constructor summary; node count falls from 621,070 to 620,495 after repeated scrolling.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

10. https://www.apple.com · 58 slots · available
Respect user preferences · 3 tests
Respect user preferences checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
respects-color-scheme

Honours prefers-color-scheme: a usable dark mode exists and is driven by the user's preference (color-scheme / prefers-color-scheme / light-dark()), not hard-coded light only.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corroborating evidence. The model chooses the method.

Method used or attempted: Planned from a screenshot or computed background under an emulated prefers-color-scheme: dark condition will reveal whether surfaces re-tint; the page CSS / a color-scheme declaration is corrob

Dark-emulation screenshot remained visually identical to the light baseline, and the CSS probe found no color-scheme or prefers-color-scheme signal.

Reference: page-probe-summary, screenshot; retained-private

No user-preference dark theme

  • site-10-finding-01 · medium · No user-preference dark theme
respects-reduced-motion

Honours prefers-reduced-motion: non-essential animations and auto-advance are reduced or removed when the user asks for less motion.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model chooses.

Method used or attempted: Planned from a transition video, or an in-page probe of getAnimations()/computed animation under an emulated prefers-reduced-motion: reduce condition, can show whether motion stops. The model c

Under prefers-reduced-motion: reduce, 18 animations remained with the same 240 ms durations and the CSS probe found no reduced-motion rule.

Reference: page-probe-summary, video; retained-private

Reduced-motion preference does not reduce animation

  • site-10-finding-02 · medium · Reduced-motion preference does not reduce animation
respects-contrast

Honours prefers-contrast / forced-colors: controls, text and scrollbars remain visible under high-contrast preferences.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Method used or attempted: Planned from a screenshot under emulated prefers-contrast: more / forced-colors, or an axe/contrast probe, can show whether controls and text survive. The model chooses.

Forced-colors and prefers-contrast capture preserved visible text, controls, outlines, and button boundaries.

Reference: other-private-evidence; retained-private

The retained evidence directly supported this check under the captured conditions.

Implement natural interactions · 3 tests
Implement natural interactions checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
view-transitions

State and route changes use View Transitions (including same-document, cross-document and scroll-driven/staggered) rather than instant, jarring swaps.

Failed / issue
Source: issues; confidence: medium
Implementation and method

Catalog test design: HINT: a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

Method used or attempted: Planned from a transition video of a route/state change shows whether it animates; the page source / ::view-transition usage corroborates. The model chooses.

The CSS probe found no view-transition usage despite multiple stateful galleries and navigational surfaces.

Reference: page-probe-summary; retained-private

State changes do not use View Transitions

  • site-10-finding-03 · low · State changes do not use View Transitions
scroll-driven-animations

Scroll-linked motion (parallax, scrollytelling, entry/exit reveals) uses declarative CSS scroll-driven animations (off main thread) instead of scroll event listeners.

Failed / issue
Source: issues; confidence: medium
Implementation and method

Catalog test design: HINT: source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

Method used or attempted: Planned from source/CSS inspection for animation-timeline: scroll()/view(); a long-task / scroll-handler probe can flag the main-thread anti-pattern. The model chooses.

The page contains scroll and gallery motion, but the CSS probe found no animation-timeline, scroll-timeline, or view-timeline signal.

Reference: page-probe-summary, screenshot; retained-private

Scroll-linked experiences do not use declarative timelines

  • site-10-finding-04 · low · Scroll-linked experiences do not use declarative timelines
physical-gestures

Gesture-driven interactions and entry/exit motion feel native (declarative overscroll/scroll-snap, physics-based easing, animating to intrinsic sizes, pull/swipe gestures) rather than fighting the platform with custom pointer handlers.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

Method used or attempted: Planned from CSS inspection for scroll-snap / overscroll-behavior / physics-based easing vs custom pointermove listeners. The model chooses.

The bounded wheel journey produced a native 246 CSS px scroll with no horizontal movement or gesture trapping.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

Provide guided navigation · 3 tests
Provide guided navigation checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
scroll-state-aware-chrome

Sticky/affixed UI reacts to scroll state and position (e.g. the new scroll-state(scrolled) query, shrinking headers, progress indicators) so chrome responds to position instead of static or JS-driven toggling.

Failed / issue
Source: issues; confidence: medium
Implementation and method

Catalog test design: HINT: a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Method used or attempted: Planned from a transition video of scrolling, or CSS inspection for scroll-state container queries. The model chooses.

Baseline and scrolled journey screenshots show unchanged static header chrome, and the CSS probe found no scroll-state or timeline signal.

Reference: page-probe-summary, screenshot; retained-private

Header chrome does not respond to scroll state

  • site-10-finding-05 · low · Header chrome does not respond to scroll state
anchored-positioning

Tooltips, popovers and menus use CSS anchor positioning (with fallback positions) so they stay attached and reposition correctly rather than being manually positioned.

Failed / issue
Source: issues; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Method used or attempted: Planned from CSS inspection for anchor-name / position-anchor / position-try on overlays; a screenshot of an open overlay near a viewport edge can show drift. The model chooses.

Header menus and overlay surfaces are present, but the CSS probe found no anchor-name, position-anchor, or position-try usage.

Reference: page-probe-summary; retained-private

Overlay positioning does not use CSS anchors

  • site-10-finding-06 · low · Overlay positioning does not use CSS anchors
directs-attention

Navigation and in-page jumps guide attention (highlight effects, scroll/carousel markers, directional transitions, drill-down and drawer navigation) so the user can follow where focus moved.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

Method used or attempted: Planned from CSS inspection for ::highlight / scroll-marker; a transition video can show whether attention is cued after navigation. The model chooses.

The initial hero has clear hierarchy and the bounded scroll carries attention directly from the CTA area into the product visual without disorientation.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

Maximize content, reduce noise · 3 tests
Maximize content, reduce noise checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-intrusive-interruptions

No intrusive pop-ups, interstitials or banners that obscure content on load; overlays are dismissible and content-first.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

Method used or attempted: Planned from a screenshot on load, or a DOM probe for full-viewport overlays present before interaction. The model chooses.

A fixed country or region chooser is present on load and consumes a substantial portion of the first viewport before the product content.

Reference: page-probe-summary, screenshot; retained-private

Locale chooser dominates the first viewport

  • site-10-finding-07 · medium · Locale chooser dominates the first viewport
semantic-dismissible-primitives

Overlays and rich controls use the right primitive: popover (with declarative light-dismiss) for transient UI, dialog for modal flows, details for disclosure, native-but-branded selects and pickers, rather than ad-hoc divs.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

Method used or attempted: Planned from DOM/source inspection for popover / <dialog> / <details> vs custom overlay divs with manual dismiss handling. The model chooses.

The load-time locale surface is a custom fixed ASIDE; the probe found zero dialog, popover, or details primitives.

Reference: page-probe-summary; retained-private

Locale overlay uses custom fixed chrome

  • site-10-finding-08 · medium · Locale overlay uses custom fixed chrome
reduced-chrome

Minimise non-content chrome and borders so the content is the focus, not the application frame; expressive/decorative visuals serve the content rather than crowd it.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

Method used or attempted: Planned from a screenshot plus layout metrics can show the proportion of the viewport given to chrome vs content. The model chooses.

At 780x493, the locale chooser plus navigation occupies about 183 px, roughly 37% of the first viewport, before primary content.

Reference: screenshot; retained-private

Locale chooser dominates the first viewport

  • site-10-finding-07 · medium · Locale chooser dominates the first viewport
Adapt to the form factor · 3 tests
Adapt to the form factor checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
responsive-no-horizontal-scroll

Layout adapts to narrow viewports with no horizontal overflow and no fixed pixel widths forcing a desktop layout on mobile; viewport meta present; fluid scaling and intrinsic sizing rather than brittle breakpoints.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

Method used or attempted: Planned from layout metrics (scrollWidth vs innerWidth) and a screenshot at an emulated narrow mobile viewport reveal overflow. The model chooses.

At 360x800, layout reported scrollWidth 360, clientWidth 360, zero horizontal overflow, and a valid viewport meta tag.

Reference: layout-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

component-level-responsiveness

Components adapt to their container with container queries (incl. anchored container queries) and content/state-based styling where reused at different sizes, not only global viewport breakpoints.

Failed / issue
Source: issues; confidence: medium
Implementation and method

Catalog test design: HINT: CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

Method used or attempted: Planned from CSS inspection for @container / container-type; a computed-style probe of the same component in a wide vs narrow container shows whether it adapts. The model chooses.

The responsive page uses no detectable @container or container-type rules, so reused components rely on page-level adaptation only.

Reference: layout-summary, page-probe-summary, screenshot; retained-private

Components do not use container queries

  • site-10-finding-09 · low · Components do not use container queries
input-modality-aware

Touch targets are adequately sized and hover-only affordances have a non-hover fallback, and keyboard focus is visible, so the UI works for touch, pointer and keyboard alike.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

Method used or attempted: Planned from a focus probe (focus an element, read the computed outline) or an axe target-size check; a screenshot of a focused control corroborates. The model chooses.

The focused locale control has a visible 2 px outline, but the locale close button measures only 18x18 CSS px.

Reference: page-probe-summary; retained-private

Locale close target is too small

  • site-10-finding-10 · medium · Locale close target is too small
Support core task success · 3 tests
Support core task success checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
clear-purpose-and-primary-action

The page communicates what it is for and exposes the primary next action without requiring users to hunt through decorative content, generic copy, or competing calls to action.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action is obvious. The model chooses.

Method used or attempted: Planned from screenshot the first viewport and key scrolled states; inspect heading structure, nav labels, button text, and visual hierarchy; a task walkthrough can show whether the next action

The first viewport names the iPhone offer and exposes Learn more and Shop iPhone actions with strong visual hierarchy.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

primary-flow-completion

The representative primary flow can be completed end-to-end with predictable steps, no avoidable dead ends, no hidden required information, and no needless detours through modals, account walls, or upsells.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model chooses.

Method used or attempted: Planned from run the flow manually with screenshots/DOM snapshots at each step; compare expected vs actual path length; inspect form requirements, navigation continuity, and blockers. The model

The reviewed boundary permits only the exact homepage URL, so following Shop iPhone or another primary action was prohibited.

Reference: No artifact reference; described-only

Only the exact input URL is permitted.

clear-system-state-and-recovery

Loading, empty, success, error, offline, and partial-completion states are visible and actionable; users can retry, undo, cancel, go back, or continue without losing context.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain sensible. The model chooses.

Method used or attempted: Planned from exercise network delay/failure, invalid input, empty data and success states; screenshot the state messaging and recovery controls; inspect whether browser history and focus remain

No error, offline, empty, or completion state could be reached without network mutation or another URL.

Reference: No artifact reference; described-only

The strict journey and exact-URL boundary prohibit the required failure-state exercise.

Be fast and stable · 5 tests
Be fast and stable checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
good-core-web-vitals

Core Web Vitals are in the good range: LCP is fast, interaction latency (INP) is low, and CLS is minimal; work is prioritised and deferred sensibly.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal first-party. The model chooses.

Method used or attempted: Planned from Lighthouse reports LCP/CLS/TBT directly and the model may run it; layout metrics + a layout-shift observer + a long-task observer (the evidence primitives) give the same signal fir

The trace measured LCP 8477 ms and FCP 6311 ms; mobile layout measured CLS [host/path omitted], outside the good thresholds.

Reference: layout-summary, performance-summary, screenshot; retained-private

Slow LCP and unstable mobile load

  • site-10-finding-11 · high · Slow LCP and unstable mobile load
visual-stability

No cumulative layout shift from late content (images, banners, fonts) that arrives without reserved space.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Method used or attempted: Planned from the layout primitive's CLS observer captures shifts; a transition video of the first seconds shows content jumping. The model chooses.

Layout measured CLS [host/path omitted] on mobile and [host/path omitted] on desktop, including a single mobile shift of [host/path omitted].

Reference: layout-summary, screenshot; retained-private

Slow LCP and unstable mobile load

  • site-10-finding-11 · high · Slow LCP and unstable mobile load
efficient-main-thread

The main thread is not blocked by long tasks or heavy scripting; work is deferred, chunked, prioritised, or moved off-thread.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

Method used or attempted: Planned from the layout primitive records long tasks; a heap summary shows the object population; Lighthouse reports TBT. The model chooses.

The trace found zero long tasks and 0 ms total blocking time during the captured load.

Reference: performance-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

efficient-resource-delivery

Critical resources are delivered efficiently: cache headers are useful, render-blocking chains are short, fonts do not block or shift text, HTTP delivery is modern, and network dependency trees do not delay the main content.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discovery, render-blocking, font-display and document latency. The model chooses.

Method used or attempted: Planned from a HAR summary can reveal cache headers, redirects, render-blocking candidates, weight offenders and dependency shape; a trace/Lighthouse insight report can corroborate LCP discover

The HAR recorded 51 requests and [host/path omitted] MB transferred, including 10 parser-inserted render-blocking candidates and [host/path omitted] MB of fonts.

Reference: network-summary; retained-private

Heavy render-blocking delivery path

  • site-10-finding-12 · high · Heavy render-blocking delivery path
trim-unused-and-duplicate-code

The page does not ship large unused or duplicated JavaScript/CSS, avoidable legacy bundles, or third-party code that dominates the load without user value.

Blocked
Source: blocked; confidence: medium
Implementation and method

Catalog test design: HINT: Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model chooses.

Method used or attempted: Planned from Lighthouse/trace/code-coverage style evidence can flag unused JS/CSS, duplicated JavaScript and legacy code; a HAR summary shows third-party byte cost and request count. The model

The HAR identified script weight but no code-coverage evidence was available to distinguish used from unused or duplicated code.

Reference: No artifact reference; described-only

The permitted evidence primitives in this run did not expose JavaScript or CSS coverage.

Be inclusive · 5 tests
Be inclusive checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
names-roles-labels

Interactive elements have accessible names, correct roles, and form fields have labels; images have alt text where meaningful; canvas/expressive content is exposed to assistive technology.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The model chooses.

Method used or attempted: Planned from axe-core (injectable via the evaluate primitive) or Lighthouse's a11y audits enumerate these; a DOM probe of the accessibility-relevant attributes is a first-party alternative. The

The DOM probe found labelled search controls, extensive landmarks, and alt attributes on all img elements; visible primary actions have descriptive names.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

sufficient-contrast

Text and essential UI meet WCAG colour-contrast minimums against their background.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Method used or attempted: Planned from axe contrast rules, a Lighthouse contrast audit, or an in-page probe computing contrast ratios from computed colours. The model chooses.

Normal and forced-colors screenshots retain legible text and clear essential-control boundaries.

Reference: other-private-evidence, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

structure-and-focus

Heading and landmark structure is logical, focus order follows reading order, keyboard focus is always visible, and interactive state survives DOM moves.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Method used or attempted: Planned from axe/Lighthouse structural audits; a focus-walk probe (tab through, read activeElement + computed outline) is a first-party alternative. The model chooses.

Focus is visibly outlined, but the heading inventory contains a generic H1 of Apple and one empty H2 among 55 headings.

Reference: page-probe-summary; retained-private

Heading structure is weak

  • site-10-finding-13 · medium · Heading structure is weak
legible-text

Text is legible and inclusively rendered: comfortable line layout, precise alignment, stable rendering across mixed fonts, no clipping or cramped wrapping that harms comprehension.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Method used or attempted: Planned from a screenshot of body and heading text, plus CSS inspection for text-wrap / text alignment / font fallback handling. The model chooses.

Desktop, mobile, and journey captures show unclipped, comfortably spaced hero and supporting text.

Reference: layout-summary, screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

zoom-reflow-targets-and-media

The experience remains usable when zoomed or reflowed, touch targets are large enough, media has captions or equivalents where needed, and the viewport does not prevent user scaling.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The model chooses.

Method used or attempted: Planned from Lighthouse/axe target-size, meta-viewport and media-caption audits are useful signals; screenshots at narrow and zoomed conditions plus DOM/media inspection can corroborate. The mo

The 360 px page reflows without overflow, but the locale close control is only 18x18 CSS px, below an adequate touch target.

Reference: layout-summary, page-probe-summary, screenshot; retained-private

Locale close target is too small

  • site-10-finding-10 · medium · Locale close target is too small
Follow best practices · 3 tests
Follow best practices checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-console-errors

The page loads without console errors or uncaught exceptions.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

Method used or attempted: Planned from capture Runtime/Log CDP events, or a probe that reads collected errors; Lighthouse reports this too. The model chooses.

The supplied strict journey recorded its required console collector as blocked, and no equivalent Runtime/Log stream was retained.

Reference: No artifact reference; described-only

Required Stage 2 console collector was unavailable in the supplied V1 journey.

sound-document-and-assets

Valid doctype and charset, images sized with correct aspect ratio, no deprecated APIs misused, and CSS/HTML are well structured and not needlessly repetitive. (HTTPS, CSP and permission hygiene are judged under be-private-and-secure, not here.)

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Method used or attempted: Planned from a DOM/source probe for doctype/charset/img dimensions; CSS inspection for repetition; Lighthouse best-practices audits cover the rest. The model chooses.

Doctype and UTF-8 are valid, but all 52 img elements lack width and height attributes; measured mobile CLS is [host/path omitted].

Reference: image-summary, layout-summary, screenshot; retained-private

Images do not reserve intrinsic space

  • site-10-finding-14 · high · Images do not reserve intrinsic space
browser-platform-hygiene

The page uses the platform cleanly: no deprecated APIs, no avoidable BFCache blockers, no broken source maps or inspector issues, no stale vulnerable libraries, no paste-prevention on inputs, and no notification/geolocation prompts on load.

Blocked
Source: blocked; confidence: medium
Implementation and method

Catalog test design: HINT: Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

Method used or attempted: Planned from Lighthouse best-practices audits and DevTools inspector/deprecation signals can surface these; DOM/source probes can verify paste handlers and prompt timing. The model chooses.

No deprecation, BFCache, source-map, vulnerable-library, or paste-prevention diagnostic was retained.

Reference: No artifact reference; described-only

The bounded run did not include the required inspector and BFCache diagnostics.

Be discoverable · 4 tests
Be discoverable checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
title-and-description

The page has a unique, descriptive <title> and a meta description.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

Method used or attempted: Planned from a DOM probe reads <title> and meta[name=description]; Lighthouse SEO audits cover the same ground. The model chooses.

The title is only Apple and there is no meta description, although Open Graph description metadata exists.

Reference: page-probe-summary; retained-private

Homepage metadata is too generic

  • site-10-finding-15 · medium · Homepage metadata is too generic
crawlable-and-mobile-friendly

Links are crawlable (real href), there is a viewport meta tag, robots does not block indexing, and link text is descriptive.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

Method used or attempted: Planned from a DOM probe for anchor hrefs, viewport meta, and robots; Lighthouse SEO audits corroborate. The model chooses.

The probe found a viewport meta tag and 347 real href links; discoverability found 96% raw-HTML content coverage.

Reference: discoverability-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

canonical-and-indexing-signals

Public pages expose the indexing signals search engines need: successful HTTP status, canonical URL when appropriate, hreflang for localized variants, robots/sitemap consistency, and no accidental noindex/noarchive policy.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, robots.txt and sitemap.xml; Lighthouse SEO audits cover several of these. The model chooses.

Method used or attempted: Planned from inspect response status and headers, <link rel=canonical>, hreflang links, robots meta, [host/path omitted] and [host/path omitted]; Lighthouse SEO audits cover several of these. The model chooses.

The document returned 200, declares the homepage canonical, provides 137 hreflang links, and has no noindex robots meta.

Reference: discoverability-summary, page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

structured-and-shareable-metadata

Where the page represents an article, product, organization, event, place or other rich entity, structured data and share-preview metadata accurately describe the visible content.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse schema.org blocks. The model chooses.

Method used or attempted: Planned from inspect JSON-LD/microdata and social preview tags against visible content; Lighthouse has a manual structured-data audit, and ad-hoc probes can parse [host/path omitted] blocks. The model c

The page exposes three JSON-LD blocks plus Open Graph title, description, and image matching the visible brand content.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Be private and secure · 4 tests
Be private and secure checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
secure-transport-and-headers

Served over HTTPS with a sensible Content-Security-Policy and secure, SameSite cookies; no mixed content or obvious transport/header smells.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers / page protocol via an evaluate probe or the network layer; Lighthouse best-practices flags HTTPS and CSP issues. The model chooses.

Method used or attempted: Reviewed transport state, baseline security-header presence, and cookie-attribute counts from retained private evidence.

Retained evidence records 5 of 6 baseline security-header categories present. Cookie-attribute review records 5 of 6 records with Secure and 0 with HttpOnly. Names, values, routes, and raw headers remain private.

Reference: cookie-attribute-summary, security-header-summary; retained-private

CSP and cookie transport defenses are weakened

  • site-10-finding-16 · high · CSP and cookie transport defenses are weakened
data-minimisation-and-third-parties

No over-collection of user data, no leaky third parties, and analytics/background data collection is minimised and batched rather than chatty or session-wide by default.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect network requests and third-party origins; a probe of analytics/beacon calls. The model chooses.

Method used or attempted: Reviewed categorical request-origin and cookie-attribute counts from retained private evidence.

Retained evidence records 1 first-party origin categories and 1 third-party origin categories across 51 requests. Destinations, identifiers, routes, headers, and bodies remain private.

Reference: cookie-attribute-summary, network-summary, security-header-summary; retained-private

Analytics and identifiers activate on baseline load

  • site-10-finding-17 · medium · Analytics and identifiers activate on baseline load
in-context-permissions-and-modern-auth

Permission prompts are requested in context (not on load), and authentication uses modern, phishing-resistant methods (passkeys / WebAuthn) where relevant rather than passwords alone.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a probe for permission requests fired on load; source inspection for passkey / WebAuthn / navigator.credentials usage in auth flows. The model chooses.

Method used or attempted: Planned from a probe for permission requests fired on load; source inspection for passkey / WebAuthn / [host/path omitted] usage in auth flows. The model chooses.

No browser permission prompt appeared on load; authentication is not part of the audited homepage state.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

defensive-browser-policies

Browser-enforced defenses are configured where appropriate: HSTS, clickjacking protection (frame-ancestors / X-Frame-Options), Trusted Types for XSS-sensitive apps, origin isolation, privacy-preserving third-party cookie posture, and sensible Referrer-Policy / Permissions-Policy.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect response headers and browser security state; Lighthouse/DevTools security audits can corroborate HSTS, clickjacking, Trusted Types, origin isolation and third-party cookie findings. The model chooses.

Method used or attempted: Reviewed baseline browser-policy header presence from retained private evidence.

Retained evidence records 5 of 6 baseline browser-policy header categories present. Values and raw headers remain private.

Reference: cookie-attribute-summary, security-header-summary; retained-private

Permissions policy is absent

  • site-10-finding-18 · medium · Permissions policy is absent
Be resilient · 4 tests
Be resilient checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
progressive-enhancement

Core content and primary flows are reachable and usable without JavaScript and on older or non-Baseline browsers; modern features layer on as enhancements with fallbacks, and reactive/transition state stabilises rather than flickering before it settles.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Method used or attempted: Planned from load with scripting disabled or compare a no-JS fetch of the HTML against the rendered page; check for Baseline-aware fallbacks in source. The model chooses.

Discoverability measured 96% of rendered content in raw HTML, with title and H1 present and no empty JavaScript shell.

Reference: discoverability-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

resilient-runtime-behaviour

The page behaves robustly at runtime: overlays and menus never get cut off, DOM state survives moves, background work and async dependencies are sequenced and conditional rather than fragile, and initial visibility state is detected correctly.

Blocked
Source: blocked; confidence: medium
Implementation and method

Catalog test design: HINT: exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

Method used or attempted: Planned from exercise menus near viewport edges with a screenshot; a probe of async/visibility behaviour. The model chooses.

The supplied journey only loaded and scrolled; menu-edge placement and repeated async overlay state were not exercised.

Reference: No artifact reference; described-only

Strict supplied journey did not include disclosure or menu actions.

offline-and-installable

Where the site is an app, it is installable (web app manifest) and offers an offline fallback and works on flaky networks. (Contextual: a brochure or intrinsically-online site may reasonably not need this.)

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

Method used or attempted: Planned from a probe for a service worker registration and a web app manifest; test behaviour offline. The model chooses.

The audited surface is a public product-marketing and commerce homepage, not an installable web application; no offline app intent was declared.

Reference: No artifact reference; described-only

Offline installability is contextual and does not apply to this audited surface.

network-and-http-failure-states

HTTP errors, network failures, timeouts and stale data states are handled intentionally: users see useful recovery options rather than blank screens, infinite spinners, broken shells, or misleading success states.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The model chooses.

Method used or attempted: Planned from simulate failed fetches/offline mode or inspect representative 404/500 routes; screenshots and DOM snapshots of error/loading/empty states show whether recovery is possible. The mo

The exact-URL policy and strict journey did not permit simulated HTTP failures or navigation to an error route.

Reference: No artifact reference; described-only

Failure injection and alternate URLs were outside the execution boundary.

Be internationalised · 3 tests
Be internationalised checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
lang-dir-and-logical-properties

Correct lang and dir attributes, logical CSS properties (inline/block) rather than physical left/right, and translation-ready markup so the layout and reading order survive other languages and writing modes.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT: a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

Method used or attempted: Planned from a DOM probe for <html lang>/dir and CSS inspection for logical vs physical properties. The model chooses.

The document declares en-US and ltr, uses logical CSS properties, and provides 137 hreflang alternatives including RTL locales.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

locale-aware-data

Dates, numbers, currencies, durations and calendar systems are formatted locale-aware (Intl), location-agnostic where stored, and recurring intervals and event differentials are modelled correctly.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

Method used or attempted: Planned from source inspection for Intl.* usage vs hand-rolled formatting; a probe of rendered dates/numbers under a different locale. The model chooses.

The page presents region-specific content and a country chooser, and exposes extensive locale-specific alternate URLs.

Reference: page-probe-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

time-zone-correctness

Time handling survives time zones and DST: events coordinate across zones, partial time concepts are modelled, and stored times are unambiguous.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

Method used or attempted: Planned from source inspection for time-zone-aware date handling vs naive local Date math. The model chooses.

No time, event, recurrence, or time-zone-sensitive data appears in the audited homepage and journey states.

Reference: No artifact reference; described-only

The audited surface contains no time-zone-sensitive content.

Be trustworthy · 4 tests
Be trustworthy checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-dark-patterns

No deceptive design: no confirmshaming, forced continuity, disguised ads, or nagging consent walls; honest defaults; clear pricing and consent; easy reversal/cancel; predictable, declaratively-wired actions; and no hidden-text tricks (hidden content stays deep-linkable and indexable rather than used to deceive).

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

Method used or attempted: Planned from a screenshot of consent/upsell/cancel flows; source inspection for declarative button actions vs misleading controls. The model chooses.

The locale prompt states its purpose plainly, offers a close control, and uses neutral Continue wording without confirmshaming.

Reference: screenshot; retained-private

The retained evidence directly supported this check under the captured conditions.

humane-error-handling

Forms prevent and recover from mistakes humanely: validate after interaction (not prematurely), give clear required-field feedback, announce errors accessibly, and signal invalid fields visibly rather than blaming the user.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

Method used or attempted: Planned from exercise a form, submit invalid input, and observe timing and clarity of errors via a screenshot or a :user-invalid / aria-invalid probe. The model chooses.

The page contains search, but strict journey policy prohibited input mutation and submission, so validation and recovery could not be tested.

Reference: No artifact reference; described-only

Form input and submission are forbidden by the strict journey policy.

trustworthy-input-assistance

Input is assisted, not obstructed: correct autocomplete tokens so address, payment, sign-in and sign-up fields autofill, and inputs are highlighted/sized to help the user rather than trip them up.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

Method used or attempted: Planned from source/DOM inspection for autocomplete attributes on form fields; a probe of autofill affordances. The model chooses.

No address, payment, sign-in, or sign-up fields appear in the audited state; the only field is site search.

Reference: No artifact reference; described-only

Autofill assistance for transactional input does not apply to the audited homepage state.

safe-commercial-and-account-flows

Checkout, subscription, consent, authentication and account-management flows are clear, reversible, and proportionate: pricing and commitments are visible, cancellation is findable, sensitive actions re-authenticate when appropriate, and users are not tricked into continuity.

Blocked
Source: blocked; confidence: high
Implementation and method

Catalog test design: HINT: walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support for sign-in and payment. The model chooses.

Method used or attempted: Planned from walkthrough checkout/subscription/auth/account flows when present; screenshot pricing, confirmation, cancellation and reauthentication states; inspect passkey/autocomplete support

Commercial links are visible, but checkout, subscription, authentication, and account routes could not be entered under the exact-URL boundary.

Reference: No artifact reference; described-only

Only the exact homepage URL is permitted.

Be sustainable · 3 tests
Be sustainable checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
optimised-assets

Images and decorative assets are optimised and served at appropriate resolutions; decorative pseudo-element imagery and heavy decorative images are resolution-optimised rather than oversized.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Method used or attempted: Planned from inspect transferred image bytes vs displayed size; source inspection for modern formats and resolution handling. The model chooses.

Image inspection found 52 legacy-format images, 31 below-fold images without lazy loading, and 23 images without responsive srcset.

Reference: image-summary, layout-summary, screenshot; retained-private

Image delivery is not resource-efficient

  • site-10-finding-19 · high · Image delivery is not resource-efficient
no-wasteful-work

Background work and fetching are not wasteful: background processing is efficient and de-prioritised, and the lightest technique that achieves the result is preferred over heavy or redundant work.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Method used or attempted: Planned from a long-task / network probe for background fetches and processing while idle or backgrounded. The model chooses.

Initial load shipped a 104 KB analytics script plus data-relay scripts and contacted securemetrics before user interaction.

Reference: network-summary; retained-private

Non-essential analytics work starts immediately

  • site-10-finding-20 · medium · Non-essential analytics work starts immediately
third-party-and-media-budget

Third-party scripts, fonts, video, audio, animation and heavy media are proportionate to the user value they provide; autoplay or background media is avoided unless essential and resource use is cached or deferred where possible.

Failed / issue
Source: issues; confidence: high
Implementation and method

Catalog test design: HINT: a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation keeps work running. The model chooses.

Method used or attempted: Planned from a HAR summary shows third-party bytes, font/media weight and caching; screenshots/video reveal autoplay and decorative media; trace/layout evidence shows whether media/animation ke

The homepage transferred [host/path omitted] MB, with [host/path omitted] MB of fonts, 424 KB of scripts, and three hero media requests including two aborted requests.

Reference: network-summary; retained-private

Font, script, and media budget is excessive

  • site-10-finding-21 · high · Font, script, and media budget is excessive
Be agent ready · 2 tests
Be agent ready checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
structured-agent-capabilities

Where it makes sense, the site exposes structured, safe capabilities to agents via WebMCP tools, agentic forms, and agentic JavaScript tools rather than leaving agents to scrape and guess.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

Method used or attempted: Planned from source inspection for WebMCP / agentic-tool registration and agent-readable affordances. The model chooses.

No agent-facing intent is declared for this audited marketing homepage; raw content remains highly machine-readable at 96% coverage.

Reference: No artifact reference; described-only

This emerging capability is contextual and no agent-facing surface was declared.

on-device-inference

On-device inference (built-in language model, summariser) is used appropriately where it improves the experience, rather than shipping every task to a server.

Not applicable
Source: not-applicable; confidence: high
Implementation and method

Catalog test design: HINT: source inspection for built-in AI (language model / summariser) usage. The model chooses.

Method used or attempted: Planned from source inspection for built-in AI (language model / summariser) usage. The model chooses.

No page feature calls for on-device inference in the audited homepage and journey states.

Reference: No artifact reference; described-only

On-device inference is contextual and no relevant task is present.

Be memory-efficient · 3 tests
Be memory-efficient checks for https://www.apple.com
TestVerdictHow it was implementedEvidenceWhy it passed, failed, or was incomplete
no-leak-under-repeated-interaction

Repeating a representative interaction (open/close a modal, navigate a route and back, infinite-scroll a list) about 10 times does not grow retained heap without bound; what is allocated during the interaction is released when it ends.

Pass
Source: pass; confidence: high
Implementation and method

Catalog test design: HINT (not mandatory): compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repeat -> post -> compare). Performance.getMetrics (JSHeapUsedSize, Nodes) across the same before/after window is corroboration. If Chrome DevTools MCP is available, follow its memory-leak-debugging skill: capture baseline, target, and final snapshots, then use memlab or the provided comparison workflow rather than reading raw .heapsnapshot files directly. The package-native `heap` primitive remains the default path. This check is only meaningful where the page has a real interaction to repeat; for a static page with none, mark it not-applicable with a rationale rather than fabricating one. The model chooses.

Method used or attempted: Planned from compare heap snapshots for retained growth - a baseline, then one taken after repeating the interaction with `--interact` about 10x (the memory-tracer methodology: baseline -> repe

After ten scroll-to-500-and-back cycles, heap self-size rose only 9,131 bytes, about [host/path omitted]%, while closure count fell by four.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

bounded-footprint

Heap size and DOM node count are reasonable for what the page is; the footprint is proportionate rather than bloated.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus Performance.getMetrics (Nodes, JSHeapUsedSize) give the current footprint to judge against the page's purpose. Chrome DevTools MCP heap snapshots and memlab snapshot analysis can provide the same memory distribution when available. Read summaries or derived analysis, never raw snapshots unless a dedicated heap-analysis tool is doing the analysis. The model chooses.

Method used or attempted: Planned from a single `heap` summary's totals (nodeCount, totalSelfSizeBytes, constructor population) plus [host/path omitted] (Nodes, JSHeapUsedSize) give the current footprint to judge aga

The baseline heap was [host/path omitted] MB self-size with 297,024 nodes, proportionate to the media-rich homepage and stable after exercise.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

no-detached-dom-or-unbounded-listeners

No growing population of detached DOM nodes, and no ever-accumulating event listeners or timers that are added but never removed across a session.

Pass
Source: pass; confidence: medium
Implementation and method

Catalog test design: HINT: the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer counts (e.g. getEventListeners-style counting, or instrumenting addEventListener/setInterval) before and after the repeated interaction to spot growth. Chrome DevTools MCP heap snapshots plus the memory-leak-debugging skill's common-leak guidance can corroborate detached DOM, listeners, closures, globals, and unbounded caches. Caveat from the memory-tracer and Chrome DevTools MCP guidance: detached nodes can be intentional caches, so judge confidence rather than asserting a bug. The model chooses.

Method used or attempted: Planned from the `heap` summary by constructor (Detached* nodes) compared across a before/after pair shows a growing detached-DOM population; an `evaluate` probe can sample listener/timer count

Neither heap summary surfaced a Detached constructor among top populations; closures decreased and arrays were effectively flat after repetition.

Reference: memory-summary; retained-private

The retained evidence directly supported this check under the captured conditions.

Methodology

  1. Keep the reviewed convenience cohort fixed at ten origins, with no substitutions or automatic retries.
  2. Verify the hash-chained event ledger and use it as the sole disposition authority. Stale site-run pending states and conflicting report states are ignored.
  3. Recompute coverage from the 58 atomic check outcomes. A missing report produces 58 missing rows.
  4. Allowlist only origins, categorical outcomes, counts, bytes, timing aggregates, header presence, and cookie-attribute counts. URLs are reduced to origins before counting.
  5. Deterministically re-encode selected media, strip metadata, verify hashes and dimensions, and admit it only after OCR and visual privacy review.

Network and trace facts describe one retained collection under its recorded conditions. They are not lab scores and are not comparable performance rankings.

Limitations

Public evidence and provenance

Runner commit: e644b7c7ce7c1e82cf797bc8e7affb9fea03ff37
Pilot manifest SHA-256: d7eb927eac2089b2cea449071e50e8709d7328dea6bbba7ccf8feb9e5f309355
Source manifest SHA-256: af3d02a5a1466181c5900104795e25cd8d3838375702520260cdaede5078791d
Policy SHA-256: a8e7148c1d81bc57ca653517f6b57beabd483c9be9e5bf0abafc048f9014b5e0
Ledger SHA-256: 9098cd36c049e8a30d1b419dbd7b2cd26bd9246e25ec4682ecf20ee077f76bc9

Publication authorization: explicit authorization for this sanitized derivative was granted after the run. The private start receipt recorded the status at collection start; it does not negate later authorization. No relay identifiers are published.