| name | site |
| description | Full website audit — UX/UI design, marketing, SEO, accessibility, performance, conversion. Crawls the site (up to 100 pages) via cmux browser, generates detailed visual critique with actionable recommendations. |
Bishx-Site: Full Website Audit
Autonomous website audit skill. Crawls the site (up to 100 pages) via cmux browser, analyzes UX/UI design,
marketing, SEO, accessibility, performance, conversion, brand consistency, and information
architecture. Produces a single detailed visual critique document with actionable recommendations.
Output is a DESIGN CRITIQUE, not code. No source code inspection, no technical implementation
details. Pure UX/marketing/visual perspective: what's wrong, what to change it to, where exactly
to move it, and why it works better — at a level of detail where a designer or another agent
can execute without a single clarifying question.
FOUNDATIONAL PRINCIPLE: Human-First Evaluation
This is the foundation of the entire audit. Every module, every check, every finding
derives from this principle. Module-specific checks (CTA placement, contrast ratios,
heading structure, page speed) are CONCRETE INSTANCES of this abstract principle —
not independent rules.
The Principle
Every element on a page exists to serve a specific human in a specific moment.
When a page is organized by the creator's logic instead of the visitor's need —
that is a failure, regardless of how technically correct the page is.
A site can score 100 on Lighthouse, pass every WCAG check, have perfect SEO —
and still fail the human who visits it.
Five Layers + Meta-Layer
Every page evaluation — in any module — passes through this lens:
LAYER 1 — FOR WHOM (is the right person being served?)
- Is the content written for the person actually looking at this page?
- Does the tone match the visitor's emotional state at this point?
(anxious on payment page → needs reassurance, not excitement;
frustrated on error page → needs empathy, not branding)
- Does the page assume only what the visitor actually knows?
(no unexplained jargon, no references to internal concepts)
Visitor States Beyond "First Visit":
When identifying WHO is on the page (HB1), consider these states:
| State | What They Need | What Fails Them |
|---|
| Multi-tab comparison (3-5 competitor sites open) | Scannable data points (price, features). Comparable format. Quick extraction. | Buried pricing. Unique terminology. Non-standard metrics. |
| Frustrated / error-recovery (broken link, failed payment, bad search) | Acknowledgment of failure. Clear recovery path. Preserved context. | Generic error. Lost form data. No alternative suggestions. |
| Urgent (emergency info, time-critical) | ONE fact in <2 seconds. Zero interaction required. | Content hierarchy that buries the answer. Popups. Loading delays. |
| Reluctant / forced (must use, no alternative) | Minimum path to completion. No unnecessary friction. | Extra steps. Upsells. Marketing copy on utility pages. |
| Referred (friend sent link, high trust, zero context) | Page that stands alone without homepage context. | Assumes prior navigation. Jargon without explanation. |
| Expert / power-user (knows domain, seeks specific data) | Jump to data point. Skip preamble. Deep search. | Forced linear reading. No anchor links. No expandable details. |
| Price-sensitive (hunting for total cost, hidden fees) | Full cost transparent upfront. Clear comparison. | "from $X" without explanation. Fees added at checkout. |
| Situationally constrained (bright sun, one hand, slow connection, tiny screen) | Graceful degradation. High contrast. Large targets. Minimal data usage. | Light gray text. Tiny buttons. Heavy assets without lazy-load. |
These are NOT separate checks. They are PERSPECTIVES for the HB1 heartbeat.
When identifying the visitor, consider: which of these states might they be in?
The page should work for ALL likely visitor states, not just the "ideal" first-time desktop visitor.
LAYER 2 — WHAT (is the right content shown?)
- Is the page organized by the VISITOR'S logic or the CREATOR'S logic?
(creator: features, departments, chronology;
visitor: problem, solution, proof, action)
- Does the depth match the visitor's readiness?
(first visit → overview; comparison phase → details; ready to act → specifics)
- Nothing unnecessary, nothing critical missing?
LAYER 3 — HOW (is the presentation right?)
- Does the format match the visitor's cognitive mode?
(scanning → headings, icons, visual signals;
reading → full text, arguments;
deciding → comparison tables, calculators;
doing → step-by-step, not prose)
- Is information in the order the visitor needs it?
LAYER 4 — FEELING (how does the page FEEL to use?)
- Control: The visitor manages their experience, not the other way around.
No autoplay without consent. No forced registration before showing content.
No scroll hijacking. No popups interrupting flow. Back button works.
The visitor can bookmark, share, and return to any state.
- Familiar patterns: The site works like the rest of the web (Jakob's Law).
Logo top-left links to home. Search in header. Cart top-right.
"Next" button on the right, "Back" on the left.
Violating web conventions costs cognitive effort — every deviation must be justified.
- Overwhelm protection: The TOTAL cognitive demand of the page is manageable.
Each element may be justified individually, but 5 CTAs + 3 banners + popup +
ticker + chat widget + animation TOGETHER = overload.
Not about individual elements — about the SUM of everything competing for attention.
LAYER 5 — BEHAVIOR (how does the site RESPOND to the visitor?)
- Predictability: Every element shows what will happen when interacted with.
Button labeled "Download" downloads, not opens a registration form.
Link to "Pricing" shows pricing, not "Contact Sales."
If the outcome can't be predicted from the label — it's a failure.
- Transparency: No hidden agenda. No pre-checked opt-ins. No fees revealed
only at checkout. No "free trial" requiring credit card. No dark patterns.
The site works IN the visitor's interest, not AGAINST it.
- Feedback: Every visitor action gets acknowledgment.
Click button → button visually responds. Submit form → confirmation appears.
Add to cart → counter updates. Delete item → "Deleted" message with undo option.
Silence after action = the visitor doesn't know if it worked.
- Forgiveness: The site assumes humans make mistakes — and makes them reversible.
Accidentally deleted → can recover. Closed tab → form data saved.
Wrong button → can go back. Wrong input → can edit without restarting.
Undo is always available for destructive actions.
- Progress: The visitor sees forward momentum toward their goal.
Long form → progress bar. Multi-step process → "Step 2 of 4."
Long article → estimated reading time. Loading → progress indicator.
"How much more?" should always have an answer.
- Respect for time: Every element justifies its time cost.
No autoplay video forcing a wait. No interstitial before content.
No "subscribe" popup 3 seconds after loading. No unnecessary steps.
Every delay or interruption says "my time matters more than yours."
META-LAYER — THE JOURNEY (does the page work in context?)
- Promise = Reality: What the entry point promised (button text, link, search result)
matches what the page delivers
- Effort ≤ Value: The page doesn't ask more than it gives
(10-field form for a 2-page PDF = violation)
- Trust matches the ask: The bigger the commitment requested (payment, personal data),
the more trust the page must earn first
- Next step is clear: The visitor knows what to do after this page
- Memorability: After the visitor leaves — did something stick?
Can they describe the site in one sentence? Can they remember how it differs
from competitors? Will they find it again? Does the shared link preview
(OG tags) represent the page well?
How Modules Apply This Principle
Each module is a DOMAIN-SPECIFIC APPLICATION of this principle:
| Module | Layers 1-3 (Content) | Layer 4-5 (Experience) | Meta (Journey) |
|---|
| UX Design | Composition guides attention to what the VISITOR needs first, not what the creator wants to show | Overwhelm protection: total visual load manageable. Familiar layout patterns respected. | Memorability: distinctive visual identity |
| Marketing Content | Copy serves visitor's need in their language, depth, and tone. Audience-content fit. | Transparency: no manipulative copy, no false urgency. Predictability: CTA labels match outcomes. | Promise of headline = content of page |
| Conversion | Path follows visitor's decision logic. Right info at each decision point. | Control: no forced gates. Feedback: every form action confirmed. Forgiveness: mistakes reversible. Progress visible. | Effort ≤ value. Trust proportional to ask. |
| SEO Technical | Page structure delivers right visitor to right page. | Predictability: search snippet matches page content. | Promise = reality at search level |
| SEO Content | Answers visitor's question at needed depth in their language. | Respect for time: content is not padded. Every paragraph earns its place. | Visitor can describe what they learned |
| Accessibility | Every visitor accesses content regardless of ability. | Control: keyboard works. Feedback: screen reader gets announcements. Forgiveness: errors explained. | The ultimate human-first layer |
| Performance | Page loads fast enough to preserve visitor's intent. | Respect for time: no unnecessary waits. Progress: loading indicators present. | Fast enough that momentum isn't broken |
| Information Architecture | Navigation follows visitor's mental model, not org chart. | Familiar patterns: standard nav placement. Predictability: link labels match destinations. | Visitor always knows where they are and where to go |
| Brand Consistency | Consistent experience = visitor focuses on content, not on "is this the same site?" | Familiar within the site: same patterns everywhere. | Trust through consistency |
For Agents: How to Use This
When evaluating ANY element on ANY page:
- First — identify WHO is looking at this page and WHAT they need right now
- Then — check if the element serves that need (Layer 1-2-3)
- Then — check if the element works in the journey context (Meta-Layer)
- Only then — apply your module's specific technical checks
If a technical check PASSES but the principle is VIOLATED — that is a finding.
If a technical check FAILS but the principle is SATISFIED — note it but lower priority.
The principle does not replace technical checks. It PRIORITIZES and CONTEXTUALIZES them.
Toolkit, Not Checklist
Module-specific checks (FAIL/WARN thresholds) are a TOOLKIT of common patterns —
not a mandatory checklist to execute mechanically on every page.
How thresholds work:
- The Foundational Principle is your PRIMARY guide. Always.
- Thresholds (e.g., ">3 CTAs = WARN") are HELPERS that catch things you might miss
- When threshold and principle AGREE — great, report the finding
- When they DISAGREE — the principle wins. Explain the disagreement briefly.
- Override IS expected: in a typical audit, ~15-20% of threshold triggers are false positives
that should be dismissed or downgraded. If you override ZERO thresholds, you are working
mechanically — that is the actual failure.
- NOT overriding when the principle says otherwise is worse than overriding incorrectly
Purpose-level evaluation:
Each check in a module describes WHAT to evaluate, not HOW to mechanically verify.
- "Check if the page's call to action is clear" — NOT "count buttons and compare to 3"
- "Check if the table is usable" — NOT "verify each cell has minimum contrast"
- "Check if the form is approachable" — NOT "count fields and compare to 5"
The specific numbers/thresholds are EXAMPLES of what commonly indicates a problem.
The agent applies judgment to the actual page, using the principle as the guide.
Example:
A pricing page with 4 CTA buttons might PASS if each serves a different plan clearly.
A pricing page with 1 CTA button might FAIL if it's invisible or misleading.
The number of CTAs is an indicator. The principle ("is the next step clear?") is the test.
Unknown Element Protocol
When an agent encounters an element, widget, or pattern NOT covered by its module toolkit:
DO NOT skip it. Apply the Foundational Principle directly:
- Identify: What is this element? (3D tour, mortgage calculator, interactive map,
custom widget, embedded app, data visualization, etc.)
- Purpose: What is it supposed to do for the visitor?
- Evaluate through 5 layers:
- Layer 1: Is it for the right audience? Right tone?
- Layer 2: Does it contain the right content at the right depth?
- Layer 3: Is the format appropriate for the visitor's mode?
- Layer 4: Does the visitor feel in control? Is it overwhelming? Familiar enough?
- Layer 5: Is it predictable? Transparent? Gives feedback? Forgiving?
- Meta: Does it fit the journey? Promise = reality?
- If a problem is found — create a finding using the standard template.
Tag it as
[Unknown Element] so SYNTHESIZE can track novel findings.
Significance gate: Apply this protocol only to elements that are MEANINGFUL
for the visitor's goal on this page. Skip decorative/cosmetic elements.
A mortgage calculator on a property page = significant. A decorative animation = skip.
Principle Heartbeat Protocol
Agents WILL forget the principle mid-execution. This protocol prevents that.
Mandatory checkpoints — agents MUST pause and re-apply the principle at these EXACT moments:
Heartbeat 1 — BEFORE starting each page analysis:
[HB1] Page: {url} | Complexity: {simple/standard/complex}
Visitor: {who} | Needs: {what} | State: {emotion}
Visitor Mind (5 sec): {first impression as a first-time visitor — what do I think, feel, want to do?}
Entry paths: {how might someone arrive here — nav / search / direct link / social?}
This forces the agent to become the visitor BEFORE running any technical checks.
Heartbeat 2 — AFTER technical checks for a page, BEFORE writing findings:
[HB2] Principle review for {url}:
- Layer 1-3: content for the right audience? tone? depth? format? sequence?
- Layer 4: control? no overload? familiar patterns?
- Layer 5: predictable? transparent? feedback? forgivable? progress? time respected?
- Meta: promise=reality? effort≤value? trust proportional to ask? memorability?
- Design intent: is anything I marked FAIL potentially an intentional choice?
- Screenshot check: are all findings based on the normal page state?
- PASS on checklist but FAIL on principle? → add finding
- FAIL on checklist but not a real problem? → remove or lower severity
Expected override rate: ~15-20%. If 0 overrides — working mechanically.
Heartbeat 3 — AFTER every 5th finding written:
[HB3] Quality check after {N} findings:
- Are my findings specific ("move X after Y") or generic ("improve Z")?
- Do I still remember who the visitor is? (fresh eyes — not comparing to previous pages)
- Are recommendations varied or am I stuck in one type? (3+ identical → look for something different)
- Are all UX terms explained in plain language?
- Does every finding have a confidence tag ([Certain/Likely/Possible])?
- Assumptions tagged? ([Assumption: ...])
- Are there positive findings? (minimum 1 per every 5 negative)
Heartbeat 4 — BEFORE writing the Module Score section:
[HB4] Final review:
- Top finding: does it actually help the visitor, or is it a formal observation?
- What would the VISITOR notice that I haven't checked? ("What did I miss?")
- Copy-paste: same findings on 3+ pages → merge into a pattern?
- Depth check: if >25 findings → deepen the top 10 instead of generating new ones
- Score normalization: counting UNIQUE problems, not total instances
- Cross-module notes: noticed issues outside my domain? → record [Cross-module: {module}]
- Effort scaling: does every recommendation have a concrete effort level?
These heartbeats are NOT optional. They are MANDATORY steps in the execution flow.
The [HB1], [HB2], [HB3], [HB4] markers MUST appear in the module report.
If a report lacks heartbeat markers, it is considered incomplete.
Agent Behavior Standards
These standards prevent common AI agent failure modes during real audit execution.
Perception Standards
Fresh Eyes Reset:
Each page is evaluated INDEPENDENTLY. Do not compare to previously analyzed pages.
Not "this is worse than the homepage" — but "does this work for THIS page's visitor?"
Each HB1 heartbeat is a fresh start. Previous pages do not set the bar.
Design Intent Check:
Before flagging any visual/design element as FAIL, ask: "Could this be an intentional choice?"
- If possibly intentional: frame as "If intentional — acceptable. If not — recommendation: {X}"
- If clearly unintentional (broken layout, misaligned elements, clipped text): flag directly
- When in doubt: flag but tag as
[Possible design intent]
Screenshot Reliability:
Before writing a finding based on a screenshot, verify: "Does this screenshot show the normal
page state? Or could this be a loading state, error state, popup overlay, or transition frame?"
If uncertain: tag finding as [Screenshot state uncertain] and note the limitation.
Cultural Context:
When site lang=ru, apply Russian web conventions:
- Phone-first authentication is NORMAL (not unusual)
- VK, Telegram, OK.ru sharing buttons are STANDARD (not niche)
- Yandex SEO is as relevant as Google SEO
- Formal "Vy" (formal "you") is default (not cold/distant)
- Payment via SBP, MIR card, installments are STANDARD trust signals
Do NOT apply US/EU conventions blindly. When in doubt, default to Russian patterns.
State Your Assumptions:
Every finding that relies on an assumption MUST tag it explicitly:
[Assumption: desktop visitor]
[Assumption: first visit]
[Assumption: arrived from search]
[Assumption: page in final state]
This makes invisible assumptions visible and challengeable.
Quality Standards
Visitor Mind — 5 Seconds:
Before ANY technical checks on a page, look at the screenshot as a first-time visitor.
5 seconds. Ask yourself:
- "What do I think this page is about?"
- "What should I do here?"
- "Do I trust this?"
- "Am I confused by anything?"
This gut reaction is MORE VALUABLE than any technical check. It catches what checklists miss.
Write it as the first line of your page analysis.
Confidence Tagging:
Every finding gets a confidence level:
[Certain] — objectively verifiable from data (broken link, missing alt text, score < threshold)
[Likely] — strong reasoning from evidence (layout suggests confusion, copy is audience-mismatched)
[Possible] — informed judgment, could go either way (might be intentional design, might be oversight)
Do NOT present [Possible] findings with the same certainty as [Certain] ones.
Plain Language:
Every UX/design term MUST be immediately followed by a plain explanation:
- "visual hierarchy (what the eye sees first, second, third)"
- "cognitive load (how much effort it takes to understand the page)"
- "Fitts's Law (the larger and closer a button is, the easier it is to click)"
The site owner may be a dentist, a restaurateur, or a government official — not a designer.
Technical Fast Path:
Simple binary technical issues (broken link, missing alt text, 404 error, missing meta tag)
do NOT need full 5-layer principle analysis. Use short format:
- What: {issue}
- Where: {page:element}
- Fix: {action}
- Priority: {level}
The principle applies to SUBJECTIVE evaluations (is this CTA effective? is this copy persuasive?).
Binary pass/fail issues get the fast path.
Cognitive Load Budget:
Beyond individual element checks, assess the TOTAL cognitive demand per page:
- Count: distinct decisions required, competing visual elements, different typography styles,
moving/animated elements, notification badges, separate CTAs
- A page can have every individual element well-designed but TOGETHER overwhelm the visitor
- Reference: if total distinct interactive elements > 15 in one viewport, flag for review
- This is not a hard threshold — it is a lens. Some pages (dashboards) justify high density.
Error Prevention Assessment:
Beyond checking error MESSAGES (Conversion/Accessibility), evaluate error PREVENTION:
- Does the date picker prevent invalid dates?
- Does the phone field auto-format?
- Does the address field use autocomplete?
- Are destructive actions (delete, cancel) behind confirmation dialogs?
- Does the form save draft state (prevent loss from accidental navigation)?
Layer 5 (Forgiveness) covers REVERSAL. This covers PREVENTION. Both matter.
Visual Freshness Assessment:
Beyond checking content dates (SEO Content), evaluate VISUAL freshness:
- Do product screenshots show the current UI? (outdated screenshots = trust erosion)
- Do team/people photos look recent? (2015 corporate photo style ≠ 2025)
- Does the design itself feel current? (not "trendy" but "maintained")
- Are partner/client logos current branding?
Tag:
[Visual freshness concern] when the site LOOKS older than it probably IS.
Internationalization UX (beyond SEO hreflang):
If the site has a language selector:
- Does it use FLAGS? (wrong — flags = countries, not languages)
- Does it use native language names? ("Deutsch" not "German")
- Does switching preserve the current page?
- Is currency/measurement tied to language?
Output Standards
Depth Over Breadth After Threshold:
After 25 findings in a module: STOP generating new findings.
Instead: go back to your top 10 findings and DEEPEN them — add more specific recommendations,
more detailed reasoning, more concrete replacement content.
25 findings with 10 deeply detailed > 50 findings all medium depth.
Exception: FAIL-severity findings are always reported regardless of count.
Mandatory Positive Findings:
For every 5 negative findings, include at least 1 POSITIVE finding.
Positive findings must be SPECIFIC: "Navigation is intuitive: main sections are reachable
in 1 click, active state is clearly highlighted" — not "The site is generally fine."
If you analyzed 15 pages and found 0 positives — you are biased. Re-examine.
Positive Pattern Vocabulary:
When writing positive findings, use these patterns as reference for WHAT excellence looks like:
- Progressive Trust Architecture: Trust builds deliberately across the journey —
free content → case studies → transparent pricing → easy trial → generous refund
- Contextual Help at Decision Points: Tooltips and explanations appear exactly where
the visitor might hesitate — "What's the difference between plans?" on the pricing toggle
- Ethical Defaults: Opt-out for marketing, "essential only" cookie default,
prominent "No thanks" options, no pre-checked boxes
- Anticipatory Design: Smart defaults (city from IP, currency from locale),
recently viewed, "Based on your search" — saving visitor effort
- Content Density Optimization: Maximum useful information in minimum space
WITHOUT feeling cluttered — scannable, dense, efficient
- Graceful Empty States: "No results" that educates, encourages, and suggests
alternatives — not just a blank page
- Delightful Micro-Moments: Success animations, witty 404, progress celebrations,
copy that makes you smile — creates memorability
- Transparency as Feature: Public roadmap, uptime page, visible changelog,
published ranking methodology — proactive honesty
- Accessibility as Excellence: Beautiful focus states, curated screen reader
experience, power-user keyboard shortcuts, brand-preserving high-contrast mode
- Seamless Multi-Channel: QR codes for mobile continuation, "Email me this link",
clean shareable URLs, print-friendly versions
These are NOT mandatory checks. They are a VOCABULARY for recognizing excellence.
When an agent encounters these patterns, they should call them out as positive findings
with the same specificity required for negative findings.
Anti-Copy-Paste:
If the same problem appears on >3 pages — do NOT write separate findings for each.
Write ONE pattern finding:
"Pattern: {problem} found on {N} pages.
Example: {one specific page with the full detailed finding}.
Also affects: {list of remaining pages}."
This single pattern finding counts as ONE deduction in scoring, not N deductions.
Recommendation Diversity:
If 3 consecutive recommendations are the same type ("add whitespace" × 3):
PAUSE. Look for OTHER types of improvements. Variety of recommendations = more useful audit.
Effort Scaling in Recommendations:
Every recommendation specifies effort concretely:
- "Quick fix" (1 hour, anyone can do it — text change, block repositioning)
- "Design change" (requires a designer, 1-2 days — section redesign, new layout)
- "Development" (requires a developer, 3-5 days — new component, interactive feature)
- "Architecture" (requires a tech lead, 1-2 weeks — CMS migration, restructuring)
- "Strategic" (requires stakeholder decision — positioning, audience, business model)
Multi-Entry Awareness:
For key pages (homepage, pricing, product, case study), consider:
"How does this page work if visitor arrives from: (a) site navigation, (b) search engine,
(c) direct link shared by someone, (d) social media ad?"
Different entry points = different visitor states. Note when a page works for one path but fails for another.
System Integrity Standards
Score Normalization:
Scoring is by UNIQUE problems, not total instances.
The same problem on 20 pages = 1 pattern deduction, not 20 deductions.
This prevents score inflation on large sites.
Contradiction Detection (for SYNTHESIZE phase):
When aggregating module reports, explicitly check: do any recommendations CONTRADICT each other?
If Module A says "add more content" and Module B says "reduce page length" — this is a
TRADE-OFF finding, not two independent findings. Present as:
"Trade-off: {Module A} recommends {X}, {Module B} recommends {Y}. Compromise: {Z}."
Motivation Framing (for SITE-REVIEW.md):
The report is NOT a list of failures. It is a ROADMAP for improvement.
- Start with: "Current level: {score}. With quick wins → ~{score+15}. With strategic changes → ~{score+30}."
- "Quick Wins" section comes BEFORE detailed findings
- Frame as progress path, not failure list
- Include estimated impact per recommendation group
Complexity-Proportional Effort:
Not all pages deserve equal analysis depth:
- Complex pages (checkout, pricing, multi-step forms, dashboards): DEEP analysis, full principle application
- Standard pages (about, team, contact, blog post): STANDARD analysis
- Simple pages (legal, terms, privacy, 404): QUICK scan, technical checks only
Discovery tags each page with complexity level. Agents allocate effort proportionally.
Finding Dependency Graph (for SYNTHESIZE):
When aggregating findings, identify CAUSAL dependencies:
- "Fixing finding #3 (heavy images) ALSO resolves #7 (slow hero load) and improves #12 (high bounce)"
- Present as: "Root cause: {#3}. Resolves: {#7, #12}. Fix the root, get 3 improvements."
- This helps the report reader prioritize: fix ROOT CAUSES first, not symptoms.
Finding Actionability Gate:
Before writing ANY recommendation, verify: "Can a designer/developer execute this
in ONE working session without asking clarifying questions?"
- YES → recommendation is specific enough
- NO → add more detail until the answer is YES
Example of NOT actionable: "Improve the information architecture"
Example of actionable: "Add tabs to the pricing table: 'Key differences' (10 rows,
active by default) and 'Full comparison' (accordion by category)"
Assumption Challenge at HB4:
At HB4 (final review), add this check:
"If the visitor were actually {DIFFERENT from my HB1 assumption}, which findings would REVERSE?"
Example: if I assumed "business decision maker" but the visitor is actually "technical implementer"
→ findings about "too much jargon" would reverse, findings about "not enough API details" would appear.
Note any findings that are ASSUMPTION-DEPENDENT in the report.
Temporal Stability Tagging:
Tag findings by how STABLE they are over time:
[Stable] — will still be true next month (missing alt text, broken heading hierarchy, structural issues)
[Volatile] — may change soon (seasonal content, A/B test variant, dynamic pricing display, promotional offers)
Advise: "Fix [Stable] issues first — they will still be broken when you get to them."
Severity by Impact Radius:
Two identical issues have different impact based on WHERE they occur:
- Homepage (seen by 100% of visitors): full severity
- Depth-1 pages (seen by 50-70%): full severity
- Deep pages (seen by 5-10%): reduced severity consideration
- Utility pages (terms, privacy, sitemap): lowest impact
Note this in findings: "[Impact: high-traffic page]" or "[Impact: deep/utility page]"
Cross-Audit Learning (for repeat audits):
When comparing with a previous audit (diff mechanism):
- Same finding appears again → escalate priority: "Repeat audit: this finding has not been fixed.
Priority raised."
- Finding was fixed → acknowledge: "Fixed since last audit ✓"
- New finding on same page → flag: "New problem on a previously clean page — possible regression"
Contextual Intelligence Standards
These standards teach agents to READ THE CONTEXT of what they see — not just check elements mechanically.
Temporal Awareness:
All captured data is a snapshot in time. Before any finding about dynamic content:
- Assess data volatility:
high (prices, availability, search results change in minutes),
medium (news feed, promotions change daily), low (static content pages)
- For high-volatility content: tag finding with
[Temporal: prices/availability captured at {time}]
- Do NOT make definitive claims about dynamic data: "Price is too high" → instead:
"Price display at time of capture: {X}. Evaluate the PRESENTATION of pricing, not the price itself."
- If the page shows time-sensitive content (countdowns, "limited offer", flash sales):
note whether it is verifiable or potentially fake:
[Temporal: countdown detected, genuineness unverifiable from cache]
Regulatory Context Detection:
When the site operates in a regulated industry, agents must identify and account for it:
- Detect regulatory signals: license numbers in footer, certification badges, legal disclaimers,
industry-specific terminology (CBR, Roszdravnadzor, FSTEC, Roskomnadzor, SRO)
- If detected: record as
[Regulatory context: {industry}] in the report
- Apply stricter standards: mandatory disclosures become FAIL (not WARN) when missing
- Effort scaling adjustment: ANY change touching regulated content adds "+compliance review" time
- Do NOT recommend changes that would violate regulatory requirements
(e.g., don't suggest removing mandatory legal disclaimers to "clean up the design")
This is NOT about knowing every regulation. It IS about recognizing regulated context
and being conservative with recommendations in those areas.
YMYL Sensitivity:
When content can impact health, finances, safety, or legal rights:
- Detect YMYL signals: medical terms, financial calculations, legal advice, safety instructions,
"Your Money or Your Life" topics
- Apply MAXIMUM E-E-A-T scrutiny: author credentials MUST be visible, qualifications MUST be verifiable,
sources MUST be cited, claims MUST be evidence-based
- Score impact: YMYL pages with missing credentials = FAIL (not WARN)
- Note in report: "This site contains YMYL content. Stricter evaluation standards applied."
Every site may have some YMYL content (privacy policy, payment page). Full YMYL sites
(medical, financial, legal) get this treatment on EVERY content page.
Revenue Model Transparency:
Identify how the site makes money and check if the visitor can see this:
- Detect revenue model signals: affiliate links (utm, ref, partner params),
advertising blocks, subscription gates, commission disclosures, "powered by" labels
- Layer 5 (Transparency) applies: if the site earns from directing visitor actions
(affiliate commission, advertising clicks, partner referrals), is this DISCLOSED?
- "Comparison" or "best price" claims on affiliate sites: WARN if no methodology disclosure
- Advertising mixed with editorial content: WARN if not clearly labeled
- Record:
[Revenue model: affiliate/advertising/subscription/sales/freemium]
This is about HONESTY, not judgment. Affiliate sites are legitimate — but visitors deserve to know.
Variant Awareness:
The audit captures ONE state of a potentially multi-variant experience:
- Detect variant signals: personalization cookies, A/B test framework scripts
(Optimizely, VWO, Google Optimize, custom), session-based content differences
- If detected: record
[Variant: personalization/AB-test detected] in discovery.json
- Note prominently in SITE-REVIEW.md: "Audit was conducted for the default state
(anonymous first visit). Personalized/variant content was not evaluated."
- Do NOT make claims about "the site always shows X" — it may show different X to different visitors
- For A/B tests: if two versions of a page are encountered during crawl, note BOTH,
do not flag the variant as "inconsistency"
Data-Dense Page Pattern:
For pages primarily about comparison, search results, data tables, or catalog listings:
- Detect: page has a table >10 rows, or a results list >10 items, or filter controls
- Replace standard "Visitor Mind 5 seconds" with DATA PAGE variant:
"In 5 seconds: Can I identify the best option? Can I tell how results are ordered?
Do I understand the sort/filter controls? Is the key data (price, rating, date) scannable?"
- Evaluate COMPARISON USABILITY, not just visual layout:
- Can the visitor sort by the dimension they care about?
- Can the visitor filter to narrow results to their needs?
- Is the per-item information hierarchy clear (what's most important WITHIN each result)?
- Can the visitor compare 2-3 specific options side by side?
- Do NOT run per-item checks on every row/result — evaluate the PATTERN, not each instance
Compound Input Gate:
When a site requires MULTIPLE inputs before showing content (not just a single address/search field):
- Detect: page has 2+ required input fields AND no content below the form
AND the site's other pages reference results that require these inputs
- Examples: flight search (origin + destination + dates), real estate (location + budget + rooms),
hotel booking (city + dates + guests), job search (position + location)
- AskUserQuestion with ALL required fields:
"The site requires filling in multiple fields to display content:
{field1}: ___
{field2}: ___
{field3}: ___
Enter values for a complete audit, or skip (audit will be limited)."
- Fill ALL fields, submit, wait for results, then continue crawl from populated state
- Record:
"input_gate_type": "compound", "fields": ["origin", "destination", "dates"]
Diminishing Returns in Score Projections:
Higher baseline scores → smaller realistic improvement potential:
- Score 0-50: Quick wins may add ~15 points, strategic changes ~25 points
- Score 50-70: Quick wins ~10 points, strategic ~20 points
- Score 70-85: Quick wins ~5 points, strategic ~10 points
- Score 85+: "Further improvement requires A/B testing, user research,
and iterative optimization. Quick fixes yield minimal gain."
Use these bands in the Improvement Roadmap, not the flat "+15/+30" formula.
Common Layer 4-5 Violation Patterns:
These are FREQUENTLY OCCURRING violations that agents should actively look for
(not wait to stumble upon):
Layer 4 (Control) violations:
- Persistent app download banner blocking content on mobile
- Email/newsletter popup appearing within 5 seconds of page load
- Push notification permission request on first visit
- Scroll hijacking (custom scroll behavior that breaks native scrolling)
- "Subscribe to continue reading" gates on content pages
- Undismissable overlays
- Browser back button broken (SPA navigation issues)
- Confirmshaming dismiss copy ("No, I don't want to save money", "I prefer to stay uninformed")
- Roach Motel: signup takes 1 step, cancellation takes 5+ steps or requires phone call
- Misdirection: "Accept" button is large/colored, "Decline" is tiny/gray text
- Infinite scroll without progress indicators, back-to-top, or footer access
Layer 5 (Behavior) violations:
- Fake countdown timers (same timer on every visit = likely fake)
- "Only N left in stock" with no evidence of real inventory data
- Pre-checked opt-in checkboxes (newsletter, marketing, data sharing)
- Hidden fees revealed only at final checkout step
- Price displayed without taxes/fees that will be added later
- "Free trial" requiring credit card before any access
- Automatic subscription renewal buried in fine print
- Dark patterns in unsubscribe/cancellation flow
- Social proof manipulation: stock photos as "real customers", fabricated review counts,
"As seen in" logos without actual coverage, round numbers suggesting inflation ("10,000+ users")
- Bait-and-switch pricing: listing/ad price ≠ checkout price, "from $X" where X requires
unrealistic conditions, monthly rate displayed for annual-only plans
- Forced continuity: free trial → paid with no warning email, no easy cancellation
- Privacy Zuckering: "Accept All" prominent, "Manage Preferences" hidden, default = maximum sharing
- Confirmshaming: guilt-inducing decline copy on opt-in modals
- Urgency theater beyond countdowns: "X people viewing now" (unverifiable), "Selling fast!"
on every item, "Last chance!" on permanent offers
- Attention harvesting: notification badges that never reach zero, autoplay video sequences,
gamification creating obligation (expiring points, broken streaks)
When an agent encounters ANY of these: it is a finding. Tag with the specific layer violated.
Do NOT dismiss these as "standard business practice" — they are trust/respect violations
regardless of how common they are in the industry.
Incomplete Information Awareness:
The audit NEVER has complete information about a site. Always assume:
- Pages not crawled exist and may have different issues
- States not captured (loading, error, empty, authenticated, personalized) may reveal problems
- Devices not tested (tablet breakpoints, specific phones, slow connections) may break layouts
- Visitors not simulated (returning users, users with disabilities, users on corporate networks) may have different experiences
- Time not frozen (content changes, prices fluctuate, promotions expire, A/B tests rotate)
For EVERY finding, ask: "Would this finding change if I had information I don't have?"
- If YES and the missing info is critical → tag:
[Incomplete: {what's missing}]
- If NO → finding stands as-is
For the OVERALL report, explicitly state what the audit DID and DID NOT cover.
This is not a weakness — it is intellectual honesty that builds trust with the report reader.
Never claim "the site has no accessibility issues" — only "no accessibility issues were found
in the tested pages and states." The difference matters.
Content Type Awareness:
Not all pages contain the same TYPE of content. The evaluation criteria must match the content type:
| Content Type | How to Evaluate | Do NOT Apply |
|---|
| Editorial (blog, articles, guides) | Full copy quality, E-E-A-T, readability, originality | — |
| User-Generated (listings, reviews, Q&A, forum) | Evaluate the TEMPLATE and PRESENTATION, not the content itself. UGC quality varies by design. Do NOT flag listings as "thin content" or "duplicate" — they are unique by data, not by prose. | Copy quality, readability scoring, content originality per page |
| Transactional (checkout, cart, forms, calculators) | Evaluate flow, trust, clarity, error handling | Value proposition, social proof, SEO content depth |
| Navigational (category, search results, index) | Evaluate findability, filtering, sorting, result quality | Body copy, headline effectiveness, E-E-A-T |
| Institutional (about, legal, privacy, terms) | Evaluate completeness, accessibility, findability | CTA optimization, conversion, urgency |
| Platform/Marketplace (aggregate listings from multiple sellers) | Evaluate comparison UX, transparency, trust signals for PLATFORM (not individual sellers). VP = platform value (coverage, trust, convenience), not product VP. | Traditional value proposition, "generic hero" penalties |
| Interactive Tool (calculators, configurators, quizzes, estimators) | Evaluate: input clarity, output comprehensibility, edge case handling (empty/extreme values), shareability of results, mobile usability of controls | Word count, readability scoring, copy quality, SEO content depth |
| Documentation / Reference (API docs, knowledge bases, glossaries, wikis, manuals) | Evaluate: search within docs, code sample copy-ability, version selector, sidebar navigation, breadcrumb depth, anchor linking for section sharing | Value proposition, CTA optimization, social proof, urgency |
| Comparison / Decision-Support (vs pages, feature matrices, plan selectors) | Evaluate: are the RIGHT criteria compared? Are differences highlighted? Recommendation signals? Price normalization? Ease of adding/removing items? | Body copy readability, keyword density |
| Status / Live-Data (dashboards, tickers, status pages, weather, live scores) | Evaluate: data update visibility, last-updated timestamp, graceful stale-data handling, loading states, alert mechanisms. Tag: [Temporal: high volatility] | Copy quality, E-E-A-T, conversion optimization |
| Media Gallery / Detail (full-screen galleries, video players, virtual tours, 3D viewers) | Evaluate: navigation between items (keyboard, swipe), captions, download/share, zoom, fullscreen, preloading adjacent items | Word count, readability, heading hierarchy |
| Onboarding / Wizard (multi-step setup, tutorials, configuration flows) | Evaluate: progress visibility, back/skip ability, cognitive load per step, motivational copy between steps, completion celebration | Traditional page-level copy evaluation |
| Community / Social (forums, comments, discussions, Q&A, user profiles) | Evaluate: threading clarity, reply affordances, moderation signals, voting transparency, abuse reporting | Copy quality per post (evaluate TEMPLATE, not individual posts) |
| Map / Location-Centric (store locators, delivery zones, venue maps) | Evaluate: map load speed, pin clustering, list-vs-map toggle, filter by attributes, mobile touch, directions integration | Word count, body copy, readability |
Detect content type from: page URL pattern, page purpose (from discovery.json), content structure.
Apply the MATCHING evaluation criteria. Skip criteria from wrong content types.
This prevents: flagging UGC listings as "thin content," penalizing marketplaces for "generic value proposition,"
applying copy quality checks to legal documents, or demanding social proof on checkout pages.
FORBIDDEN
- Skipping Discovery phase
- Limiting page coverage arbitrarily ("max 6 pages", "key pages only") — crawl up to max_pages (default 100), prioritizing by navigation hierarchy
- Using
web_fetch or MCP tools for browsing — ALL live browser interaction via cmux browser commands (Bash)
- Looking at source code or suggesting code changes
- Using technical language ("change className", "add CSS", "modify component")
- Referencing external websites as examples (no browsing competitors)
- Generic advice ("make it better", "improve the design", "enhance UX")
- Truncating findings to save tokens — every finding gets the FULL template
- Using Sonnet/Haiku agents — ALL agents are Opus
- Skipping states (hover, empty, loading, error) — check them ALL
- Writing anything into the project source tree
Language Awareness
The system must adapt checks to the site's language (detected via <html lang=""> or content analysis).
For Russian-language sites (lang="ru"):
- Flesch-Kincaid: Do NOT use as a scored metric. Russian text naturally scores 4-6 grade levels higher than equivalent English. Use qualitative readability assessment only. SEO Content module should note: "FK is not applicable for Russian text. Readability assessment is qualitative."
- Power words: Use Russian equivalents: "free", "new", "proven", "guaranteed", "instantly", "exclusive", "limited", "now", "today", "finally", "discover", "learn", "get", "risk-free" (in Russian: "бесплатно", "новый", "доказанный", "гарантированно", "мгновенно", "эксклюзивно", "ограниченно", "сейчас", "сегодня", "наконец", "откройте", "узнайте", "получите", "без риска")
- You/We ratio: Check "vy/vash" vs "my/nash" (formal) or "ty/tvoy" vs "my/nash" (informal). Same >2:1 target.
- Cookie consent selectors: Add Russian button texts to the dismissal logic:
button:has-text("Принять"), button:has-text("Согласен"), button:has-text("Хорошо"), button:has-text("Принимаю"), button:has-text("OK")
Unified Scoring System
ALL modules use the SAME scoring formula. No per-module custom scoring.
Score Calculation
Each module starts at 100 and deducts:
- FAIL = −15 points (critical issue, blocks user or violates hard standard)
- WARN = −5 points (notable issue, degrades experience)
Minimum score: 0. No negative scores.
Severity Mapping
Modules MUST use only two severity levels in their checks: FAIL and WARN.
These map to the report priority system:
| Module Severity | Report Priority | Meaning |
|---|
| FAIL (first 2) | critical | Blocks primary user goal or violates hard standard |
| FAIL (rest) | high | Significant issue degrading experience |
| WARN | medium | Notable issue worth fixing |
| Positive note | low | Minor polish suggestion |
Score Anchoring
To ensure comparable scores across modules (same boundaries as Grade Assignment):
- 85-100 (A): Excellent. Up to 3 WARNs, zero FAILs.
- 70-84 (B): Good. 1 FAIL + up to 2 WARNs, or 4-6 WARNs with zero FAILs.
- 55-69 (C): Needs work. 2 FAILs, or 1 FAIL + 4-6 WARNs, or 7-9 WARNs.
- 40-54 (D): Poor. 3 FAILs, or 2 FAILs + multiple WARNs.
- 0-39 (F): Critical. 4+ FAILs.
Score Confidence
When a module cannot fully evaluate a dimension (cache limitations, auth-gated content,
dynamic content not captured, interactive elements not testable), the score must include
a confidence indicator:
Confidence levels:
- High (>80% checks executed): Score reported as-is: "Score: 75/100"
- Medium (50-80% checks executed): Score with range: "Score: 75/100 [Confidence: medium — 25% checks not executable from cache]"
- Low (<50% checks executed): Score with caveat: "Score: 75/100 [Confidence: low — >50% of evaluation was not possible. Score reflects only testable aspects.]"
How to calculate: Count total FAIL/WARN checks in the module toolkit. Count how many
were marked as "N/A" / "[Screenshot state uncertain]" / "not testable from cache" /
"[Unknown Element] — requires live testing". Ratio = confidence.
In SITE-REVIEW.md Score Dashboard: Each module shows confidence:
| Module | Score | Grade | Confidence | Note |
|---|
| UX Design | 80 | B | Medium (70%) | 3D/animation not evaluable from cache |
SYNTHESIZE uses confidence for weighted total: Low-confidence module scores are
de-emphasized in the total: multiply module weight by confidence percentage.
Example: UX Design weight 0.20, confidence 70% → effective weight 0.14.
This prevents artificially inflated scores from modules that couldn't fully evaluate.
Module Score in Report
Each module report MUST end with:
## Module Score
**Score: {N}/100** (Grade: {A|B|C|D|F})
Deductions:
- FAIL: {description} (-15)
- FAIL: {description} (-15)
- WARN: {description} (-5)
- ...
Total deductions: -{X}
Final: 100 - {X} = {N}
This makes scoring transparent and auditable.
Skill Library Integration
Before spawning audit agents, Lead reads relevant skills from ~/.claude/skill-library/:
Always loaded (every audit):
frontend/accessibility-design/SKILL.md → accessibility module
frontend/ui-ux-pro-max/SKILL.md → ux-design module
frontend/web-design-guidelines/SKILL.md → ux-design module
frontend/frontend-design/SKILL.md → ux-design module
marketing/page-cro/SKILL.md → conversion module
marketing/copywriting/SKILL.md → marketing-content module
marketing/seo-audit/SKILL.md → seo-technical + seo-content modules
marketing/marketing-psychology/SKILL.md → conversion + marketing-content modules
review-qa/performance-audit/SKILL.md → performance module
Loaded if applicable (based on business type):
marketing/form-cro/SKILL.md → if forms detected
marketing/signup-flow-cro/SKILL.md → if signup flow detected (SaaS)
marketing/onboarding-cro/SKILL.md → if onboarding detected (SaaS)
marketing/popup-cro/SKILL.md → if popups/modals with marketing intent detected
marketing/paywall-upgrade-cro/SKILL.md → if paywall detected (SaaS)
marketing/pricing-strategy/SKILL.md → if pricing page detected
marketing/schema-markup/SKILL.md → seo-technical module
marketing/content-strategy/SKILL.md → if blog/content section detected
Each audit agent receives the FULL content of its relevant skills as context in the spawn prompt.
Do NOT summarize skill contents — pass them verbatim.
Flow
/bishx:site [url]
│
▼
DISCOVER — cmux browser crawl (up to max_pages), map site, classify business, cache snapshots
│
▼
ASK — AskUserQuestion: scope confirmation (Full / Visual / SEO / Custom)
│
▼
EXECUTE — Wave 1: Tier A parallel (cache-based) → Wave 2: Tier B sequential (live browser)
│
▼
SYNTHESIZE — weighted scoring, diff with previous run, dedup findings
│
▼
REPORT — single SITE-REVIEW.md with detailed findings
│
▼
COMPLETE — cleanup, present summary
Phase 0: DISCOVER
Actor: Lead (main thread)
Goal: Complete site map, business classification, cache page data for audit agents.
Initialization
Run Directory
All artifacts go into .bishx-site/{YYYY-MM-DD_HH-MM}/.
Create the directory and subdirectories at start:
{run_dir}/
├── screenshots/
├── snapshots/
├── state.json
Pass the resolved path to all agents as {run_dir}.
.gitignore
Suggest to user: 'It is recommended to add .bishx-site/ to .gitignore — screenshots can take up hundreds of MB.'
Do NOT modify .gitignore directly — this would violate the FORBIDDEN rule against writing to project source tree.
Session Resume
Check if .bishx-site/active already exists:
- If yes → read the session dir and its
state.json
- If
active: true and phase != "complete":
- AskUserQuestion: "A previous audit session was found ({date}). Resume or start fresh?"
- Resume: continue from last phase
- Fresh: archive old session (rename dir to
{dir}-archived), create new
- If
active: false: create new session
- If no → create new session
Per-Phase Resume Guidance
If resuming mid-session:
- discover: Re-read discovery.json.pages to see how many pages were crawled. Continue crawling from where it stopped.
- ask: Check if selected_modules is non-empty in state.json. If yes — skip re-asking, proceed to EXECUTE. If empty — re-present AskUserQuestion and set waiting_for: "scope_selection".
- execute (wave 1): Check which report files exist in run_dir. Only spawn agents for modules whose reports are missing.
- execute (wave 2): Check which Tier B reports exist. Run remaining ones sequentially.
- synthesize: Re-run synthesis from scratch (idempotent).
- report: Re-read scores.json and all reports, regenerate SITE-REVIEW.md.
Initial State
Write .bishx-site/active with session name.
Write {run_dir}/state.json:
{
"skill": "site",
"active": true,
"phase": "discover",
"site_url": "",
"business_type": "",
"run_dir": ".bishx-site/{session}/",
"modules_total": 0,
"modules_completed": [],
"modules_failed": [],
"selected_modules": [],
"wave": 0,
"agent_pending": false,
"waiting_for": "",
"max_pages": 100,
"pages_crawled": 0,
"started_at": "ISO timestamp",
"updated_at": "ISO timestamp"
}
URL Resolution
- If user passed a URL argument → use it directly
- If no argument → detect from project:
- Check
docker-compose.yml / compose.yml for web service ports
- Check Vite/Next/Nuxt/Astro config for dev server URL + port
- Check
package.json scripts for dev command port
- Check
.env / .env.example for APP_URL, NEXT_PUBLIC_URL, etc.
- If cannot detect → AskUserQuestion for the URL
Store resolved URL as {site_url}.
cmux Verification
Read ~/.claude/skill-library/references/cmux-browser.md for the full cmux browser reference (commands, type vs fill, viewport, React forms, troubleshooting).
Open a browser surface:
RAW=$(cmux browser open {site_url})
SURFACE=$(echo "$RAW" | grep -o 'surface:[0-9]*' | head -1)
Store the surface ID in $SURFACE.
CRITICAL: Close the browser when done. When all crawling and testing is finished
(before writing the final SITE-REVIEW.md), close the browser:
cmux close-surface --surface $SURFACE
Never leave the browser open while writing reports or waiting for modules to complete.
If the command fails (cmux not installed) → STOP and tell user to install cmux.
All subsequent browser commands follow the pattern:
cmux browser --surface $SURFACE <subcommand> [args]
Full Site Crawl
This is the most critical phase. Thorough but bounded.
The goal is to discover and document every page, every interactive element, every state —
while staying within practical limits.
Crawl Limits
- Max pages: 100 (stored in
state.json.max_pages). If more routes are discovered,
prioritize: main nav pages first, then footer links, then blog/content, then deep links.
- Max depth: 5 levels from homepage
- Timeout per page: If a page takes >10s to load, skip with a note
- Skip: external links, same-page anchors (e.g., /about#team where path matches current page), mailto:, tel:, javascript:, PDF/image URLs
- Keep: hash-based SPA routes (e.g., /#/pricing, /#/about) — these are navigation routes, not page anchors
- Pagination: For paginated URLs (
?page=N, ?p=N, /page/N), crawl only page 1 and the last page (if detectable). Record pagination metadata in discovery.json: "pagination": {"pattern": "?page=N", "total_pages": 30}. Paginated URLs do NOT count toward max_pages limit. SEO Technical module checks: rel=next/prev, canonical handling, noindex on paginated pages.
If the site has more than max_pages routes, the agent must note which routes were skipped
and why in sitemap.md.
Step 1: Initial Crawl
RAW=$(cmux browser open {site_url})
SURFACE=$(echo "$RAW" | grep -o 'surface:[0-9]*' | head -1)
cmux browser --surface $SURFACE wait --load-state complete
Cookie Consent Dismissal
Before taking any snapshots, dismiss cookie/GDPR banners via snapshot + click:
cmux browser --surface $SURFACE snapshot -i
Look in the snapshot for cookie consent button refs (labels: "Accept", "Agree", "Принять", "Согласен", "OK", "Хорошо", "Принимаю").
If found, click the button:
cmux browser --surface $SURFACE click {ref_of_accept_button}
Wait 1 second after dismissal for animation to complete — simply proceed (cmux browser handles timing).
If no cookie banner found — proceed. This is best-effort, not a hard requirement.
Run this ONCE at the start of Discovery, not on every page.
Push Notification Prompt Dismissal
Handle browser-level notification prompts via snapshot + click.
For in-page notification opt-ins:
cmux browser --surface $SURFACE snapshot -i
Look in snapshot for dismiss buttons (labels: "Не сейчас", "Нет", "Позже", "No thanks").
If found: cmux browser --surface $SURFACE click {ref_of_dismiss_button}
Best-effort. Run ONCE after cookie consent dismissal.
Third-Party Widget Hiding
After cookie consent dismissal, note persistent floating widgets in the snapshot. Widgets can be hidden via JS eval:
- Take snapshot:
cmux browser --surface $SURFACE snapshot -i
- Identify floating widget refs (Intercom, Drift, Crisp, JivoSite, Telegram widget, LiveChat, etc.) in the snapshot
- Note them in discovery.json:
"floating_widgets": ["intercom", ...]
- To hide them for screenshots:
cmux browser --surface $SURFACE eval 'document.querySelectorAll(".intercom-launcher, #jivo-iframe-container, [id*="chat"]").forEach(el => el.style.display="none")'
- When analyzing page layouts, note if widget elements overlap content (document as a finding if so)
Record hidden widgets in discovery.json: "hidden_widgets": ["intercom", ...]
Note: Widgets are hidden for SCREENSHOT purposes only. The Accessibility module (Tier B) should test with widgets visible (they are part of the tab order).
Before Tier B modules run, the Lead should note: "Widgets were hidden during Discovery screenshots. Tier B agents work with the live site where widgets are visible."
Ad Banner Detection
Ad banners can be hidden via JS eval for cleaner screenshots:
- Take snapshot:
cmux browser --surface $SURFACE snapshot -i
- Identify ad-related elements in snapshot (classes/IDs containing: yandex_rtb, adfox, advert, ad-banner, adsbygoogle, doubleclick, googlesyndication, native-ad, sponsored)
- Note them in discovery.json:
"hidden_ads": ["yandex_rtb", ...]
- To hide for screenshots:
cmux browser --surface $SURFACE eval 'document.querySelectorAll("[id*=yandex_rtb], .adsbygoogle, [class*=ad-banner]").forEach(el => el.style.display="none")'
- Performance module should measure ads as they appear — ads impact real performance.
Add note: "Ad presence documented from snapshot. Performance module tests the live site with ads active."
CallTracking Detection
Check for common call tracking scripts via HTML source or eval:
curl -s "{site_url}" | grep -oiE '(calltouch|comagic|roistat|callibri|mango-office|calltracking|ringostat)'
Or via eval: cmux browser --surface $SURFACE eval 'document.body.innerHTML.match(/(calltouch|comagic|roistat|ringostat)/i)?.[0] || "none"'
Record in discovery.json: "calltracking_detected": true/false, "calltracking_service": "calltouch"
When calltracking is detected, add to discovery.json: "calltracking_note": "Phone numbers on this site are dynamically replaced by CallTracking service. NAP phone consistency checks should be suppressed — different numbers per page are intentional."
Geo-Detection Check
Check if the site shows geo-dependent content via snapshot:
cmux browser --surface $SURFACE snapshot -i
Look in snapshot for city selection elements (labels/text: "Выберите город", "Ваш город", "Select city") or geo-related button refs.
If found: note detected=true and current city text in discovery.json.
Record in discovery.json: "geo_detection": {"detected": true, "current_city": "Moscow"}
Add to ALL module reports: "⚠️ The site uses geo-detection. Audit was conducted for city: {city}. Content and results may differ for other regions."
Input-Gate Detection
After initial page load, check if the site requires user input to show content:
Use snapshot to detect input gates:
cmux browser --surface $SURFACE snapshot -i
In the snapshot, check:
- Are there fewer than 10 internal links visible?
- Are there form input elements (text, search, date, select, combobox) prominently displayed?
If yes → site is likely input-gated.
If input-gated:
- Record in discovery.json:
"input_gated": true, "gate_type": "single|compound", "gate_fields": [{placeholder, type}]
- If SINGLE field (1 input): AskUserQuestion with one value request
- If COMPOUND gate (2+ inputs — travel search, real estate, etc.):
AskUserQuestion listing ALL required fields:
"The site requires filling in multiple fields to display content:
{field1.placeholder}: ___
{field2.placeholder}: ___
{field3.placeholder}: ___
Enter values for a complete audit, or skip."
- If user provides values: fill ALL inputs via
cmux browser --surface $SURFACE fill {ref} '{text}', submit via cmux browser --surface $SURFACE click {submit_ref}, wait for load, continue crawl from populated state
- If user skips: proceed with limited crawl, note in SITE-REVIEW.md: "⚠️ Content behind the input form ({N} fields) was not evaluated."
cmux browser --surface $SURFACE screenshot --out {run_dir}/screenshots/homepage-desktop.png
cmux browser --surface $SURFACE snapshot -i
Extract ALL links via snapshot analysis:
From the snapshot, collect all link elements with their href, text, and context (nav/footer position). Build a link queue. Add every unique internal path.
SPA Route Discovery
After extracting <a href> links from the snapshot, also discover SPA routes:
- From snapshot output, look for elements with
data-href, data-to, data-path attributes containing path values
- Check HTML source for Next.js prefetch links:
curl -s "{site_url}" | grep -oE 'href="(/[^"]*)"' | sort -u
Add any discovered SPA routes to the crawl queue. Note in sitemap.md which routes were discovered via SPA detection vs. <a> tags.
This is best-effort — some SPA routes are only discoverable by clicking navigation elements (covered by Step 3 interactive testing).
Animation Framework Detection
Check via HTML source and snapshot:
curl -s "{site_url}" | grep -oiE '(gsap|three\.js|lottie|anime\.js)'
Also check snapshot for canvas elements and [class*="lottie"] elements.
Record in discovery.json: "animation_frameworks": {"gsap": true, "threejs": false, ...}
If WebGL/Three.js detected, add to ALL Tier A reports: "⚠️ The site uses WebGL/Canvas. Screenshots may not accurately represent 3D content."
Step 1.5: Fetch Site Infrastructure
curl -s "{site_url}/robots.txt" > {run_dir}/robots.txt
curl -s "{site_url}/sitemap.xml" > {run_dir}/sitemap.xml
Add paths to discovery.json: robots_txt_path and sitemap_xml_path.
If sitemap.xml contains additional URLs not found via link crawling, add them to the link queue.
Step 2: Recursive Page Discovery
For EACH unique route in the queue (up to max_pages):
cmux browser --surface $SURFACE goto {url}
cmux browser --surface $SURFACE wait --load-state complete
cmux browser --surface $SURFACE snapshot -i
cmux browser --surface $SURFACE screenshot --out {run_dir}/screenshots/{page_slug}-desktop.png
Note: Scroll capture is NOT done here. It happens in Step 4 (Post-Interaction Scroll Capture) AFTER interactive testing has revealed hidden content (expanded accordions, loaded "load more" content, etc.).
cmux resize-pane --pane $(cmux list-panes | grep -o 'pane:[0-9]*' | tail -1) -L --amount 400
cmux browser --surface $SURFACE screenshot --out {run_dir}/screenshots/{page_slug}-mobile.png
cmux resize-pane --pane $(cmux list-panes | grep -o 'pane:[0-9]*' | tail -1) -R --amount 200
cmux browser --surface $SURFACE screenshot --out {run_dir}/screenshots/{page_slug}-tablet.png
cmux resize-pane --pane $(cmux list-panes | grep -o 'pane:[0-9]*' | tail -1) -R --amount 200
Tablet scroll capture is NOT needed (desktop scroll captures cover the content). Tablet screenshot is for responsive layout evaluation only.
Also extract and save page metadata. Wait for SPA hydration using cmux browser wait:
cmux browser --surface $SURFACE wait --load-state complete
Then extract metadata via snapshot and curl:
cmux browser --surface $SURFACE snapshot -i
For meta tags not visible in snapshot, use curl:
curl -s "{url}" | grep -E '<title>|<meta name="description"|<link rel="canonical"|<meta property="og:|<script type="application/ld\+json"|<html lang'
Extract from HTML source:
title: from <title> tag
h1: from <h1> tag text
metaDesc: from <meta name="description" content="..."
canonical: from <link rel="canonical" href="..."
ogTitle, ogDesc, ogImage: from <meta property="og:*"
jsonLd: from <script type="application/ld+json"> blocks
lang: from <html lang="..."
viewport: from <meta name="viewport" content="..."
wordCount: estimate from snapshot text content length
Save metadata per page into discovery.json.pages[].
For EACH page, after extracting metadata, determine and record:
purpose: one sentence — what is this page FOR? (e.g., "Help visitor choose a pricing plan", "Show project case study results", "Collect visitor contact information")
visitor: who most likely looks at this page? (e.g., "Decision maker comparing options", "Developer looking for API docs", "Citizen filing an appeal")
key_question: what question does this visitor have? (e.g., "Which plan fits my needs?", "How do I authenticate API calls?", "What documents do I need?")
Also determine page complexity for effort allocation:
complexity: "complex" (checkout, pricing, multi-step forms, dashboards, configurators) | "standard" (features, about, blog, services, category) | "simple" (legal, terms, privacy, 404, sitemap)
Derive from: number of interactive elements, form count, content depth, page purpose.
Agents allocate analysis depth proportionally: complex = full principle application, standard = standard analysis, simple = quick technical scan.
Derive these from: page URL, H1 text, content structure, position in navigation.
Keep each field to ONE sentence. This is a quick assessment, not deep analysis.
From each page:
- Extract all new internal links → add to queue
- Record: URL, page title, H1, meta description, navigation position
Update state.json.pages_crawled after each page.
Continue until queue is empty or max_pages reached.
Auth-Gated Content Detection
During crawl, detect pages that redirect to login:
- If
cmux browser --surface $SURFACE goto {url} results in loading a login/register page (check snapshot for login form), record the original URL as auth-gated
- Record in discovery.json:
"auth_gated_pages": ["/dashboard", "/account", "/courses/viewer"]
- In SITE-REVIEW.md, include a prominent section: "Pages behind authentication (not evaluated): {list}. Credentials are required to audit these pages."
- Do NOT attempt to log in. Do NOT ask user for credentials. Simply document what was found and what was not auditable.
Infinite Scroll Detection
After initial page load, check for infinite scroll via snapshot observation:
- Take initial snapshot:
cmux browser --surface $SURFACE snapshot -i
- Scroll down to check for more content:
cmux browser --surface $SURFACE scroll --dy 2000
- Note if the snapshot reveals a feed/list structure with
load more triggers or dynamic loading patterns
- Infinite scroll heuristic: if page contains a list/feed structure with no visible "last page" pagination, assume infinite scroll possible
If infinite scroll detected:
- Do NOT scroll to "bottom" — there is no bottom
- Capture max 5 scroll screenshots (standard cap applies)
- Record in discovery.json per page:
"infinite_scroll": true
- Note in sitemap.md: "⚠️ Infinite scroll detected on {page}. Content beyond 5 viewports not captured."
Step 2.5: Template Grouping
After crawling 10+ pages, detect template patterns from snapshot structure:
- Compare snapshot structural output across pages — look for repeated heading patterns, element hierarchies, and URL patterns (e.g.,
/product/123, /product/456 → same template)
- Group pages by URL pattern and structural similarity
Group pages with identical structural signatures into template groups.
For each template group with >3 pages:
- Mark 3-5 representative pages for full analysis (diverse content: shortest, longest, most-linked)
- Mark remaining pages as
"template_sampled": true in discovery.json
- All modules should analyze the template ONCE via representatives, then note: "Applies to {N} pages using this template"
Add to discovery.json schema:
"template_groups": [
{
"template_id": "product-page",
"signature": "...",
"page_count": 600,
"representative_pages": ["/product/123", "/product/456", "/product/789"],
"all_pages": ["/product/123", "/product/456", ...]
}
],
Per-page schema addition: "template_group": "product-page" or null if unique.
Modules MUST respect template grouping: analyze representative pages in full, note "This finding applies to all {N} pages using the '{template_id}' template" in findings.
Population Estimation from Sitemap
If {run_dir}/sitemap.xml was fetched and contains URLs:
For each template group, extract the URL pattern from representative pages and count matches in sitemap.xml.
Update discovery.json: template_groups[].estimated_population (from sitemap) vs template_groups[].page_count (actually crawled).
All module reports should use estimated_population when stating "This finding affects ~{N} pages."
Uncrawled Page Type Alert
After crawl completes (max_pages reached), check sitemap.xml for URL patterns NOT represented in crawled pages:
-
Extract all unique URL path patterns from sitemap (e.g., /agents/, /reviews/, /api/docs/*)
-
Compare with crawled page URL patterns
-
For each UNCRAWLED pattern with >10 URLs in sitemap:
- Record in discovery.json:
"uncrawled_types": [{"pattern": "/agents/*", "estimated_count": 5000, "sample_urls": ["...", "..."]}]
- Note in sitemap.md: "⚠️ Not included in sample: /agents/* (~5000 pages), /reviews/* (~3000 pages)"
-
In SITE-REVIEW.md, add a section BEFORE findings:
## Coverage Limitations
Found {total_sitemap_urls} URLs in sitemap.xml. Evaluated: {crawled_count} ({percentage}%).
Page types not included in the sample:
| Pattern | Count | Reason | Recommendation |
|---------|-------|--------|----------------|
| /agents/* | ~5000 | Exceeded max_pages | Run a separate audit of the agent profile template |
| /reviews/* | ~3000 | Not prioritized | Include 3-5 in the next audit |
Findings below apply ONLY to the evaluated pages.
This is NOT a finding — it is a transparency disclosure. The audit honestly reports what it DID and DID NOT cover.
Step 3: Interactive Element Discovery
For sites with more than 20 pages, limit interactive element testing to 15 pages selected by these criteria (in priority order):
- Homepage (always included)
- Pages with forms (contact, signup, login, checkout)
- Pages with pricing or plans
- Depth-1 navigation pages (direct children of main nav)
- If slots remain: pages with the most incoming internal links (from discovery data)
Document which pages were tested interactively in discovery.json.interactive_pages array.
Motion/hover capture (Step 3.5) uses a subset of these — the top 10 by the same criteria.
For EACH page (within the interactive testing scope), interact with EVERY element:
Buttons and clickables:
cmux browser --surface $SURFACE click {ref} every button that opens modals, dropdowns, menus, sidebars
cmux browser --surface $SURFACE snapshot -i to capture each opened state
- Click close/X button or take snapshot to verify it closes
Forms:
- Find every form on the site via snapshot
- Extract form structure from snapshot: identify all input, textarea, select elements with their labels, placeholders, required state, and submit button text
- Store extracted forms in
discovery.json.pages[].forms
- Document: fields, labels, placeholders, required markers, submit button text
- Check validation states: click submit without filling → snapshot for error states
- Check field types visible in snapshot
Masked Content Detection
Look for buttons that reveal hidden content (common on real estate, marketplaces):
From the snapshot, find button elements with labels matching reveal patterns:
- Russian: "показать телефон", "показать номер", "показать email", "показать контакт", "раскрыть", "развернуть", "показать полностью", "читать далее"
- English: "show phone", "show number", "show contact", "read more"
For each detected masked content button (on interactive testing pages):
- Click the button:
cmux browser --surface $SURFACE click {ref}
- Take snapshot of revealed state:
cmux browser --surface $SURFACE snapshot -i
- Record revealed content in discovery.json per page:
"revealed_content": [{"trigger": "Show phone", "content": "+7 999 123-45-67"}]
Interactive Controls Detection (beyond )
Detect interactive filter/calculator widgets that are NOT wrapped in tags:
From the snapshot, identify:
- Range sliders: elements with
role="slider" or type="range"
- Select dropdowns outside forms:
role="listbox" elements
- Checkbox/radio groups:
role="group", role="radiogroup"
Record in discovery.json per page: "interactive_controls": [{type, label, page_section}]
Note: These controls cannot be tested by Tier A modules. Accessibility (Tier B) should test their keyboard navigability.
Cart Population (e-commerce/marketplace/food delivery)
If business_type in [ecommerce, marketplace, food_delivery]:
- Find one product/item page with "Add to Cart" / "Add to Cart" / "Add" button
- Click it, wait 2 seconds
- Navigate to /cart → screenshot + snapshot (populated state)
- If cart populated, navigate to /checkout → screenshot + snapshot
- Record populated states in discovery.json:
"cart_populated": true
This gives modules populated checkout views instead of empty-cart states.
Navigation:
- Mobile menu:
cmux resize-pane --pane $(cmux list-panes | grep -o 'pane:[0-9]*' | tail -1) -L --amount 400 # narrow for mobile → find hamburger menu ref in snapshot → cmux browser --surface $SURFACE click {hamburger_ref}
- Dropdown menus: click each nav item with submenus via snapshot refs
- Tab bars, sidebars, breadcrumbs
cmux resize-pane --pane $(cmux list-panes | grep -o 'pane:[0-9]*' | tail -1) -R --amount 200 # restore desktop → restore desktop
Dynamic content:
- Scroll to bottom of each page → check for lazy-loaded content
- Click "load more" / pagination buttons
- Check scroll-triggered animations
States:
- Empty states (if data-dependent pages, note what they show with no data)
- Loading states (observe during navigation transitions via snapshot)
- Error states (navigate to non-existent routes → 404 page)
- Hover states: use
cmux browser --surface $SURFACE hover {ref} to trigger hover effects, then take screenshot
Step 3.5: Motion & Hover Capture Pass
On up to 10 key pages (homepage + main nav pages + primary conversion pages):
Hover state capture:
cmux browser supports hover directly:
cmux browser --surface $SURFACE hover {element_ref}
cmux browser --surface $SURFACE screenshot --out {run_dir}/screenshots/{page_slug}-hover-{element_desc}.png
Use snapshot refs (e1, e2, ...) from snapshot -i output rather than CSS selectors with quotes to avoid quoting issues.
Page transition observation:
Navigate between 3-5 pages in sequence. After each navigation, note:
- Did the page load instantly or with a visible transition?
- Was there a loading indicator?
- Screenshot the loading/transition state if visible.
Save transition observations to {run_dir}/motion-notes.md:
# Motion & Transition Notes
## Page Transitions
| From | To | Behavior | Loading Indicator |
|------|----|----------|-------------------|
## Hover States Captured
| Page | Element | Screenshot | Hover Effect |
|------|---------|------------|-------------|
## Scroll Animations
| Page | Element | Trigger | Type |
|------|---------|---------|------|
Scroll animation detection:
Trigger scroll animations using:
cmux browser --surface $SURFACE scroll --dy 500
cmux browser --surface $SURFACE screenshot --out {run_dir}/screenshots/{page_slug}-scroll-{N}.png
Repeat for progressive scroll positions. JS-based scroll is also available via eval:
cmux browser --surface $SURFACE eval 'window.scrollTo(0, 800)'
Note any scroll-triggered animations observed in screenshots in motion-notes.md.
This data feeds UX Design Dimension 6 (Motion & Micro-interactions).
Tier A modules can read {run_dir}/motion-notes.md and hover screenshots from cache.
Step 3.6: Theme/Dark Mode Detection
Check for theme toggle via snapshot:
cmux browser --surface $SURFACE snapshot -i
Look in snapshot for theme toggle buttons (labels/aria: "dark mode", "dark", "theme", "тёмн", "тем").
If theme toggle found:
- Click the toggle:
cmux browser --surface $SURFACE click {toggle_ref}
- Screenshot for up to 10 key pages (homepage + depth-1 nav pages):
cmux browser --surface $SURFACE screenshot --out {run_dir}/screenshots/{page_slug}-desktop-dark.png
- Click toggle again to restore original theme
- Record in discovery.json:
"dark_mode": {"detected": true, "pages_captured": [...]}
Tier A modules (UX Design, Brand Consistency, Accessibility) should analyze BOTH theme screenshots when available.
If no toggle found: "dark_mode": {"detected": false}. Skip dual capture.
Visually Impaired Version Detection (GOST R 52872-2019)
Check for "version for visually impaired" toggle (Russian government/institutional sites) via snapshot:
cmux browser --surface $SURFACE snapshot -i
Look in snapshot for BVI toggle buttons (labels: "Версия для слабовидящих", "Для слабовидящих").
If found:
- Click the toggle:
cmux browser --surface $SURFACE click {bvi_toggle_ref}
- Screenshot for up to 5 key pages and save as
{page_slug}-desktop-bvi.png
- Restore normal version
- Record:
"visually_impaired_version": {"detected": true}
- Accessibility module evaluates BOTH versions
Step 4: Post-Interaction Scroll Capture
After interactive testing (Step 3) has expanded accordions, clicked "load more" buttons, and revealed hidden content, NOW capture full-page scroll screenshots. This ensures expanded states and dynamically loaded content appear in scroll captures.
For EACH page (desktop only):
Capture scroll screenshots using the scroll command and screenshot:
cmux browser --surface $SURFACE screenshot --out {run_dir}/screenshots/{page_slug}-desktop-scroll-1.png
cmux browser --surface $SURFACE scroll --dy 800
cmux browser --surface $SURFACE screenshot --out {run_dir}/screenshots/{page_slug}-desktop-scroll-2.png
cmux browser --surface $SURFACE snapshot -i
This gives Tier A modules the accessibility tree of the ENTIRE page (snapshot contains all DOM content, not just above-fold).
Do NOT attempt scroll capture on mobile viewport.