design and review on-call rotations and escalation including coverage and follow-the-sun arrangements, primary and secondary tiers with acknowledgement and response expectations, shift handoff and transfer of open state, page load budget and out-of-hours…
Skills in this repository
MadewellRD/skills-lab - Page 10
SkillsMP has collected 1,540 skills from MadewellRD/skills-lab. Open a skill to review its source and details.
MadewellRD/skills-labShowing 40 of 1,540 collected skills.
write and facilitate blameless postmortems and incident reviews with a timeline anchored to timestamped evidence, contributing factors across technical detection and process dimensions, counterfactual discipline that resists the single root cause story, a…
run a production readiness review or support acceptance review with a gate set scored pass, waived, failed, or not assessed against named evidence, a launch or acceptance decision with conditions, waivers carrying a named owner and an expiry date, hand-back…
run the recurring service reliability review covering error budget attainment and burn across the period, error budget policy adjudication including whether a freeze applies and its exception path, incident page and toil trends, open postmortem actions and…
design timeout and retry budgets, circuit breakers, bulkheads, load shedding and admission control, backpressure, cache and stale-serve fallbacks, idempotency and replay safety, and graceful degradation modes, and record whether each control is configured,…
write and review operational runbooks keyed to the alert or failure mode that triggers them, with the first mitigating action stated before diagnosis, diagnostic decision trees carrying exact queries and dashboards, escalation and rollback branches, access…
map critical user journeys to the services on their path, assign criticality tiers using an explicit tiering rule, attribute ownership and pager rotation per service, and name every service with no owner, no tier, or a rotation that resolves to nobody. use…
specify service level indicators per critical user journey including the good-event and valid-event definitions, the measurement point and its bias, the implementation query or export behind each indicator, and the split between measured, partially measured,…
set service level objectives with windows, define and compute the error budget and its balance, account for burn rate over multiple windows, write an error budget policy with consequences that bind someone, and separate objectives agreed with the owner from…
orchestrate site reliability engineering work across slis, slos, error budgets, capacity planning, load and performance testing, chaos and resilience testing, dependency and failure-mode analysis, disaster recovery, backup and restore, alerting quality,…
account for operational toil with measured hours and how they were established, classify each recurring task as automatable partially automatable or inherent, define the elimination path per task through automation self-service or a design change that removes…
define accessibility and seo requirements for web surfaces including wcag, semantic html, keyboard navigation, focus, screen-reader behavior, metadata, structured data, canonicals, sitemaps, robots, crawlability, localization, and page-type search…
plan web backend integration across api contracts, auth, sessions, bff layers, cms connections, data models, caching, rate limits, pagination, failure modes, and ownership. use when a web surface depends on services, databases, third-party systems, identity,…
define cms and content operations for web properties including structured content models, editorial workflow, publishing rules, approvals, localization, migration, governance, content debt, metadata ownership, and day-two content maintenance. use for…
create implementation-ready frontend engineering plans for web surfaces including rendering strategy, routing, layouts, components, state, forms, data fetching, framework constraints, accessibility hooks, performance controls, and coding-agent handoffs. use…
design web information architecture including sitemap, route hierarchy, navigation, url taxonomy, content relationships, wayfinding, content hierarchy, and findability. use when the assistant needs to structure websites, portals, dashboards, docs sites,…
create web-specific product requirements, page and route scope, user journeys, acceptance criteria, success metrics, analytics intent, source facts, risks, open questions, and downstream handoff notes for websites, web apps, landing pages, portals,…
define ux/ui design system guidance for web surfaces including component inventory, responsive behavior, design tokens, interaction states, accessibility-aware patterns, brand consistency, and design governance. use before frontend implementation or…
orchestrate complete web development workflows across product, information architecture, ux/ui design systems, frontend engineering, backend integration, cms/content operations, web security/secops, performance, accessibility, seo, testing, observability,…
coordinate post-launch web maintenance and growth including analytics-informed backlog, experiments, conversion optimization, content refresh, seo iteration, accessibility remediation, performance regression follow-up, dependency upgrades, refactors,…
design production web observability including rum, synthetic checks, availability, frontend errors, api errors, latency, core web vitals, analytics events, dashboards, alerts, ownership, incident hooks, launch monitoring, and post-launch review. use before…
plan and gate web performance using core web vitals, performance budgets, rendering cost, hydration, bundle size, images, fonts, scripts, cdn, caching, data fetching, latency, field and lab measurement, launch thresholds, and regression controls. use before…
plan web release and deployment across ci/cd, preview environments, build commands, environment promotion, hosting, cdn, edge config, cache invalidation, feature flags, launch checklist, rollback, post-release validation, and release-to-observability handoff.…
review and plan web security and secops controls across auth, sessions, browser storage, cookies, csp, security headers, cors, csrf, dependency risk, secrets, third-party scripts, cdn/edge hardening, abuse prevention, monitoring hooks, and incident readiness.…
define web testing and qa strategy across browsers, devices, responsive layouts, visual regression, forms, auth flows, integrations, accessibility checks, seo checks, performance checks, smoke tests, regression tests, release signoff, and defect triage. use…
design AI agent architecture including planning boundaries, execution loops, memory and state strategy, tool routing, approval gates, retries, delegation, and halt behavior.
design observability for AI agents and workflows including traces, prompts, model calls, tool calls, retrieval events, approvals, errors, eval probes, cost, latency, and safety signals.
orchestrate AI engineering workflows from capability intent through model, prompt, tool, agent, retrieval, eval, safety, inference, observability, release, and incident stages using connector-grounded evidence, workflow packets, stage advancement, and halt…
triage AI production incidents involving hallucination spikes, safety failures, prompt injection, tool misuse, data leakage, model regressions, cost spikes, latency degradation, eval regressions, or user harm reports.
assess readiness to release AI capabilities across requirements, evals, safety review, red-team status, inference ops, observability, rollback, docs, support handoff, and owner approval.
review AI capability risks including misuse, policy compliance, privacy, security, hallucination harm, data leakage, autonomy, tool-use risk, user impact, and mitigations.
optimize AI system cost and latency using model routing, caching, prompt compression, context pruning, batching, streaming, parallelism, retrieval tuning, and fallback tiers while preserving quality and safety gates.
plan and review AI datasets for source selection, labeling, balancing, privacy, deduplication, train and eval splits, drift, provenance, consent, and retention.
design AI evaluation plans with goals, datasets, rubrics, grading methods, thresholds, regression slices, safety checks, human review, and reporting requirements.
analyze completed AI eval runs, regression deltas, failure clusters, grading reliability, threshold status, release blockers, and rerun recommendations.
assess and plan fine tuning only when prompt, retrieval, tool, model routing, and eval evidence justify training a specialized model.
plan production inference operations including deployment topology, rate limits, quotas, retries, caching, streaming, fallbacks, batching, timeouts, secrets, logging, and SLOs.
select model candidates, routing constraints, fallback behavior, and model tradeoffs for AI capabilities using task fit, quality, latency, cost, safety, modality, context, and deployment evidence.
design prompt systems, instruction hierarchy, context assembly, prompt contracts, refusal and defer behavior, prompt evaluation fixtures, prompt injection defenses, and prompt observability hooks for AI capabilities.
plan and analyze adversarial AI testing for jailbreaks, prompt injection, data exfiltration, harmful instructions, over-permissioned tools, and policy evasion.