run vulnerability management across consolidated and deduplicated findings from every scanner, risk-based prioritization with each severity carrying its scale, asset criticality and exposure weighting, remediation ownership and service level assignment, patch…
Skills in this repository
MadewellRD/skills-lab - Page 29
SkillsMP has collected 1,540 skills from MadewellRD/skills-lab. Open a skill to review its source and details.
MadewellRD/skills-labShowing 40 of 1,540 collected skills.
design and review production alerting including symptom-based paging on user impact, multi-window multi-burn-rate rules tied to error budget spend, page versus ticket versus dashboard routing, deduplication grouping inhibition and dependency-aware…
check backup coverage against the data inventory, record mechanism schedule retention and immutability with their sources, write the restore procedure and its measured time from a dated drill, verify integrity and deletion resistance including ransomware and…
build the demand model and its drivers, measure headroom against the binding saturation signal rather than average cpu, find quota connection partition thread and licence ceilings, state provisioning lead time, compute failover headroom at real peak, and…
make production change safe through rollout strategy and staged exposure, canary analysis signals and promotion thresholds, bake time per stage, rollback triggers and whether rollback has ever actually succeeded, freeze policy and its exception path, schema…
design chaos experiments with an explicit steady-state hypothesis, choose the fault and its scope, contain blast radius with abort criteria written before injection, run game days that exercise responders and runbooks, and promote controls that fail the…
map the dependency graph along each critical user journey with hard, soft, and degraded-ok coupling, find single points of failure and shared-fate risk across zones clusters datastores identity dns and control planes, analyze failure modes with trigger…
set rto and rpo per tier with their sources, state the failover mode actually implemented rather than the one designed, define regional evacuation and failback, derive the dependency recovery order across identity dns data tier and control plane, run and…
run a live production incident with severity declaration, incident commander and operations communications and scribe roles, evidence capture before any mitigating restart, reversible mitigation preferred over diagnosis, internal and customer communication…
derive a workload model from production traffic, design load stress soak spike and breakpoint profiles with what each proves, state environment fidelity gaps and what they invalidate, measure the saturation point and the behavior past it, and define the…
design and review on-call rotations and escalation including coverage and follow-the-sun arrangements, primary and secondary tiers with acknowledgement and response expectations, shift handoff and transfer of open state, page load budget and out-of-hours…
write and facilitate blameless postmortems and incident reviews with a timeline anchored to timestamped evidence, contributing factors across technical detection and process dimensions, counterfactual discipline that resists the single root cause story, a…
run a production readiness review or support acceptance review with a gate set scored pass, waived, failed, or not assessed against named evidence, a launch or acceptance decision with conditions, waivers carrying a named owner and an expiry date, hand-back…
run the recurring service reliability review covering error budget attainment and burn across the period, error budget policy adjudication including whether a freeze applies and its exception path, incident page and toil trends, open postmortem actions and…
design timeout and retry budgets, circuit breakers, bulkheads, load shedding and admission control, backpressure, cache and stale-serve fallbacks, idempotency and replay safety, and graceful degradation modes, and record whether each control is configured,…
write and review operational runbooks keyed to the alert or failure mode that triggers them, with the first mitigating action stated before diagnosis, diagnostic decision trees carrying exact queries and dashboards, escalation and rollback branches, access…
map critical user journeys to the services on their path, assign criticality tiers using an explicit tiering rule, attribute ownership and pager rotation per service, and name every service with no owner, no tier, or a rotation that resolves to nobody. use…
specify service level indicators per critical user journey including the good-event and valid-event definitions, the measurement point and its bias, the implementation query or export behind each indicator, and the split between measured, partially measured,…
set service level objectives with windows, define and compute the error budget and its balance, account for burn rate over multiple windows, write an error budget policy with consequences that bind someone, and separate objectives agreed with the owner from…
orchestrate site reliability engineering work across slis, slos, error budgets, capacity planning, load and performance testing, chaos and resilience testing, dependency and failure-mode analysis, disaster recovery, backup and restore, alerting quality,…
account for operational toil with measured hours and how they were established, classify each recurring task as automatable partially automatable or inherent, define the elimination path per task through automation self-service or a design change that removes…
define accessibility and seo requirements for web surfaces including wcag, semantic html, keyboard navigation, focus, screen-reader behavior, metadata, structured data, canonicals, sitemaps, robots, crawlability, localization, and page-type search…
plan web backend integration across api contracts, auth, sessions, bff layers, cms connections, data models, caching, rate limits, pagination, failure modes, and ownership. use when a web surface depends on services, databases, third-party systems, identity,…
define cms and content operations for web properties including structured content models, editorial workflow, publishing rules, approvals, localization, migration, governance, content debt, metadata ownership, and day-two content maintenance. use for…
create implementation-ready frontend engineering plans for web surfaces including rendering strategy, routing, layouts, components, state, forms, data fetching, framework constraints, accessibility hooks, performance controls, and coding-agent handoffs. use…
design web information architecture including sitemap, route hierarchy, navigation, url taxonomy, content relationships, wayfinding, content hierarchy, and findability. use when Gemini needs to structure websites, portals, dashboards, docs sites, ecommerce…
create web-specific product requirements, page and route scope, user journeys, acceptance criteria, success metrics, analytics intent, source facts, risks, open questions, and downstream handoff notes for websites, web apps, landing pages, portals,…
define ux/ui design system guidance for web surfaces including component inventory, responsive behavior, design tokens, interaction states, accessibility-aware patterns, brand consistency, and design governance. use before frontend implementation or…
orchestrate complete web development workflows across product, information architecture, ux/ui design systems, frontend engineering, backend integration, cms/content operations, web security/secops, performance, accessibility, seo, testing, observability,…
coordinate post-launch web maintenance and growth including analytics-informed backlog, experiments, conversion optimization, content refresh, seo iteration, accessibility remediation, performance regression follow-up, dependency upgrades, refactors,…
design production web observability including rum, synthetic checks, availability, frontend errors, api errors, latency, core web vitals, analytics events, dashboards, alerts, ownership, incident hooks, launch monitoring, and post-launch review. use before…
plan and gate web performance using core web vitals, performance budgets, rendering cost, hydration, bundle size, images, fonts, scripts, cdn, caching, data fetching, latency, field and lab measurement, launch thresholds, and regression controls. use before…
plan web release and deployment across ci/cd, preview environments, build commands, environment promotion, hosting, cdn, edge config, cache invalidation, feature flags, launch checklist, rollback, post-release validation, and release-to-observability handoff.…
review and plan web security and secops controls across auth, sessions, browser storage, cookies, csp, security headers, cors, csrf, dependency risk, secrets, third-party scripts, cdn/edge hardening, abuse prevention, monitoring hooks, and incident readiness.…
define web testing and qa strategy across browsers, devices, responsive layouts, visual regression, forms, auth flows, integrations, accessibility checks, seo checks, performance checks, smoke tests, regression tests, release signoff, and defect triage. use…
design AI agent architecture including planning boundaries, execution loops, memory and state strategy, tool routing, approval gates, retries, delegation, and halt behavior.
design observability for AI agents and workflows including traces, prompts, model calls, tool calls, retrieval events, approvals, errors, eval probes, cost, latency, and safety signals.
orchestrate AI engineering workflows from capability intent through model, prompt, tool, agent, retrieval, eval, safety, inference, observability, release, and incident stages using connector-grounded evidence, workflow packets, stage advancement, and halt…
triage AI production incidents involving hallucination spikes, safety failures, prompt injection, tool misuse, data leakage, model regressions, cost spikes, latency degradation, eval regressions, or user harm reports.
assess readiness to release AI capabilities across requirements, evals, safety review, red-team status, inference ops, observability, rollback, docs, support handoff, and owner approval.