Run any Skill in Manus with one click

google-cloud-waf-reliability

Stars28,260

Forks2,942

UpdatedMay 13, 2026 at 13:53

Generates reliability-focused guidance for Google Cloud workloads based on the Google Cloud Well-Architected Framework. Use to evaluate a workload, identify reliability requirements, and provide actionable recommendations for building resilient, highly available systems.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

Run Skill in Manus

Source

davila7

davila7/claude-code-templates

View GitHub Repository View Creator Repositories

Download

Run Skill in Manus

Related occupationsSOC

Based on SOC occupation classification

Network and Computer Systems AdministratorsComputer and Mathematical Occupations·SOC 15-1244

SKILL.md

readonly

name	google-cloud-waf-reliability
description	Generates reliability-focused guidance for Google Cloud workloads based on the Google Cloud Well-Architected Framework. Use to evaluate a workload, identify reliability requirements, and provide actionable recommendations for building resilient, highly available systems.
source	google/skills (Apache 2.0)

Google Cloud Well-Architected Framework skill for the Reliability pillar

Overview

The Reliability pillar of the Google Cloud Well-Architected Framework provides principles and recommendations to help you design, deploy, and manage reliable, resilient, and highly available workloads in Google Cloud. A reliable system consistently performs its intended functions under defined conditions, is resilient to failures, and recovers gracefully from disruptions, thereby minimizing downtime, enhancing user experience, and ensuring data integrity.

Core principles

The recommendations in the reliability pillar of the Well-Architected Framework are aligned with the following core principles:

Define reliability based on user-experience goals: Measurement of reliability should reflect the actual experience of the system's users rather than merely relying on infrastructure metrics. Focus on outcomes that matter most to users. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/define-reliability-based-on-user-experience-goals
Set realistic targets for reliability: Determine appropriate Service Level Objectives (SLOs) that balance the cost and complexity of maximizing availability against business requirements. Utilize error budgets to manage feature velocity. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/set-targets
Build highly available systems through resource redundancy: Eliminate single points of failure by duplicating critical components across zones and regions to maintain operations during localized outages. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/build-highly-available-systems
Take advantage of horizontal scalability: Design system architectures to scale horizontally (adding more instances) to seamlessly accommodate load fluctuations and improve overall fault tolerance. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/horizontal-scalability
Detect potential failures by using observability: Implement thorough monitoring, logging, and alerting systems to proactively detect, diagnose, and address anomalies before they cause user-facing issues. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/observability
Design for graceful degradation: Architect systems to maintain critical functionality, even if at reduced performance or with limited features, when dependencies fail or the system experiences extreme stress. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/graceful-degradation
Perform testing for recovery from failures: Build confidence in system resilience by continuously simulating failures and verifying the effectiveness of automated and manual recovery procedures. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/perform-testing-for-recovery-from-failures
Perform testing for recovery from data loss: Regularly test backup and restore protocols to ensure rapid recovery from data corruption or loss, remaining within the defined Recovery Time Objective (RTO) and Recovery Point Objective (RPO). Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/perform-testing-for-recovery-from-data-loss
Conduct thorough postmortems: Foster a blameless culture by investigating outages comprehensively to understand root causes, followed by implementing measures that prevent recurrence. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/conduct-postmortems

Relevant Google Cloud products

The following are examples of Google Cloud products and features that are relevant to reliability:

Compute: Compute Engine Managed Instance Groups (MIGs), Google Kubernetes Engine (GKE), Cloud Run
Networking: Cloud Load Balancing, Cloud CDN, Cloud DNS
Storage and databases: Cloud Storage (multi-region), Cloud SQL High Availability, Spanner, Filestore, Firestore
Operations: Cloud Monitoring, Cloud Logging, Google Cloud Managed Service for Prometheus
Disaster recovery: Backup and DR Service, Filestore backups

Workload assessment questions

Ask appropriate questions to understand the reliability-related requirements and constraints of the workload and the user's organization. Choose questions from the following list:

How does your organization define and measure the reliability of your systems in relation to user experience?
How does your organization approach setting reliability targets for your services?
What is your organization's strategy for ensuring high availability through resource redundancy?
How does your organization leverage horizontal scalability to maintain performance and reliability?
How does your organization utilize observability (metrics, logs, traces) to gain insights and detect potential failures?
How does your organization manage alerting based on observability data to ensure timely responses to significant issues without causing alert fatigue?
What measures does your organization take to ensure systems can gracefully degrade during high load or partial failures?
How frequently and comprehensively does your organization test for recovery from system failures (e.g., regional failovers, release rollbacks)?
What is your organization's approach to testing for recovery from data loss?
How does your organization conduct and utilize postmortems after incidents?

Validation checklist

Use the following checklist to evaluate the architecture's alignment with reliability recommendations:

User-focused SLIs and SLOs are explicitly defined and actively monitored.
The architecture avoids single points of failure through cross-zone or cross-region redundancy.
Autoscaling is enabled to handle variable demand without manual intervention.
Application and infrastructure health checks are configured to trigger automated failovers.
Regular backup schedules are in place, and restoration processes are routinely tested.
The system architecture incorporates patterns like circuit breakers, retries with exponential backoff, and rate limiting to support graceful degradation.
Game days or chaos engineering practices are regularly held to validate failure recovery.
A formalized, blameless postmortem process exists to ensure organizational learning from operational incidents.

More from this repository

same repository

context-architecture

davila7/claude-code-templates

Audit and incrementally retrofit an existing codebase so its intent and behavior are equally legible to people and AI agents. Applies Context Architecture's eight principles: place AGENTS.md at boundaries, bind every context claim to a mechanism (lint, types, tests, review), name boundaries, and find context that has rotted. Use when an agent reimplements code that already exists, invents structure, follows stale or deleted docs, propagates a deprecated pattern, or resolves ambiguity at random, or when asked to make a repository "agent-ready", "AI-legible", or to add or fix AGENTS.md / CLAUDE.md files.

2026-06-2128.3k

x-twitter-scraper

davila7/claude-code-templates

Use when the user wants to integrate with the X (Twitter) API via Xquik to search tweets, look up user profiles, extract followers, run giveaway draws, monitor accounts, or access trending topics. Also use when the user mentions 'Xquik,' 'Twitter API,' 'X API,' 'tweet scraper,' 'follower extraction,' or 'Twitter monitoring.' Covers REST API, webhooks, and MCP server setup.

2026-06-2128.3k

pdf-fill-studio

davila7/claude-code-templates

Fill any PDF locally and place each value precisely in a visual editor. Use when the user wants to fill out a PDF form, enter data into a PDF, complete a tax/insurance/bank form, or position text on a flat/scanned PDF. Handles flat (field-less) PDFs, per-character (comb) fields, and native AcroForm fields; leaves the signature blank for the user to sign.

2026-05-2928.3k

android-cicd

davila7/claude-code-templates

Automated Android CI/CD pipeline to Google Play — supports TWA, React Native, Flutter, and native Android. Run npx android-cicd to set up keystore generation, GitHub Secrets, and a multi-stage workflow (internal/alpha/beta/production) with auto-bump versionCode.

2026-05-2328.3k

bleu

davila7/claude-code-templates

Use this skill whenever a developer wants to turn an idea into a complete, production-ready, end-to-end system plan BEFORE writing any code. Trigger on 'plan this system', 'design the architecture for', 'help me blueprint', 'deep plan for X', 'break this idea into components', 'expand into action points', 'full implementation plan', or when the user pastes a project idea wanting architecture, components, pipelines, and file-level execution mapped out. Casual phrasing also triggers: 'help me think this through end-to-end', 'plan before coding'. Also covers living-workspace patterns: self-improving knowledge bases, reflection loops with auditor agents, four-agent teams, schema-as-code, wiki health scoring. **Resume triggers**: 'where did we leave off', 'continue this plan', 'resume my blueprint' - rehydrates state from disk via SESSION.md/NEXT.md/decisions/. Web research is mandatory every invocation.

2026-05-2328.3k

building-blog

davila7/claude-code-templates

Use when adding a blog to a Next.js + Sanity site, building a blog section from scratch, integrating Sanity CMS for editorial content, or setting up an SEO-optimized article surface. Triggers on phrases like 'add a blog', 'build the blog', 'create blog section', 'set up blog with Sanity', 'integrate Sanity CMS', 'blog architecture', or any new-blog scoping conversation on a Next.js project.

2026-05-2328.3k

name	google-cloud-waf-reliability
description	Generates reliability-focused guidance for Google Cloud workloads based on the Google Cloud Well-Architected Framework. Use to evaluate a workload, identify reliability requirements, and provide actionable recommendations for building resilient, highly available systems.
source	google/skills (Apache 2.0)

Google Cloud Well-Architected Framework skill for the Reliability pillar

Overview

Core principles

The recommendations in the reliability pillar of the Well-Architected Framework are aligned with the following core principles:

Define reliability based on user-experience goals: Measurement of reliability should reflect the actual experience of the system's users rather than merely relying on infrastructure metrics. Focus on outcomes that matter most to users. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/define-reliability-based-on-user-experience-goals
Set realistic targets for reliability: Determine appropriate Service Level Objectives (SLOs) that balance the cost and complexity of maximizing availability against business requirements. Utilize error budgets to manage feature velocity. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/set-targets
Build highly available systems through resource redundancy: Eliminate single points of failure by duplicating critical components across zones and regions to maintain operations during localized outages. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/build-highly-available-systems
Take advantage of horizontal scalability: Design system architectures to scale horizontally (adding more instances) to seamlessly accommodate load fluctuations and improve overall fault tolerance. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/horizontal-scalability
Detect potential failures by using observability: Implement thorough monitoring, logging, and alerting systems to proactively detect, diagnose, and address anomalies before they cause user-facing issues. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/observability
Design for graceful degradation: Architect systems to maintain critical functionality, even if at reduced performance or with limited features, when dependencies fail or the system experiences extreme stress. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/graceful-degradation
Perform testing for recovery from failures: Build confidence in system resilience by continuously simulating failures and verifying the effectiveness of automated and manual recovery procedures. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/perform-testing-for-recovery-from-failures
Perform testing for recovery from data loss: Regularly test backup and restore protocols to ensure rapid recovery from data corruption or loss, remaining within the defined Recovery Time Objective (RTO) and Recovery Point Objective (RPO). Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/perform-testing-for-recovery-from-data-loss
Conduct thorough postmortems: Foster a blameless culture by investigating outages comprehensively to understand root causes, followed by implementing measures that prevent recurrence. Grounding document: https://docs.cloud.google.com/architecture/framework/reliability/conduct-postmortems

Relevant Google Cloud products

The following are examples of Google Cloud products and features that are relevant to reliability:

Compute: Compute Engine Managed Instance Groups (MIGs), Google Kubernetes Engine (GKE), Cloud Run
Networking: Cloud Load Balancing, Cloud CDN, Cloud DNS
Storage and databases: Cloud Storage (multi-region), Cloud SQL High Availability, Spanner, Filestore, Firestore
Operations: Cloud Monitoring, Cloud Logging, Google Cloud Managed Service for Prometheus
Disaster recovery: Backup and DR Service, Filestore backups

Workload assessment questions

Ask appropriate questions to understand the reliability-related requirements and constraints of the workload and the user's organization. Choose questions from the following list:

How does your organization define and measure the reliability of your systems in relation to user experience?
How does your organization approach setting reliability targets for your services?
What is your organization's strategy for ensuring high availability through resource redundancy?
How does your organization leverage horizontal scalability to maintain performance and reliability?
How does your organization utilize observability (metrics, logs, traces) to gain insights and detect potential failures?
How does your organization manage alerting based on observability data to ensure timely responses to significant issues without causing alert fatigue?
What measures does your organization take to ensure systems can gracefully degrade during high load or partial failures?
How frequently and comprehensively does your organization test for recovery from system failures (e.g., regional failovers, release rollbacks)?
What is your organization's approach to testing for recovery from data loss?
How does your organization conduct and utilize postmortems after incidents?

Validation checklist

Use the following checklist to evaluate the architecture's alignment with reliability recommendations:

User-focused SLIs and SLOs are explicitly defined and actively monitored.
The architecture avoids single points of failure through cross-zone or cross-region redundancy.
Autoscaling is enabled to handle variable demand without manual intervention.
Application and infrastructure health checks are configured to trigger automated failovers.
Regular backup schedules are in place, and restoration processes are routinely tested.
The system architecture incorporates patterns like circuit breakers, retries with exponential backoff, and rate limiting to support graceful degradation.
Game days or chaos engineering practices are regularly held to validate failure recovery.
A formalized, blameless postmortem process exists to ensure organizational learning from operational incidents.