Skip to main content
在 Manus 中运行任何 Skill
一键导入
GitHub 仓库

gke-mcp

gke-mcp 收录了来自 GoogleCloudPlatform 的 25 个 skills,并提供仓库级职业覆盖和站内 skill 详情页。

已收集 skills
25
Stars
162
更新
2026-07-13
Forks
81
职业覆盖
1 个职业分类 · 已分类 100%
仓库浏览

这个仓库中的 skills

gke-ai-troubleshooting-jobset-interruption
网络与计算机系统管理员

Systematically diagnose GKE JobSet interruptions, restarts, and preemptions for AI/ML training workloads. Identifies preemption events, maintenance interruptions, bad host VMs, unhealthy pods, and coordinator worker failures.

2026-07-13
gke-ai-troubleshooting-handle-disruption-gpu-tpu
网络与计算机系统管理员

Diagnose and predict node disruption during Compute Engine host maintenance for GPU and TPU workloads.

2026-07-13
gke-ai-troubleshooting-tpu-connection-failure-vbar-oom
网络与计算机系统管理员

Diagnose and prevent `vbar_control_agent` segfaults and OOMs caused by race conditions during TPU device resets and frequent metrics collection (e.g. every 3s). Use when TPU slice initialization fails or `vbar_control_agent` crashes on TPU v6e nodes.

2026-07-10
verify-unused
网络与计算机系统管理员

Verifies if a GKE or Kubernetes cluster is unused (no active compute, external exposure, or persistent data) before allowing deletion. Evaluates external exposure (LoadBalancer Service, Ingress, Gateway, MultiClusterIngress), persistent data (Bound PVC), and active compute (Running/Pending Pods in user namespaces) with low-overhead queries and fail-close timeouts.

2026-07-03
gke-tpu-dynamic-slices-monitoring
网络与计算机系统管理员

Monitor and manage GKE TPU Dynamic Slices custom resources. Use when checking slice lifecycle states, troubleshooting failed slice creations (e.g. SliceCreationFailed, FAILED), running single or multi-slice workloads, or safely deleting/disabling slices.

2026-07-02
gke-tpu-metrics-monitoring
网络与计算机系统管理员

Monitor and troubleshoot GKE TPU workloads using GKE system metrics and PromQL.

2026-07-01
gke-skill-creator
网络与计算机系统管理员

Dynamically generates specialized GKE skills for complex troubleshooting, operational workflows, architectural setup, or performance/cost optimization. Trigger this skill whenever the user faces a novel or non-obvious GKE challenge, needs custom cluster management workflows, or standard agent capabilities fall short, even if they don't explicitly ask to create a skill.

2026-06-18
gke-ai-troubleshooting-skill-creation-guide
网络与计算机系统管理员

Expert instructions for building high-quality GKE troubleshooting skills. Codifies Step 0 context rules, zero-hallucination signatures, and explicit LQL/PromQL query requirements.

2026-06-10
gke-productionize
网络与计算机系统管理员

Assists in preparing applications and clusters on GKE for production.

2026-04-29
gke-app-onboarding
网络与计算机系统管理员

Workflows for containerizing and deploying applications to GKE for the first time.

2026-04-29
gke-workload-security
网络与计算机系统管理员

Workflows for auditing and hardening the security of GKE workloads.

2026-04-21
gke-cost-analysis
网络与计算机系统管理员

Answer natural language questions about GKE-related costs by leveraging BigQuery export and cost allocation data.

2026-04-15
gke-cluster-creator
网络与计算机系统管理员

Guides the user through creating GKE clusters using pre-defined templates (Standard, Autopilot, GPU/AI).

2026-04-13
gke-cluster-lifecycle
网络与计算机系统管理员

Guidance on managing the lifecycle and upgrades of Google Kubernetes Engine (GKE) clusters.

2026-04-13
gke-cost-optimization
网络与计算机系统管理员

Guidance on optimizing costs for Google Kubernetes Engine (GKE) clusters.

2026-04-13
gke-multi-tenancy
网络与计算机系统管理员

Guidance on implementing multi-tenancy and governance in Google Kubernetes Engine (GKE) clusters.

2026-04-13
gke-storage
网络与计算机系统管理员

Guidance on managing storage in Google Kubernetes Engine (GKE) clusters.

2026-04-13
gke-backup-dr
网络与计算机系统管理员

Workflows for configuring Backup for GKE and disaster recovery.

2026-04-10
gke-networking-edge
网络与计算机系统管理员

Workflows for configuring edge networking, ingress, and security on GKE.

2026-04-10
gke-observability
网络与计算机系统管理员

Workflows for setting up and auditing observability (logging, monitoring, tracing) on GKE.

2026-04-10
gke-reliability
网络与计算机系统管理员

Workflows for ensuring high availability and reliability of GKE workloads.

2026-04-10
gke-workload-scaling
网络与计算机系统管理员

Specific workflows for scaling GKE workloads using HPA and VPA, as well as best practices for autoscaling configuration.

2026-04-10
custom-golden-image-discovery
网络与计算机系统管理员

Expert at discovering golden base images for GKE custom nodes using technical specs or context clues.

2026-03-17
gke-compute-class-creator
网络与计算机系统管理员

Guide for creating GKE ComputeClass resources. Use this skill when users want to define custom node configurations, autoscaling priorities, or hardware requirements (e.g., Spot VMs, GPUs, specific machine families) for their GKE workloads.

2026-02-03
gke-inference-quickstart
网络与计算机系统管理员

Deploy optimized AI/ML inference workloads on GKE using Google's Inference Quickstart (GIQ). Covers model discovery, manifest generation, and deployment using native MCP tools and CLI.

2026-01-22