Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة
مستودع GitHub

devops-sre-skills

يحتوي devops-sre-skills على 17 من skills المجمعة من bregman-arie، مع تغطية مهنية على مستوى المستودع وصفحات skill داخل الموقع.

skills مجمعة
17
Stars
16
محدث
2026-03-20
Forks
5
التغطية المهنية
2 فئات مهنية · 100% مصنفة
مستكشف المستودعات

Skills في هذا المستودع

triage-aws-accessdenied
محللو أمن المعلومات

Identify why an AWS API call is denied and what policy element blocks it.

2026-03-20
triage-eks-node-notready
مديرو الشبكات وأنظمة الحاسوب

Diagnose EKS worker nodes in NotReady and determine safe remediation.

2026-03-20
triage-cloud-cost-spike
مديرو الشبكات وأنظمة الحاسوب

Identify the primary drivers of a sudden cloud cost increase and implement safe mitigations.

2026-03-20
triage-gcp-quota-exceeded
مديرو الشبكات وأنظمة الحاسوب

Diagnose GCP quota errors and identify the quota, scope, and fastest safe mitigation.

2026-03-20
triage-kubernetes-node-pressure
مديرو الشبكات وأنظمة الحاسوب

Diagnose node-level memory/disk/pid pressure and determine safe mitigations.

2026-03-20
triage-pending-pods
مديرو الشبكات وأنظمة الحاسوب

Diagnose pods stuck in Pending and identify scheduling constraints.

2026-03-20
triage-kubernetes-service-dns
مديرو الشبكات وأنظمة الحاسوب

Diagnose in-cluster DNS resolution failures and isolate root causes.

2026-03-20
triage-error-budget-burn
مديرو الشبكات وأنظمة الحاسوب

Investigate rapid SLO burn and identify whether the driver is errors, latency, or availability.

2026-03-20
triage-suspected-secret-exposure
محللو أمن المعلومات

Contain and respond to a suspected credential/secret exposure without increasing blast radius.

2026-03-20
investigate-terraform-drift
مديرو الشبكات وأنظمة الحاسوب

Determine why actual infrastructure differs from Terraform state and choose a safe reconciliation path.

2026-03-20
triage-argo-cd-app-outofsync
مديرو الشبكات وأنظمة الحاسوب

Identify why an Argo CD application is OutOfSync and resolve safely.

2026-03-20
sev1-first-15-minutes
مديرو الشبكات وأنظمة الحاسوب

Execute the initial incident workflow to stabilize, communicate, and delegate.

2026-03-20
diagnose-crashloopbackoff
مديرو الشبكات وأنظمة الحاسوب

Triage pods restarting repeatedly and identify the most likely root cause.

2026-03-20
diagnose-imagepullbackoff
مديرو الشبكات وأنظمة الحاسوب

Determine why a pod cannot pull its container image and resolve safely.

2026-03-20
triage-latency-regression
مديرو الشبكات وأنظمة الحاسوب

Identify what changed and where latency increased, using logs/metrics/traces.

2026-03-20
example-skill-name
مديرو الشبكات وأنظمة الحاسوب

One-line description of when this skill is used.

2026-03-20
recover-terraform-state-lock
مديرو الشبكات وأنظمة الحاسوب

Safely assess and recover from a stuck Terraform state lock.

2026-03-20