mit einem Klick
ai-agents-sandbox
ai-agents-sandbox enthält 37 gesammelte Skills von lloydchang, mit Repository-Berufsabdeckung und Skill-Detailseiten auf SkillsMP.
Skills in diesem Repository
Use this skill to automate the full SaaS tenant lifecycle: provisioning, configuration, scaling, suspension, and deprovisioning across multi-cloud environments. Triggers: requests to onboard a new tenant, offboard or deprovision a tenant, resize/scale a tenant's resources, clone an environment for testing, or audit tenant resource allocation and billing tags.
Discover and visualize infrastructure resources with interactive HTML output. Use when exploring new environments, understanding resource relationships, or creating infrastructure documentation.
Orchestrate and coordinate multiple AI agents for complex workflows. Use when managing agent interactions, coordinating multi-agent tasks, or implementing agent communication patterns.
Use this skill to collect, forward, and query audit logs and security events from cloud infrastructure, Kubernetes, and application layers into a SIEM. Triggers: any request to set up audit logging, query who accessed secrets, vault access, audit who accessed, accessed secrets, configure log forwarding to Sentinel or Splunk, investigate a security event, search audit logs for suspicious activity, generate a SOC compliance evidence package, configure detection rules, or produce an audit trail for a specific user/resource/time.
Manage Backstage software catalog, components, and API documentation. Use when creating catalog entities, managing component metadata, or organizing software inventory.
Use this skill to forecast resource capacity needs, identify headroom risks, recommend scaling actions, and produce capacity plans for cloud infrastructure. Triggers: any request to forecast compute or storage needs, check whether current resources can handle projected growth, plan for a product launch or traffic spike, assess autoscaler configuration, model different scaling scenarios, or produce a quarterly capacity report.
Use this skill to automate the change management lifecycle: risk scoring, change advisory board (CAB) coordination, change freeze window enforcement, rollback planning, and audit trail generation. Triggers: any request to raise a change request, score a change for risk, check if a deployment is blocked by a change freeze, coordinate an emergency change, generate the weekly change calendar, or produce a change success rate report.
Use this skill to run chaos engineering experiments and load tests to validate platform resilience, measure breaking points, and verify auto- healing behaviour. Triggers: any request to run a chaos experiment, inject a fault, load test an endpoint or system, measure platform performance under stress, validate autoscaler response, simulate a zone failure, test circuit breakers, or produce a resilience test report.
Use this skill to monitor, trigger, diagnose, and remediate CI/CD pipelines across GitHub Actions, Azure DevOps, Jenkins, and ArgoCD. Triggers: any request to check pipeline status, investigate a build failure, re-trigger a deployment, analyse flaky tests, enforce pipeline standards, generate a deployment summary report, or show DORA metrics, deployment frequency, change failure rate, mean time to restore.
Start and monitor compliance checks for SOC2, GDPR, HIPAA standards. Use when verifying infrastructure compliance, preparing for audits, or ensuring regulatory requirements are met.
Use this skill to run automated security and compliance scans across cloud infrastructure, IaC code, containers, and APIs. Triggers: any request to run a security audit, check compliance posture (SOC2, ISO27001, CIS benchmarks), scan Terraform/Kubernetes manifests, review IAM permissions, detect secrets in code, assess CVE exposure, or generate a compliance report for executive or auditor review. Also handles before major deployments, quarterly security audits, and new compliance vulnerability notifications.
Use this skill to manage container registries, image lifecycle, vulnerability scanning, and image promotion pipelines. Triggers: any request to set up or manage Azure Container Registry (ACR), push or pull images, scan for CVEs, promote images from dev to prod registries, enforce image signing, clean up old images, configure replication, manage access controls, or audit what images are running in production.
Use this skill to analyse, track, and reduce cloud infrastructure spend across Azure, AWS, and GCP. Triggers: any request to review cloud costs, identify waste or idle resources, right-size over-provisioned workloads, generate a cost report, set up budget alerts, optimise Reserved Instance or Savings Plan coverage, analyse cost per tenant, or produce recommendations for reducing monthly cloud spend.
Analyze and optimize cloud infrastructure costs using specialized subagent. Use when reviewing spending, identifying savings opportunities, or planning cost reduction strategies.
Use this skill to manage cloud database lifecycle operations including provisioning, scaling, backup/restore, failover, high-availability configuration, performance tuning, and version upgrades for Azure Database for PostgreSQL, Azure SQL, and MongoDB. Triggers: any request to provision a database, restore from backup, trigger or test a failover, scale compute or storage, tune query performance, investigate slow queries, set up read replicas, rotate credentials, or generate a database health report.
Use this skill to validate deployments before and after they go live, and to execute automated rollbacks when issues are detected. Triggers: any request to validate a deployment, run smoke tests, check rollout health, perform a canary or blue-green promotion decision, trigger or assess a rollback, or review deployment reliability metrics.
Use this skill to implement and operate an Internal Developer Portal (IDP) and self-service catalog for platform capabilities. Triggers: any request to set up or manage a Backstage developer portal, create a self-service template for a new service or environment, onboard a developer team to the platform, build a service catalog, automate the golden-path service scaffolding, or reduce toil for engineering teams that need platform resources.
Use this skill to design, implement, test, and execute disaster recovery procedures for workloads. Triggers: any request to create a DR plan, execute a region failover, test RTO/RPO objectives, run a DR drill, validate backup integrity, restore a failed environment, or assess and improve the current disaster recovery posture.
Use this skill to implement, operate, and troubleshoot GitOps workflows using ArgoCD and Flux. Triggers: any request to set up GitOps for a new cluster or tenant, configure app-of-apps patterns, investigate a sync failure, enforce drift detection, promote releases across environments, manage ArgoCD ApplicationSets, configure Flux kustomizations, or audit what version is running where across the fleet.
Use this skill to detect, triage, and execute runbooks incidents. Triggers: any incident, alert, P1, P2, P3, P4, page, outage, cluster error, 503, 5xx, service degraded, degradation, anomaly, or request to investigate; execute a runbook step-by-step; create a post-mortem; automate top incident response patterns; or reduce mean-time-to-resolution (MTTR) for recurring issues.
Use this skill to automatically collect, aggregate, and generate KPI and quarterly progress reports for teams. Triggers: any request to generate a platform report, build a quarterly review, calculate DORA metrics, summarise operational health, produce an executive dashboard, or track progress against OKRs and roadmap milestones.
Use this skill to manage the full Kubernetes cluster lifecycle across Azure AKS, AWS EKS, and GCP GKE. Triggers: any request to provision, upgrade, scale, harden, or decommission a Kubernetes cluster; manage node pools; configure RBAC or network policies; perform version upgrades with zero downtime; diagnose cluster health; or enforce Kubernetes operational standards across a multi-cloud fleet.
Use this skill to design, provision, and operate network infrastructure across Azure, AWS, and GCP. Triggers: any request to provision or create VNets/VPCs, spoke networks, peering connections, private endpoints, DNS zones, load balancers, WAF rules, firewall policies, NSGs, ExpressRoute/Direct Connect circuits, troubleshoot connectivity between tenants, diagnose why service cannot reach endpoint, services, or clouds.
Use this skill to deploy, configure observability stack for teams: metrics (Prometheus/Grafana), logging (ELK/Loki), distributed tracing (Jaeger/Tempo), and alerting pipelines. Triggers: any request to set up monitoring for a new tenant or service, configure log aggregation, create dashboards, set up distributed tracing, build alerting rules, investigate a missing metric or log gap, or produce an observability health assessment.
Use this skill as the top-level orchestrator for the Cloud AI agent. Coordinates all other skills to handle complex multi-step tasks. Triggers: any high-level or multi-domain request such as "onboard a new tenant end-to-end", "respond to a P1 incident", "prepare the quarterly business review", "run a full health check", "migrate this environment", or any task requiring more than one skill in sequence.
Use this skill to define, enforce, and audit governance policies across cloud infrastructure and Kubernetes using Open Policy Agent (OPA), Azure Policy, AWS Service Control Policies, and Kubernetes Gatekeeper. Triggers: any request to create or update a governance policy, enforce tagging standards, restrict resource types or regions, audit policy compliance, generate a policy violation report, set up guardrails for developer self-service, or implement a platform governance framework.
Use this skill to automatically generate, update, and maintain operational runbooks, architecture decision records (ADRs), platform documentation, and wiki pages. Triggers: any request to write or update a runbook, document an incident pattern, create an ADR, generate API or infrastructure documentation from code/config, produce an onboarding guide, or keep documentation in sync with current platform state.
Use this skill to manage secrets, API keys, connection strings, and TLS certificates across cloud secret stores and Kubernetes clusters. Triggers: any request to rotate a secret, renew a certificate, audit secret access, detect expiring certs, inject secrets into workloads, set up cert-manager, migrate secrets between environments, or detect hardcoded credentials in infrastructure or application code.
Perform comprehensive security analysis with dynamic context injection. Use when scanning for vulnerabilities, analyzing security posture, or responding to security incidents.
Use this skill to install, configure, and operate a service mesh (Istio or Linkerd) across Kubernetes clusters. Triggers: any request to enable mTLS between services, configure traffic management (canary, circuit breaker, retries, timeouts), set up mutual TLS enforcement, generate a service dependency map, debug inter-service connectivity, configure observability through the mesh, or enforce zero-trust service-to-service communication.
Use this skill to define, monitor, and report on platform SLAs and SLOs for uptime, deployment success, incident response, and performance. Triggers: requests to check SLA status, calculate SLO error budgets, set up alerting thresholds, generate SLA compliance reports, detect SLA breaches, or review operational reliability metrics across tenants and environments.
Use this skill to draft, structure, and send stakeholder communications for teams. Triggers: any request to write an incident notification, executive status update, platform change announcement, risk escalation, SLA breach notification, quarterly business review summary email, or cross-team alignment memo. Also triggers when asking for help communicating progress, risks, or outages to leadership or customers.
Create, manage, and monitor Temporal workflows with AI agent orchestration. Use when developing workflow definitions, monitoring execution, or troubleshooting workflow issues.
Use this skill to automate the full SaaS tenant lifecycle: provisioning, configuration, scaling, suspension, and deprovisioning across multi-cloud environments. Triggers: requests to onboard a new tenant, offboard or deprovision a tenant, resize/scale a tenant's resources, clone an environment for testing, or audit tenant resource allocation and billing tags.
Use this skill to automate cloud infrastructure provisioning, modification, and teardown using Terraform, CDK, CloudFormation, ARM templates, Google Cloud Infrastructure Manager Terraform Blueprint across AWS, Azure, and GCP. Also validates IaC changes against company standards during code reviews or before merging infrastructure changes. Triggers: any request to provision, destroy, plan, or validate cloud infrastructure; generate or review Terraform, CDK, CloudFormation, ARM templates, or Blueprints; manage state files; run drift detection; or enforce infrastructure-as-code standards.
Orchestrate and monitor Temporal AI Agent workflows. Use when managing multiple concurrent workflows, checking status, or coordinating complex multi-agent operations.
Use this skill to plan and execute migrations of cloud workloads: cloud-to- cloud (Azure to AWS, etc.), region migrations, subscription moves, Kubernetes cluster upgrades with data migration, database migrations, and SaaS tenant migrations to a new platform tier or cloud environment. Triggers: any request to migrate a workload, move a tenant to a new cluster or region, consolidate environments, perform a blue-green environment switch, or validate a migration plan before execution.