Skip to main content

nrp-cluster

Build, deploy, and monitor Kubernetes batch jobs on the NRP Nautilus cluster. Use this skill for pipeline execution workflows, S3/W&B setup checks, jobs that call NRP-hosted open-weight LLMs or Brave Search through managed credentials, and job-level troubleshooting.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
braingeneers/mission_control
آخر نشاط في المصدر
١٠ أغسطس ٢٠٢٦ في ١٨:١٧
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
١
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.

مستكشف الملفات
9 ملفات

عرض SKILL.md

SKILL.md
تعليمات المصدر · معاينة للقراءة فقط
name
nrp-cluster
description
Build, deploy, and monitor Kubernetes batch jobs on the NRP Nautilus cluster. Use this skill for pipeline execution workflows, S3/W&B setup checks, jobs that call NRP-hosted open-weight LLMs or Brave Search through managed credentials, and job-level troubleshooting.
# NRP Cluster Management Use this skill when working on compute pipelines that run as Kubernetes jobs on the NRP Nautilus cluster. ## Start Here 1. Read these required references before making recommendations: - `references/agent-context.md` - `references/readme.md` - `references/nrp_setup.md` - For a job that calls an NRP-hosted LLM, also read `references/hosted-llms.md` and verify the current official NRP documentation linked there. - For a job that calls Brave Search, also read `references/brave-search.md` and verify the current official API and pricing documentation linked there. 2. Confirm the credentials required by the workload are available without printing them: - `kubelogin` plugin and a valid kubeconfig for NRP cluster access (PowerShell/WSL path differs by environment). - `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` only when the job uses S3. - `WANDB_API_KEY` only when the run uses W&B. - The existing `nrp-llm-api-key/nrp_llm_api_key` Kubernetes Secret key only when the job calls the NRP-hosted LLM API. - The existing `brave-search-api/brave-search-api` Kubernetes Secret key only when the job calls the Brave Search API. 3. Confirm user intent: - Build/push image - Update job config - Deploy job - Monitor job state - Troubleshoot and cleanup 4. Choose a workflow before editing commands: - `/.agent/workflows/build-push.md` for image publishing. - `/.agent/workflows/deploy-job.md` for job creation and initial validation. - `/.agent/workflows/monitor-job.md` for pod/job monitoring and cleanup. 5. Treat all credentials as ephemeral environment values. Never request or print raw secret contents. ## What this skill covers - Container image build and push with Linux/amd64 target for NRP. - K8s Job submission using `envsubst` and a templated `jobdefinition.yaml`. - Runtime-aware job policy checks (resource limits, GPU policy, sleep bans). - Pod/job status checks, log retrieval, and failed-job cleanup. - Template-level customization guidance (pipeline scripts, S3 prefixes, command entrypoint). - OpenAI-compatible NRP-hosted LLM calls, model discovery, Kubernetes Secret injection, cache isolation, and retry/resume design. - Brave Search calls, task-scoped Kubernetes Secret injection, and shared-quota request budgeting. ## Policy Reminders (high priority) - No `sleep infinity` in batch jobs. - Resource limits should track requests within policy boundaries. - GPU utilization should stay above 40% on GPU jobs. - Max concurrent jobs should not exceed 400. - `backoffLimit` and `restartPolicy` policy choices should reflect single-attempt jobs. ## Reference Loading Read only the files needed for your task: - `references/nrp_setup.md` for platform setup, prerequisites, and troubleshooting. - `references/readme.md` for the template architecture and quick start. - `references/agent-context.md` for agent-specific command patterns. - `references/hosted-llms.md` for jobs that call NRP-hosted open-weight models. - `references/brave-search.md` for jobs and workflows that call Brave Search. - `/.agent/workflows/build-push.md` to build/publish. - `/.agent/workflows/deploy-job.md` to submit jobs. - `/.agent/workflows/monitor-job.md` to inspect and clean up jobs/pods. ## Workflow ### 1) Build and push image If image freshness is unknown, run `turbo` workflow: - `/.agent/workflows/build-push.md` ### 2) Configure job intent Before submission, ask for or confirm: - Job name (`NAME`) - Job prefix (`JOB_PREFIX`) - Resource requests and limits - GPU count/type if required - Optional run command override - NRP-hosted LLM model and request policy if the job calls the managed API; discover active model IDs instead of assuming a stale model name. - An explicit per-run Brave request budget if the job calls Brave Search; include retries and pagination in the shared monthly usage estimate. ### 3) Deploy - Use `/.agent/workflows/deploy-job.md`. - Verify image and credentials are available before submission. - Confirm `kubectl` access and cluster context are valid. - For hosted LLM use, inject the existing API key with `secretKeyRef`; never retrieve its value into the submission command or request a GPU solely for remote inference. - For Brave Search use, inject the existing API key only into calling tasks with `secretKeyRef`; never retrieve its value from the secrets repository into a submission command. ### 4) Monitor and debug - Use `/.agent/workflows/monitor-job.md`. - Capture job/pod status and recent logs after submission. - Delete finished/failed jobs when requested. ## Escalation Escalate for issues that are clearly out-of-repo: - Kubeconfig or OIDC token acquisition failures. - NRP quota/policy enforcement questions. - Persistent cluster scheduling failures unrelated to job YAML. - Registry permission errors preventing image push.
عرض على GitHub