Skip to main content

nrp-cluster

Build, deploy, and monitor Kubernetes batch jobs on the NRP Nautilus cluster. Use this skill for pipeline execution workflows, S3/W&B setup checks, jobs that call NRP-hosted open-weight LLMs or Brave Search through managed credentials, and job-level troubleshooting.

설치로 이동

소스 정보

저장소
braingeneers/mission_control
최근 소스 활동
2026년 8월 10일 18:17
감지된 SKILL.md 언어
영어
스타
1
포크
0

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
9 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
nrp-cluster
description
Build, deploy, and monitor Kubernetes batch jobs on the NRP Nautilus cluster. Use this skill for pipeline execution workflows, S3/W&B setup checks, jobs that call NRP-hosted open-weight LLMs or Brave Search through managed credentials, and job-level troubleshooting.
# NRP Cluster Management Use this skill when working on compute pipelines that run as Kubernetes jobs on the NRP Nautilus cluster. ## Start Here 1. Read these required references before making recommendations: - `references/agent-context.md` - `references/readme.md` - `references/nrp_setup.md` - For a job that calls an NRP-hosted LLM, also read `references/hosted-llms.md` and verify the current official NRP documentation linked there. - For a job that calls Brave Search, also read `references/brave-search.md` and verify the current official API and pricing documentation linked there. 2. Confirm the credentials required by the workload are available without printing them: - `kubelogin` plugin and a valid kubeconfig for NRP cluster access (PowerShell/WSL path differs by environment). - `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` only when the job uses S3. - `WANDB_API_KEY` only when the run uses W&B. - The existing `nrp-llm-api-key/nrp_llm_api_key` Kubernetes Secret key only when the job calls the NRP-hosted LLM API. - The existing `brave-search-api/brave-search-api` Kubernetes Secret key only when the job calls the Brave Search API. 3. Confirm user intent: - Build/push image - Update job config - Deploy job - Monitor job state - Troubleshoot and cleanup 4. Choose a workflow before editing commands: - `/.agent/workflows/build-push.md` for image publishing. - `/.agent/workflows/deploy-job.md` for job creation and initial validation. - `/.agent/workflows/monitor-job.md` for pod/job monitoring and cleanup. 5. Treat all credentials as ephemeral environment values. Never request or print raw secret contents. ## What this skill covers - Container image build and push with Linux/amd64 target for NRP. - K8s Job submission using `envsubst` and a templated `jobdefinition.yaml`. - Runtime-aware job policy checks (resource limits, GPU policy, sleep bans). - Pod/job status checks, log retrieval, and failed-job cleanup. - Template-level customization guidance (pipeline scripts, S3 prefixes, command entrypoint). - OpenAI-compatible NRP-hosted LLM calls, model discovery, Kubernetes Secret injection, cache isolation, and retry/resume design. - Brave Search calls, task-scoped Kubernetes Secret injection, and shared-quota request budgeting. ## Policy Reminders (high priority) - No `sleep infinity` in batch jobs. - Resource limits should track requests within policy boundaries. - GPU utilization should stay above 40% on GPU jobs. - Max concurrent jobs should not exceed 400. - `backoffLimit` and `restartPolicy` policy choices should reflect single-attempt jobs. ## Reference Loading Read only the files needed for your task: - `references/nrp_setup.md` for platform setup, prerequisites, and troubleshooting. - `references/readme.md` for the template architecture and quick start. - `references/agent-context.md` for agent-specific command patterns. - `references/hosted-llms.md` for jobs that call NRP-hosted open-weight models. - `references/brave-search.md` for jobs and workflows that call Brave Search. - `/.agent/workflows/build-push.md` to build/publish. - `/.agent/workflows/deploy-job.md` to submit jobs. - `/.agent/workflows/monitor-job.md` to inspect and clean up jobs/pods. ## Workflow ### 1) Build and push image If image freshness is unknown, run `turbo` workflow: - `/.agent/workflows/build-push.md` ### 2) Configure job intent Before submission, ask for or confirm: - Job name (`NAME`) - Job prefix (`JOB_PREFIX`) - Resource requests and limits - GPU count/type if required - Optional run command override - NRP-hosted LLM model and request policy if the job calls the managed API; discover active model IDs instead of assuming a stale model name. - An explicit per-run Brave request budget if the job calls Brave Search; include retries and pagination in the shared monthly usage estimate. ### 3) Deploy - Use `/.agent/workflows/deploy-job.md`. - Verify image and credentials are available before submission. - Confirm `kubectl` access and cluster context are valid. - For hosted LLM use, inject the existing API key with `secretKeyRef`; never retrieve its value into the submission command or request a GPU solely for remote inference. - For Brave Search use, inject the existing API key only into calling tasks with `secretKeyRef`; never retrieve its value from the secrets repository into a submission command. ### 4) Monitor and debug - Use `/.agent/workflows/monitor-job.md`. - Capture job/pod status and recent logs after submission. - Delete finished/failed jobs when requested. ## Escalation Escalate for issues that are clearly out-of-repo: - Kubeconfig or OIDC token acquisition failures. - NRP quota/policy enforcement questions. - Persistent cluster scheduling failures unrelated to job YAML. - Registry permission errors preventing image push.
GitHub에서 보기