Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
A direct command skips the review prompt. Inspect the source before running it.
Any Agent-Skills (SKILL.md)-compatible agent โ Claude Code, Codex, Cursor, Trae, Gemini CLI, etc.
Needs a shell + SSH (or a platform CLI/API) to drive the remote box; scripts are bash/python. A few
durable-monitoring recipes assume a host background-task runner + scheduler โ map to the running
agent's equivalents (references/monitoring_patterns.md ยง7). Companion skills (verifying-dl-experiments,
superpowers:*, huggingface-skills:*) are optional separate installs.
Deploy and babysit long-running GPU jobs on rented boxes you don't own, across any platform, and
get the result off the box before the meter or a preemption kills it. The core insight: you are a
short-term tenant on someone else's machine โ so the job is to detach the work, make the result
outlive the instance, and stop the meter safely, not to provision a cluster.
This skill is platform-agnostic at the core, platform-specific at the edges: a fixed set of
operating principles + a 6-phase lifecycle that hold everywhere, plus one profile per platform
(profiles/<platform>.md) that owns every concrete path, proxy, billing verb, and spot semantic. Its
defensible value is the union the big orchestrators skip: Chinese cgroup-isolated rentals + bare-SSH
cheap boxes + the disk-budget / monitoring / teardown reality that is the job on metered hardware.
When to Use This Skill
Use whenever the user deploys, trains, monitors, or troubleshoots a long-running GPU job on a RENTED
or remote instance they do not own โ training, eval, ablation sweeps, batch inference, or large data
processing โ on AutoDL, RunPod, vast.ai, Lambda, Paperspace, Chinese platforms (ๆๆบไบ/็ฉๆฑ ไบ/Featurize/
ๆฝ็ฟๆ่), a bare SSH box, Slurm, or Kubernetes; single OR multi-instance. Triggers (multilingual):
"่ฟ็จ GPU ่ฎญ็ป", "GPU ็ง่ต", "GPU rental", "็งๅก", "spot ๆขๅ ", "spot preemption", "ๆญ็น็ปญ่ฎญ",
"resumable training", "tmux ่ฎญ็ปๅฎๆค", "้ฒ SSH ๆญ็บฟ", "scp/rsync ไธไผ ", "ๅคๅฎไพ ablation",
"่ฟ็จ GPU ็ๆง", "็้ฑๅ ณๆบ/้ๆฏๅฎไพ", "stop vs terminate billing", "checkpoint ็ฃ็ๆปก",
"CUDA OOM/ๆพๅญไธ่ถณ", "loss NaN/loss spike", "loss ไธไธ้/ไธๆถๆ", "overfit ๅ batch",
"FSDP/DeepSpeed ้ ็ฝฎ", "ๅคๅก่ฎญ็ป hang", "dataloader worker/ๆฐๆฎๅขๅนฟ bug". NOT for purely local
single-GPU training, in-instance multi-GPU DDP (use torchrun/accelerate), managed multi-cloud
price-shopping (use SkyPilot's skill), or zero-ops serverless (use Modal).
When NOT to use โ and what to use instead
Situation
Use instead
Local single-GPU, or multi-GPU DDP inside one box
torchrun / accelerate directly
Managed multi-cloud price-shopping + auto spot-recovery across Western clouds
SkyPilot (has its own Agent Skill) โ then come back here to make your code resume-correct so its recovery actually works
Open BYOC dev enprojectnments
dstack
Zero-ops serverless inference
Modal
"Is this metric / ablation delta real?"
REQUIRED:verifying-dl-experiments (this skill owns running the job; that one owns whether the number is true)
This skill is for the blind spot those tools leave: AutoDL + Chinese platforms, bare SSH/Slurm/K8s
rentals, and the operational gotchas (inode caps, mirror stalls, cgroup OOM, silent sync, spot grace
windows, irreversible teardown) that survive whichever provisioner you use.
Operating principles (the WHY โ 10 invariants)
These hold on every metered, isolated, rented GPU; only the paths/CLI change. One line each; the deep
form with cross-platform nuance is in references/principles.md (read it before Phase 0).
Minimize paid wall-clock. The meter runs the whole time โ smoke locally on CPU before renting, launch detached, release the instant verification passes.
Cheap checks before expensive compute. A 1โ2 batch CPU smoke (logger off) kills import/config/shape/scale bugs for ~free. (Smoke content โ verifying-dl-experiments.)
Trust artifacts you loaded, not log lines that claim success. "synced/saved/done" lies under a silently-failed write; a watcher's own state is also a claim โ reconcile it against the real process/artifact.
Know what survives stop vs destroy. Per platform, identify exactly which mount survives a stop and which survives a terminate โ the data you need often lives on the volatile one. (The single biggest portability trap.)
Storage fails on the dimension you're not watching. Disk dies on inodes before bytes; the real hog hides in a symlinked cache; clean by value (keep tiny evidence, drop big scratch); monitor df -i, not just df -h.
Never mutate inputs under a live run. A running job holds its scripts in memory by byte-offset; overwriting one mid-run re-executes blocks. Version filenames.
Design for retry โ failure is probabilistic, transfers are flaky, mirrors are route-specific. Make wrappers idempotent + resumable; retry the identical config; wrap bulk transfers in timeout+resume loops; a mirror/proxy speeds ONE route โ validate on the same route the real transfer uses.
Checkpoint-to-durable + idempotent resume is the universal spine. File checkpoint to the platform's durable location + unconditional load-latest-on-startup is the one mechanism that survives an SSH drop, a Slurm walltime kill, a K8s reschedule, a spot preemption, and a Colab disconnect. The detach primitive (tmux/sbatch/Job/commit) is the swappable plug; this is the invariant.
Cost and destructive actions are the user's call. Never auto-release/terminate, never delete durable files without confirmation; if cleanup can't free space, ask to expand the disk rather than silently shrink the experiment.
Teach the user the platform, don't just drive it. Most users don't know a platform's non-obvious conveniences (one-click SSH-key registration, GPU-availability notifications, built-in panels) or its danger clocks (auto-release/auto-delete timers on a stopped box โ AutoDL releases a ๅ ณๆบ instance after 15 days โ data disk gone; a stop that keeps billing; low-balance purge). Surface them on first contact โ #9 stops the agent doing the dangerous thing, #10 warns the human before the clock fires. Per-platform list โ each profile's Surface to the user block.
Monitoring physics (substrate for #3): foreground Bash hard-caps at 600 s; run_in_background has no cap and notifies on exit; a never-exiting watcher never notifies; an unquoted | in a poll regex reads stdin and hangs forever. The four-layer monitoring architecture is built on these facts โ references/monitoring_patterns.md.
Code discipline (the wrapper & training scripts you write)
Two rules govern the launch/wrapper/training code this skill has you write โ corollaries of #1 and #8, not new invariants:
Reuse before writing. Take the lowest rung that already works before adding code: the base image's pre-installed stack + platform features โ a framework/library utility (torchrun / accelerate / HF) โ your existing scripts/ templates โ minimal new code. On a metered box a needless pip install also burns paid wall-clock and can break the image's ABI โ Phase 1's rule (the prebuilt image is the env; don't conda create on a rental) is exactly this principle applied to dependencies.
Floor โ minimum bounds scope, not correctness. Shrinking code must never drop what makes an expensive run survivable: checkpoint-to-durable + idempotent resume (#8), atomic writes, the error handling that prevents losing a long run, or seed/determinism logging. Keep one minimal self-check for non-trivial logic.
Pick your platform profile FIRST
Read the matching profile before Phase 0 โ it owns every path, proxy, credential location, billing
verb, and spot rule the phases below delegate to. Each follows the same 8-field schema
(profiles/_schema.md).
New here? The path is: (1) find your platform in the table below โ (2) read that profile's LAUNCH
section (it walks rent โ register SSH key โ reach the box) โ (3) come back and run the 6 phases from Phase 0.
Already have a box you can ssh into? Skip straight to Phase 0.
You're onโฆ
Profile
Kind
Detach primitive
Meter-stop verb
AutoDL (deepest, battle-tested)
profiles/autodl.md
ssh-rental
tmux
ๅ ณๆบ (stops meter, keeps disk โ the AutoDL exception)
RunPod
profiles/runpod.md
ssh-rental
tmux
terminate (stop still bills 2ร; destroys volume disk)
per-platform (data disk often bills while stopped)
Bare SSH box / Slurm / K8s / Colab-Kaggle
profiles/generic-ssh.md
ssh / slurm / k8s
tmux / sbatch / Job / commit
manual (a forgotten box bills 24/7)
Profile confidence: AutoDL is battle-tested from the author's daily use; the other six profiles are
built from each platform's official docs + community reports (cited inline, verified <month>) and not
yet independently live-tested โ lean on the Phase-0 live measurements and re-verify any teardown/
billing fact against current docs before betting money or data (references/self-improvement.md ยง5).
Mental verb model (one API across all platforms; the profile binds each verb to real commands):
up (rent+reach) โ push (code/data on) โ run (detached + checkpointing) โ watch (durable monitor) โ pull (results off + verify) โ down (stop the meter).
Default workflow (6 phases)
Skip phases already done. Each phase delegates substrate to the profile and ends in a runnable check.
Phase 0 โ Enprojectnment audit. Read the profile's STORAGE survival-matrix + region/DC-lock. Measure live:
df -h && df -i <data-mount>, cgroup memory.max, nvidia-smi. Pre-compute the checkpoint disk budget
(ckpt_size ร N + scratch). โ verify:nvidia-smi shows the expected GPU and df -i is not near 100%.
Phase 1 โ SSH + credentials. Set the alias/env per the profile (the prebuilt image/base IS the env โ
do not conda create on a rental). Never rented before? the profile's LAUNCH section walks rent โ register SSH key โ connect. Push secrets via stdin, never onto a shared/durable FS
(references/ssh_transport.md). โ verify:ssh <alias> 'python -c "import torch;print(torch.cuda.is_available())"'.
Phase 2 โ Wrapper + CPU-smoke gate. Build an idempotent run_one/run_queue from scripts/ (parameterized
from the profile's OVERRIDES; size batch/workers to the box for a standalone run, but PIN them across cells for a fair comparison โ references/training/throughput-profiling.md). Run the cheap CPU smoke locally BEFORE renting โ it kills the dumb,
expensive failures (e.g. python -m <your.train.module> --limit-batches 2 --epochs 1 โ substitute your own entrypoint; this gate needs your training code plugged in). โ verify: that smoke exits 0 on 2 batches with the logger disabled.
Phase 3 โ Detached launch. Launch via the profile's detach primitive; probe briefly (log head + alive +
no traceback), then hand back โ never a blocking foreground sleep. โ verify: within 60 s, the detach
session is alive and the first log line shows the expected step/epoch.
Phase 4 โ Durable monitoring. For anything over ~1โ2 h, deploy the four-layer architecture
(references/monitoring_patterns.md): on-box self-completion chain + session patrol loop + event sentinels +
recovery handbook. On Claude Code, fire the L2 patrol via /loop 30m (or ScheduleWakeup) running scripts/health_patrol.sh.template; a host with no local recurring runner wires the on-box self-push instead (references/monitoring_patterns.md ยง7). A session-bound watcher alone dies with the session. Classify each outcome โ
fixed remediation; never blind-retry. โ verify: the patrol reports even when nothing changed.
Phase 5 โ Aggregate + verify + teardown. Checked-sync to durable storage (gate the success line on the
copy result โ principle #3), then load-and-verify each artifact (scripts/verify_local.py), THEN the profile's
meter-stopping action. โ verify:verify_local.py reports 100% OK before any teardown.
Iron Law โ teardown gate: NO release / terminate / destroy / file-delete until checkpoints are
pulled to local AND verified by load, and the user has explicitly approved the cost-affecting action.
"It looked done in the log" is not evidence (principle #3). On most platforms the meter-stopping action is
irreversible (deletes the disk) โ confirmation matters more, not less.
Parallel ablation fan-out
For N ablation cells: one job per cell, an isolated write path per job (no shared mutable output), launched
across instances/queues. REQUIRED:superpowers:dispatching-parallel-agents supplies the independence
predicate (don't fan out onto shared state) and the mandatory post-fan-out reconciliation. FS-shared deployment
pattern โ references/parallel_ablation.md.
Quick reference โ the four facts that bite per platform
Full detail in each profile; this table is the at-a-glance.
Platform
Survives stop
Survives destroy
Spot grace
China mirror needed
AutoDL
/root + data + FS
FS only
n/a
yes (/etc/network_turbo, hf-mirror)
RunPod
volume disk (bills 2ร)
Network Volume only
~5 s SIGTERMโKILL
no (hf_transfer)
vast.ai
disk (bills forever)
nothing
~0 s (abrupt)
no
Lambda
n/a (no stop)
nothing
n/a (on-demand)
no
China (ๆๆบไบ/็ฉๆฑ ไบ/โฆ)
varies; data disk bills
per-platform persistent vol
n/a
yes
generic-SSH/Slurm/K8s
you own it
you own it
Slurm SIGTERMโKillWait (def 30 s)
only if in China
Common gotchas (top 8 inline โ full catalog in references/)
The universal ones that cost the most GPU-hours. Symptom โ fix; root cause + the rest in
references/gotchas_universal.md (run grep -i '<keyword>' references/gotchas_universal.md to jump).
SSH drops on pkill -9 (exit 255 + "Connection reset") โ normal; re-ssh to verify, don't panic.
tmux holds the script in memory โ editing it mid-run re-executes blocks; version the filename.
cgroup OOM with no traceback (bare Killed / exit 137) โ num_workers ร big-tensor; size workers vs memory.max, not CPU count.
Silent sync failure โ cp โฆ 2>/dev/null; echo synced lies on a full/inode-exhausted FS; gate the success line on the actual copy result.
Spot preemption grace is tiny (~5 s โ ~0 s on the platforms profiled here; AWS-style 2-min grace only on clouds not profiled) โ a SIGTERM-flush handler is NOT a safety net; checkpoint on a timer to durable storage, load-latest unconditionally (references/spot-resilience.md).
"Stop" rarely stops the meter โ only terminate/destroy does, and it's irreversible (deletes the disk). Know the verb from the profile before you click, and on RunPod a stopped Pod can even restart with zero GPUs.
CRLF breaks .sh on Linux โ author on Windows โ .gitattributes*.sh text eol=lf; on-box unblock sed -i 's/\r$//'.
When training itself breaks (the model, not the platform)
Platform ops is only half the job โ once the box is running, training breaks in its own ways. The
references/training/ layer is the debug knowledge for the run itself. Boundary: this layer owns
"make it run, fast, and not crash"; verifying-dl-experiments owns "is the number real" โ
cross-link it for collapse / leakage / metric-validity. Every entry is symptom โ root cause โ fix with
cited current docs.
references/training/oom-memory.md โ CUDA/VRAM + host-RAM OOM and the fit-it ladder (grad-accum โ bf16 โ activation-checkpointing โ expandable_segments โ FSDP/ZeRO โ CPU/NVMe offload โ LoRA/QLoRA); OOM-at-a-specific-step (first backward / val / longest batch); the memory snapshot + visualizer.
Companion skills (separate installs; REQUIRED reading where present)
These are separate Agent Skills, not bundled here โ install them for the full experience. On an
agent where a companion isn't installed, treat its pointer below as an optional cross-reference; this
skill still works standalone.
huggingface-skills:hf-cli โ the transport verbs (hf download --resume, hf upload-large-folder, hf cache verify); this skill owns the China-mirror swap + stall-retry (references/china-network.md).
huggingface-skills:huggingface-trackio โ hosted tracker so metrics survive teardown (gotcha U20); poll trackio alerts as a structured monitor instead of brittle ssh-tail.
superpowers:verification-before-completion โ the Iron Law's general form; gates every "training done / synced / teardown complete" claim.
superpowers:dispatching-parallel-agents โ independence predicate + reconciliation for ablation fan-out.
Getting better over time (capture new gotchas + personalize)
This skill is static, but every run can teach it something โ without corrupting it.
Protocol โ references/self-improvement.md. In short: when a run surfaces a gotcha the catalog
lacks, only sediment a root-caused, reproduced, generalizable one (a one-off flake is a hypothesis,
not a gotcha โ principle #3); route it โ user/project-specific โ the host's memory system,
generalizable โ propose adding to references/gotchas_universal.md / the profile ยง7 /
references/training/ (and offer an upstream PR); never silently rewrite a skill file โ draft the
symptom โ root cause โ fix and let the user approve. On first use, capture the user's platforms +
paths + tracker entity into memory so later runs are pre-parameterized. Platform facts carry a verified <month> stamp โ re-verify any teardown/billing fact against current docs before betting money or data.
Limitations
Does not replace a real cloud orchestrator or managed provisioner; use it to make rented-box work survivable, not to optimize multi-cloud procurement.
Platform billing, stop, destroy, and data-retention behavior can drift; re-check current provider docs before destructive or money-impacting actions.
Requires user-owned credentials, SSH/API access, and explicit confirmation before teardown, deletion, or other irreversible cleanup.
Companion skills named above are not bundled here; treat them as optional references unless installed in the current agent enprojectnment.
Bundled resources
Load only what the current phase needs.
references/principles.md โ the 10 invariants expanded, with the cross-platform nuance behind each.
references/lifecycle_checklist.md โ the 6-phase runbook as a per-platform checklist.
references/gotchas_universal.md โ universal + mixed gotchas (TOC + grep index at top).
references/monitoring_patterns.md โ the four-layer durable-monitoring architecture + robust ssh-poll template.
references/training/ โ the DL-training debug layer (8 files: oom-memory, distributed-launch, precision-stability, throughput-profiling, checkpoint-resume, by-domain, convergence-debugging, data-pipeline) โ see "When training breaks" above.
references/self-improvement.md โ the feedback loop: capture a new gotcha (at a bar) into memory or the catalog, personalize on first run, keep platform facts fresh.
scripts/ โ wrapper templates (run_one/run_queue), monitors (mem_monitor, gpu_health, reap_vram_zombies), the read-only patrol (health_patrol.sh.template), transfer/aggregation (download_loop, aggregate_to_fs, setup-china-mirrors), the load-and-verify checker (verify_local.py), and the verified-stamp freshness linter (check_staleness.py).
profiles/<platform>.md โ the per-platform substrate (one per platform; _schema.md defines the 8 fields).
examples/autodl_sweep/ โ one complete, runnable worked case end to end.