Skip to main content

combined-power-thermal-profiling

Orchestrate a full platform profiling session on an Intel host — apply a power envelope and thermal policy, start the power/thermal monitor, drive a bounded stress load, then summarize the result as a single enclosure report. Chains set-power-profile → set-thermal-profile → monitor-power-thermal → generate-platform-stress and emits one consolidated report (min/mean/max of PkgTmp / PkgWatt / GFXWatt plus throttle/headroom verdict). Ideal for qualifying whether an enclosure can sustain a chosen profile under load before it ships.

소스 정보

저장소
open-edge-platform/edge-node-infrastructure-blueprint
최근 소스 활동
2026년 9월 1일 09:44
감지된 SKILL.md 언어
영어
스타
1
포크
9

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
combined-power-thermal-profiling
description
Orchestrate a full platform profiling session on an Intel host — apply a power envelope and thermal policy, start the power/thermal monitor, drive a bounded stress load, then summarize the result as a single enclosure report. Chains set-power-profile → set-thermal-profile → monitor-power-thermal → generate-platform-stress and emits one consolidated report (min/mean/max of PkgTmp / PkgWatt / GFXWatt plus throttle/headroom verdict). Ideal for qualifying whether an enclosure can sustain a chosen profile under load before it ships.
## Purpose `combined-power-thermal-profiling` is an **orchestrator** skill: it runs the full **apply → monitor → stress → summarize** loop end to end and emits a single enclosure report, instead of requiring the operator to invoke the four power-tuning skills separately and stitch the results together. It does **not** wrap a new script. It sequences the existing reference skills: 1. **CONSTRAIN** — `set-power-profile` (RAPL PkgWatt/SysWatt cap + `intel_lpmd`) and, when requested, `set-thermal-profile` (thermald trip points). 2. **OBSERVE** — `monitor-power-thermal` (turbostat → `pt_mon.txt`), bounded to the stress window. 3. **LOAD** — `generate-platform-stress` (stress-ng CPU + iGPU), bounded by `duration`. 4. **SUMMARIZE** — parse the captured trace and emit one enclosure report: min/mean/max of `PkgTmp` / `PkgWatt` / `GFXWatt`, whether the package power held at the target, and whether the thermal trips engaged / throttling occurred. Each underlying skill still runs its own preconditions, dry-run, confirmation gate, and validation. This orchestrator adds cross-skill sequencing, a single combined confirmation, and the consolidated report. Use it to answer one question: **"can this enclosure sustain profile X under load Y without throttling?"** ### Cross-profile dependency and reboot lifecycle The order is functional, not cosmetic. `set-thermal-profile` reads the live package RAPL PL1 while it generates `thermal-conf.xml`, then embeds that value as thermald's PPCC maximum. At thermald startup, that PPCC value is written back to the package power control. Therefore this skill must always apply **power first, thermal second**. Reapply thermal after every power-profile change so PPCC represents the current power target. The RAPL PL1/PL2 caps are volatile and reset to firmware defaults on reboot. The `intel_lpmd` configuration, thermald configuration, and thermald systemd drop-in remain on disk. After every reboot, re-run this skill (or manually apply power then thermal in that order) before a qualified workload; this reinstates the RAPL cap and refreshes thermald's persisted PPCC maximum. ## Terminology Acronyms and terms used throughout this skill. See the underlying skills for the full glossaries. | Term | Meaning | |---|---| | enclosure | The physical chassis / thermal environment being qualified (fanless kiosk, sealed edge box, etc.). | | profiling session | One apply → monitor → stress → summarize pass at a fixed power (and thermal) profile. | | power profile | The PkgWatt/SysWatt envelope applied by `set-power-profile` (LowPower … MaxPerformance or Custom). | | thermal profile | The thermald trip points applied by `set-thermal-profile` (cool/warm/hot/thermal-max or custom). | | PkgTmp / PkgWatt / GFXWatt | Package temperature / package power / integrated-GPU power sampled by the monitor. | | headroom | How far the measured PkgWatt/PkgTmp sits below the profile target / thermal trips under sustained load. | | throttle | Firmware clamps sustained power below the target, or the platform holds at a passive trip — the enclosure cannot cool the profile. | | enclosure report | The single consolidated report this skill emits at the end of the session. | | dry_run | Preview mode: resolve and show the full plan for all stages without applying, monitoring, or stressing. | ## Trigger Phrases - profile the enclosure / profile this platform - run a full profiling session - qualify this enclosure for <profile> - apply, monitor, stress and summarize - can this chassis sustain <profile> under load - run the power/thermal profiling loop - combined power profiling - burn-in and report / thermal qualification run - one-shot power and thermal profiling ## Required Inputs - enib_home: absolute path to this repository root (default: current workspace root). On a host provisioned with Infrastructure Blueprint, the developer source tree lives at `/opt/edge/developer`, so `enib_home` is `/opt/edge/developer` on the target system. - profile: power profile to apply — one of `LowPower`, `BalancedLow`, `BalancedHigh`, `Performance`, `MaxPerformance`, or `Custom` (default: `BalancedHigh`). Passed to `set-power-profile`. - pkg_watt / sys_watt / burst_ratio / pl1_tau: optional power-envelope overrides, forwarded verbatim to `set-power-profile` (same rules/validation as that skill; `pkg_watt` only for `Custom`). - thermal_profile: thermal policy to apply — one of `cool`, `warm`, `hot`, `thermal-max`, `custom`, or `none` (default: `none`). `none` is allowed only when thermald is inactive or disabled; an active thermald service must receive an explicit profile after every power-profile change so its PPCC maximum is refreshed. When `custom`, also supply `fan_c`/`proc_c`/`clamp_c`. Passed to `set-thermal-profile`. - fan_c / proc_c / clamp_c: custom thermal trip points (only when `thermal_profile=custom`), forwarded to `set-thermal-profile`. - duration: bounded stress/monitor window in stress-ng time syntax, e.g. `60s`, `3m` (default: `3m`). **Required to be bounded** — an open-ended session is rejected (see Safety Rules). - cpus / load / gpu: stress parameters forwarded to `generate-platform-stress` (defaults: all CPUs, `100`%, `12` GPU workers). - interval: monitor sampling interval in seconds (default: `2`), forwarded to `monitor-power-thermal`. - log_path: where the monitor trace is written (default: `<enib_home>/tools/power-tuning/pt_mon.txt`). - dry_run: `true` | `false` (default: `false`). When `true`, every stage runs its own dry-run only; nothing is applied, monitored, or stressed. - auto_confirm: `true` | `false` (default: `false`). When `true`, skip the single combined confirmation gate and each sub-skill's gate. ## Preconditions Run silently without user prompts. This orchestrator's preconditions are the **union** of the sub-skills' preconditions; delegate to each and aggregate. - [ ] This skill file exists and is readable: - `test -f <enib_home>/skills/combined-power-thermal-profiling/SKILL.md` - [ ] All four sub-skill files exist and are readable: - `test -f <enib_home>/skills/set-power-profile/SKILL.md` - `test -f <enib_home>/skills/set-thermal-profile/SKILL.md` (only when `thermal_profile != none`) - `test -f <enib_home>/skills/monitor-power-thermal/SKILL.md` - `test -f <enib_home>/skills/generate-platform-stress/SKILL.md` - [ ] All four reference scripts exist and are executable: - `test -x <enib_home>/tools/power-tuning/set_power_profile.sh` - `test -x <enib_home>/tools/power-tuning/set_thermal_profile.sh` (only when `thermal_profile != none`) - `test -x <enib_home>/tools/power-tuning/pt_mon.sh` - `test -x <enib_home>/tools/power-tuning/stress_gen.sh` - [ ] Required tools present: `command -v turbostat`, `command -v stress-ng`, `command -v rdmsr && command -v wrmsr`, and (when `thermal_profile != none`) `test -x /usr/sbin/thermald`. On any miss, stop with the same install hint the owning sub-skill gives. - [ ] No stress-ng or turbostat instance is already running (would skew the capture): - `pgrep -x stress-ng` and `pgrep -x turbostat` — if either returns a PID, stop and instruct the user to stop it first (`sudo pkill -x stress-ng` / `sudo pkill -x turbostat`) before re-triggering. - [ ] When `thermal_profile=none`, verify thermald is not active: `systemctl is-active --quiet thermald`. If it is active, stop before applying power and require the user to select an explicit thermal profile. thermald owns the package RAPL cooling device and can restore its persisted PPCC maximum, making an unchanged thermal policy incompatible with a new power target. - [ ] **Sudo probe (MANDATORY unless `dry_run=true`).** The session applies power/thermal changes and runs `sudo turbostat`. Run `sudo -n true`; if exit is non-zero, do NOT proceed — stop and instruct the user to run `sudo -v` (or add the scoped `NOPASSWD` entries the sub-skills document), then re-trigger. Never collect a password via prompts, env vars, scripts, or logs. See [AGENTS.md](../../AGENTS.md#sudo-handling-must-follow-for-all-skills-that-invoke-sudo). - [ ] Host is x86_64 with an Intel CPU (sanity check; non-fatal warning if not): `uname -m` and `grep -m1 -o 'GenuineIntel' /proc/cpuinfo`. - [ ] (Informational) Detect psys/SysWatt support so the report can annotate a `0.00` reading (as in `monitor-power-thermal`). Prompt only for missing required inputs: - [ ] Do not prompt when `profile`, `thermal_profile`, or the stress/monitor knobs are omitted — use the defaults above. - [ ] Only for `thermal_profile=custom`: if any of `fan_c`/`proc_c`/`clamp_c` is missing, ask for the three trip points (required by `set-thermal-profile`). Input validation (fail closed before running anything): - [ ] `profile` and the power overrides validate against `set-power-profile`'s rules; `thermal_profile` and any custom trips validate against `set-thermal-profile`'s rules; `cpus`/`load`/`gpu` against `generate-platform-stress`'s ranges; `interval` is a positive integer. - [ ] `duration` matches `^[0-9]+(s|m|h)?$` **and is present** — a bounded window is mandatory for this orchestrator. ## Steps **Terminal command rules (MUST follow for every command, inherited by every stage):** - Always invoke scripts by **absolute path** — never prefix with `cd`. - Never combine `cd` with any output redirection (`>`, `>>`, `2>`, `2>&1`, `| tee`) in the same compound command — VS Code blocks it with an approval dialog. - Never use `$(...)` command substitution in terminal commands — VS Code blocks them with an approval dialog. The scripts handle all internal computation themselves. 1. **Resolve the full session plan (no writes yet).** Build the argument lines for all four stages from the resolved inputs and render a single **Planned Session** table: power profile + envelope, thermal profile + trips (or `none` with thermald confirmed inactive), monitor interval/duration/log path, and stress cpus/load/gpu/duration. 2. **Dry-run every mutating stage** (read-only, no sudo). Run `set-power-profile` with `--dry-run`, and `set-thermal-profile` with `--dry-run` when `thermal_profile != none`; capture each resolved plan verbatim and fold any firmware/cTDP clamp or thermal notes into the Planned Session table. 3. **Single combined confirmation gate** — pause before any write: - If `dry_run=true`: stop here and record `CONFIRMATION=dry_run_only`. Do not apply/monitor/stress. - Else if `auto_confirm=true`: log `AUTO_CONFIRM=true`, propagate `auto_confirm=true` to each sub-skill, and continue. - Else: present the Planned Session table and ask **once**: "Run the full profiling session — apply <profile> (+ <thermal_profile> thermal), monitor for <duration>, and stress <cpus> CPUs @ <load>% + <gpu> GPU workers on this host? (yes/no)". On anything other than `yes`/`y`, stop and record `CONFIRMATION=declined`. A `yes` authorizes all stages; still surface (do not re-prompt for) each sub-skill's plan as it runs. 4. **CONSTRAIN — apply the envelope** (only after confirmation), in order: - Invoke **set-power-profile** with the resolved `profile` and any `pkg_watt`/`sys_watt`/`burst_ratio`/`pl1_tau`, `auto_confirm=true`. Capture its report and exit code. Abort the session if it fails. - If `thermal_profile != none`: invoke **set-thermal-profile** immediately after power with the resolved `thermal_profile` (+ custom trips / `--charge` if given), `auto_confirm=true`. It captures the newly applied RAPL PL1 as thermald PPCC. Capture its report and exit code. Abort if it fails. 5. **OBSERVE — start the monitor**, bounded to the stress window. Invoke **monitor-power-thermal** with `duration` (≈ the stress `duration`, plus a few seconds of lead/tail), `interval`, and `log_path`. For non-default interval, duration, or log path, that skill runs `turbostat` directly rather than passing unsupported options to `pt_mon.sh`. Start it **before** the stress load so the trace captures the ramp. Record the log path and PID. 6. **LOAD — drive the stress**, synchronously and bounded. Invoke **generate-platform-stress** with the resolved `cpus`/`load`/`gpu` and the bounded `duration`, `auto_confirm=true`. Let it complete; capture its report (including the GPU-load verification when `gpu > 0`) and exit code. 7. **Stop the monitor** if it is still running (bounded runs self-terminate; otherwise `sudo pkill -x turbostat`). Confirm the trace file is non-empty (`test -s <log_path>`). 8. **SUMMARIZE — build the enclosure report.** Parse `<log_path>` for the min/mean/max of `PkgTmp`, `PkgWatt`, and `GFXWatt` over the stress window. Compare against the applied profile target and the thermal trips to derive the verdict (see Expected Result Summary). Emit the single consolidated report. ## Validation Validation section is criteria-only. Do not render the pass/fail results table here. - Preconditions passed (all required sub-skill files/scripts executable; tools present; no pre-existing stress-ng/turbostat; sudo probe = 0 when a run is intended). - Inputs validated against each sub-skill's rules; `duration` present and bounded. - A single Planned Session table was rendered from the stage dry-runs before the combined confirmation gate. - Confirmation gate outcome recorded as one of: `confirmed`, `auto_confirm`, `declined`, `dry_run_only`. - Stages executed only when the outcome is `confirmed` or `auto_confirm`, and in order: power → thermal → monitor → stress → summarize. Thermal is omitted only when thermald was confirmed inactive for `thermal_profile=none`. - Each executed sub-skill reported success (exit `0`); the session aborted (and rolled back per Rollback) on the first failure. - The monitor trace is non-empty and covers the stress window; the enclosure report's min/mean/max were parsed from it. - `SysWatt=0.00` is annotated as a known firmware limitation (per psys detection), NOT a session failure. ## Rollback - The session is composed of the sub-skills' own reversible actions; roll back in reverse order of application. - **Stress** leaves no persistent state — if aborted, stop it: `sudo pkill -x stress-ng`. - **Monitor** is read-only — stop it (`sudo pkill -x turbostat`); the only artifact is the trace at `<log_path>`. - **Power profile**: the RAPL cap is runtime-only (reverts on reboot); to revert immediately re-run `set-power-profile` with a lower profile, or restore the `intel_lpmd` `.orig` config (see that skill's Rollback). - **Thermal profile** (if applied): re-run `set-thermal-profile` with a different profile (previous config backed up to `.bak`), restore the `.bak`, or run it with `disable=true` to return to kernel default thermal control. This config persists across reboot. Do not restore a prior thermald configuration after applying a new power target without reapplying the thermal profile, because its saved PPCC maximum may be stale. - To restore the original local settings rather than only the previous profile, use configuration copies captured before the first apply. The thermal `.bak` files are replaced on later applies, and the power script creates `.orig` only for model-specific `intel_lpmd` files it replaced; a newly created generic `intel_lpmd_config.xml` has no automatic original backup. - After a reboot, do not assume the persisted thermald PPCC still represents the active power cap: reapply power first, then thermal, before the next profile run. - If any stage fails mid-session, stop the already-started monitor/stress, report which stages applied, and propose the matching rollback for each applied stage. ## Safety Rules - **Bounded only.** Refuse to run with an open-ended (missing/zero) `duration` — an orchestrated apply-and-load session must self-terminate. Direct the user to the individual skills for open-ended runs. - Never collect a sudo password and never prompt for sudo approval; passwordless sudo is expected to be pre-configured for the underlying scripts/`turbostat`. Never collect a password via prompts, env vars, scripts, or logs. - **Never combine `cd` with output redirection** and **never use `$(...)`** in terminal commands — VS Code blocks both with approval dialogs. - Warn before a session that combines a high power profile (`MaxPerformance`) or a hot thermal profile (`thermal-max`) with full-load stress on a thermally constrained / fanless enclosure — this is exactly the case that can throttle or overheat; recommend starting cooler and stepping up. - Do not stack stressors or monitors: honor the pre-existing-instance precondition. - Delegate all writes to the sub-skills; this orchestrator does not modify anything outside what those skills manage (`tools/power-tuning/`, the `intel_lpmd` config, `/etc/thermald/`, and the monitor trace). - Confirm once before the whole session; a single `yes` authorizes all stages. Do not silently apply a profile the user only asked to preview. ## Expected Result Summary Emit the single **enclosure report** as the following tables. ### Session Metadata | Field | Value | |---|---| | Preconditions | PASS/FAIL | | Host | `<uname -m>` + CPU model name | | Enclosure / label | `<user-supplied or 'unspecified'>` | | Power profile | `<profile>` (+ any `pkg_watt`/`sys_watt`/`burst_ratio`) | | Thermal profile | `<thermal_profile>` (Fan/Proc/clamp °C), or `none` (thermald inactive) |
GitHub에서 보기
이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기