| name | hyperloom-qwen3-8b-3h |
| description | Run a 3-hour Hyperloom Qwen3-8B optimization session without the Kernel Agent. Use when the user wants a short, no-kernel Hyperloom demo on the local AMD ROCm environment. |
Hyperloom Qwen3-8B 3h No-Kernel Run
Read .env first and resolve HYPERLOOM_SKILL_PATH. Read and follow the optimizer skill at @${HYPERLOOM_SKILL_PATH} before launching. If HYPERLOOM_SKILL_PATH is missing, fall back to @hyperloom/inference_optimizer/SKILL.md (wheel install) or @src/hyperloom/inference_optimizer/SKILL.md (source checkout). This skill provides the concrete workload and launch constraints for a short Qwen3-8B demo.
Run Mode
Resolve the run mode before launching Hyperloom:
- If
HYPERLOOM_RUN_MODE=baremetal or it is unset, run this demo directly on the host.
- If
HYPERLOOM_RUN_MODE=docker, this skill owns the Docker setup. Ask the user whether they want a vllm or sglang Docker image unless HYPERLOOM_IMAGE is already set. Use a ROCm image that already contains the selected framework; do not install the framework inside Docker.
In docker mode:
- If
hyperloom-setup already ran, do not re-run setup on the host.
- Read
HYPERLOOM_DOCKER_TARGET_HOST from .env when present. If it names a
host different from $(hostname), first SSH to that host and continue this
Docker setup there; do not start Docker on the login/current host.
- Always run setup inside the container after
docker run.
- Pass
--install-framework none --yes in the container (ROCm/framework comes from
the image). Do not use --skip-base-check — let Phase 1 preflight validate
the container environment.
- Do not run
python -m hyperloom.inference_optimizer.cli optimize on the host.
Suggested Docker images:
vllm: docker.io/vllm/vllm-openai-rocm:v0.27.1
sglang MI300X: docker.io/lmsysorg/sglang-rocm:v0.5.18-rocm724-mi30x-20260825
sglang MI355X: docker.io/lmsysorg/sglang-rocm:v0.5.18-rocm724-mi35x-20260825
In Docker mode, start a long-running container on HYPERLOOM_DOCKER_TARGET_HOST
(or the current host when it is unset) before running setup or optimize:
export REPO_ROOT="$(pwd -P)"
docker run -d \
--name "${HYPERLOOM_CONTAINER_NAME:-hyperloom-local}" \
--shm-size "${HYPERLOOM_SHM_SIZE:-64g}" \
--entrypoint tail \
--device /dev/kfd \
--device /dev/dri \
--group-add video \
-v "$REPO_ROOT:$REPO_ROOT" \
"$HYPERLOOM_IMAGE" \
-f /dev/null
Mount the Hyperloom workspace at the same absolute path (-v "$REPO_ROOT:$REPO_ROOT") so paths in .env, logs, and session artifacts stay valid. If USER_DATA_PATH or a pre-downloaded model directory is outside the workspace, add matching -v host_path:host_path mounts before starting the container.
Then run the setup backend inside the container:
docker exec -w "$REPO_ROOT" "${HYPERLOOM_CONTAINER_NAME:-hyperloom-local}" bash -lc \
'REPO_ROOT="$(pwd -P)"; PYTHONPATH="$REPO_ROOT" python3 -m hyperloom.inference_optimizer.setup -- --install-framework none --yes'
After that, run all remaining commands for this demo inside the same container with docker exec -w "$REPO_ROOT" ...; do not run python -m hyperloom.inference_optimizer.cli optimize on the host in Docker mode. When the demo is finished, ask the user whether to stop the container. If they say yes, run:
docker stop "${HYPERLOOM_CONTAINER_NAME:-hyperloom-local}"
Environment
-
MODEL_PATH=<optional; if unset, download Qwen/Qwen3-8B from Hugging Face with the Python steps below, then set MODEL_PATH to that local path>
-
FRAMEWORK=<provided by the existing environment or repository-root .env; do not invent it>
-
GPU_TYPE=<do not set; omit --gpu-type and let Hyperloom auto-detect from ROCm/system info>
Required optimize CLI flags:
-
--tp 1
-
--conc 64
-
--isl 1024
-
--osl 1024
-
--precision bf16
-
--target-gain 30
-
--max-hours 3
-
--max-minutes-framework-pct 0.44
-
--max-minutes-sweep-pct 0.01
-
--no-framework-agent
-
--no-kernel
-
--no-enable-conc-sweep
-
--no-enable-roofline
Before launch, read the repository-root .env file if it exists and load the needed environment variables from it, such as LLM API keys/base URLs, FRAMEWORK, and HF_TOKEN. Do not copy secret values into the prompt, terminal output, reports, or logs. Do not modify USER_DATA_PATH.
Before resolving or downloading any model, always ask the user which model path to use. Present the currently resolved option when MODEL_PATH is already set, and always offer a custom local path plus the demo default. Do not continue until the user chooses one.
Use this decision flow:
- If the user chooses the existing
MODEL_PATH, inspect that path and use it only when it contains config.json; otherwise ask again for a valid path or the demo default.
- If the user provides a custom local path, export
MODEL_PATH to that path and require config.json before launch.
- If the user chooses the demo default, set
MODEL_PATH=${REPO_ROOT}/.cache/hyperloom-models/Qwen3-8B and download Qwen/Qwen3-8B there when config.json is not already present.
Do not assume the Hugging Face CLI exists; resolve or download the selected model with Python:
python -m pip install -U huggingface_hub
export REPO_ROOT="$(pwd -P)"
export MODEL_PATH="${MODEL_PATH:-${REPO_ROOT}/.cache/hyperloom-models/Qwen3-8B}"
python - <<'PY'
import os
from pathlib import Path
from huggingface_hub import snapshot_download
target = Path(os.environ["MODEL_PATH"]).expanduser()
if (target / "config.json").is_file():
print(f"Using existing model at {target.resolve()}")
else:
snapshot_download(
repo_id="Qwen/Qwen3-8B",
local_dir=str(target),
local_dir_use_symlinks=False,
)
print(target.resolve())
PY
Pre-launch Runtime Install
Before the first optimize launch, run the full runtime installer in the same
environment that will launch the optimizer. This is required even for this
--no-kernel demo: preflight loads kernel-agent.env.sh before it can reach the
later Ray/Magpie/InferenceX auto-install checks.
For Docker mode, run this inside the container. For bare-metal mode, run it on
the host:
export REPO_ROOT="$(pwd -P)"
_dotenv_prev="$(export -p | grep -v -e '=""$' -e "=''\$")"
set -a; . "${REPO_ROOT}/.env"; set +a
eval "$_dotenv_prev"
unset _dotenv_prev
export USER_DATA_PATH="${USER_DATA_PATH:?USER_DATA_PATH missing}"
export PYTHONPATH="${REPO_ROOT}:${PYTHONPATH:-}"
ulimit -Sn 65536 || true
INSTALL_SH="${REPO_ROOT}/hyperloom/inference_optimizer/assets/install.sh"
if [ ! -f "$INSTALL_SH" ]; then
INSTALL_SH="${REPO_ROOT}/src/hyperloom/inference_optimizer/assets/install.sh"
fi
bash "$INSTALL_SH"
. "$USER_DATA_PATH/runtime/kernel-agent.env.sh"
export PYTHONPATH="${REPO_ROOT}:${PYTHONPATH:-}"
If hyperloom/inference_optimizer/assets/install.sh is not present (source
checkout layout), use src/hyperloom/inference_optimizer/assets/install.sh.
User-visible Progress
Keep the user informed with concise status updates throughout the demo. Do not
dump full debug logs into chat; report the important values and paths so the user
can tell that work is progressing.
Before launch, report the launch plan:
- model path and whether it is an existing local model or a downloaded default;
- run mode (
baremetal or docker) and target host/container when applicable;
- framework, TP, concurrency, ISL, OSL, precision, max hours, and required demo
flags;
USER_DATA_PATH and where runtime artifacts will be written.
After the runtime install, report whether it succeeded and the path to
kernel-agent.env.sh. After starting the optimizer, report:
- optimizer PID;
- run log path;
- launch-info JSON path;
- resolved session directory;
state.json path;
- initial health check result.
During monitoring, print a short summary at each 300-second check:
- process alive/stopped;
- phase and
stop_reason;
- baseline throughput, current best throughput, and cumulative gain when present;
- latest benchmark result or candidate decision when available;
- the most relevant recent log lines, excluding secrets.
When the run finishes, report the final status, final report path, best result,
and the stop reason. Never print API keys, tokens, or custom header values.
Launch Requirements
- Run the pre-launch runtime install above and source
$USER_DATA_PATH/runtime/kernel-agent.env.sh before launching.
- Keep
PYTHONPATH="$REPO_ROOT:${PYTHONPATH:-}" in the launch shell so robustness
and critic subprocesses can import hyperloom.agents after changing cwd.
- Run in background with
setsid nohup.
- Pass all required optimize CLI flags in the
python -m hyperloom.inference_optimizer.cli optimize command. Do not rely on .env alone for TP, CONC, ISL, OSL, or PRECISION; CLI defaults can otherwise override the intended workload.
- Include
--max-minutes-framework-pct 0.44 and --max-minutes-sweep-pct 0.01
in the optimize command. With FRAMEWORK_AGENT and KERNEL_AGENT disabled,
Hyperloom redistributes their shares so most of the short run budget is
reserved for OPTIMIZE while still leaving SWEEP/CLOSE time to exit cleanly
near the deadline.
- Include
--no-framework-agent in the optimize command so the
FRAMEWORK_AGENT phase is skipped.
- Include
--no-kernel in the optimize command so the Kernel Agent phase is skipped.
- Include
--no-enable-conc-sweep in the optimize command so the SWEEP-phase post-optimization concurrency sweep is skipped.
- Include
--no-enable-roofline in the optimize command so PRELUDE uses the lighter profile path instead of roofline analysis.
- Report the session ID, log path, PID, and initial health check result.
- Monitor the process every 300 seconds until work is done.
- To recover an unexpected crash, only run
optimize --resume-from "$SESSION_DIR" against the same session dir. After the first launch, never start a new optimize; that creates a new <UTC_ts> session and is forbidden.
- If
stop_reason in the current session state.json is final, stop and exit.