Skip to main content

debug-openshell-cluster

Debug why an OpenShell gateway deployment is unhealthy, unreachable, or unable to create sandboxes. Use for gateway health failures, Docker/Podman runtime issues, Helm failures, Kubernetes scheduling, TLS or auth, gateway interceptors, supervisor middleware startup or runtime failures, external compute-driver sockets, VM drivers, or sandbox startup. Trigger keywords - debug gateway, gateway failing, deployment failing, helm install failing, cluster health, gateway health, gateway not starting, health check failed, sandbox pending, docker driver, podman driver, kubernetes driver, external driver, compute driver socket, gateway interceptor, supervisor middleware, middleware failed, vm driver.

Informations de source

Dépôt
NVIDIA/OpenShell
Dernière activité de la source
26 septembre 2026 à 01:50
Langue détectée de SKILL.md
anglais
Étoiles
8 810
Forks
1 280

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Explorateur de fichiers
2 fichiers

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
debug-openshell-cluster
description
Debug why an OpenShell gateway deployment is unhealthy, unreachable, or unable to create sandboxes. Use for gateway health failures, Docker/Podman runtime issues, Helm failures, Kubernetes scheduling, TLS or auth, gateway interceptors, supervisor middleware startup or runtime failures, external compute-driver sockets, VM drivers, or sandbox startup. Trigger keywords - debug gateway, gateway failing, deployment failing, helm install failing, cluster health, gateway health, gateway not starting, health check failed, sandbox pending, docker driver, podman driver, kubernetes driver, external driver, compute driver socket, gateway interceptor, supervisor middleware, middleware failed, vm driver.
# Debug OpenShell Gateway Deployment Diagnose a gateway and its selected compute platform. Do not assume OpenShell provisions Kubernetes or runs a k3s container. OpenShell targets a reachable gateway endpoint backed by Docker, Podman, Kubernetes, the experimental VM driver, or an operator-managed out-of-tree compute driver. Use `openshell` first to identify the active endpoint. Then use the platform tools that match the gateway's compute driver: `docker`, `podman`, `kubectl`/`helm`, or VM driver logs. ## Overview The target deployment flow is: 1. Operator starts or deploys the gateway with system packages, systemd, or Helm. The CLI does not start, stop, or destroy gateway services. 2. Operator configures the compute driver. 3. Operator provides the CLI and supervisor authentication material required by the deployment mode: edge or OIDC user auth, optional CLI mTLS, and gateway-minted sandbox JWTs. 4. The CLI registers a reachable gateway endpoint with `openshell gateway add`. 5. The gateway creates sandboxes through the selected compute driver. The `openshell-gateway` composition crate explicitly installs its compiled Docker, Podman, Kubernetes, and VM registrations at startup; `openshell-server` does not link compute-driver crates. Custom gateway binaries may include a subset of those registrations. With no configured driver, the gateway probes only installed registrations in priority order (Kubernetes, Podman, then Docker); VM has no probe and remains opt-in. Confirm the binary's registered drivers when auto-detection reports that no suitable driver is available. If configuration selects a driver that was not compiled in, the gateway treats the name as an external driver and reports a missing `socket_path` unless an endpoint is configured. On Windows, custom binaries can include MXC independently. Registrations for Docker, Podman, Kubernetes, and VM are rejection stubs when included; they do not enable those runtimes on Windows. See the [compute driver reference](https://docs.nvidia.com/openshell/latest/how-it-works/sandboxes/runtimes) for selective-build options and external-driver configuration. For local evaluation only, TLS may be disabled and the gateway can be reached through `http://127.0.0.1:<port>`. ## Prerequisites - The `openshell` CLI must be available for endpoint checks. - Know the active gateway name and endpoint, or be able to inspect local gateway metadata. - Know the compute platform: Docker, Podman, Kubernetes, VM, or an out-of-tree driver. - For Kubernetes: `kubectl` must target the cluster that hosts OpenShell and Helm version 3 or later must be available. - For Docker or Podman: the runtime socket must be reachable from the gateway host. Use `openshell --help` and nested `--help` output as the authority for the installed CLI version. Use the published [installation guide](https://docs.nvidia.com/openshell/latest/about/installation.md), [compute-driver reference](https://docs.nvidia.com/openshell/latest/how-it-works/sandboxes/runtimes), [gateway configuration reference](https://docs.nvidia.com/openshell/latest/how-it-works/gateways/configuration), and [Kubernetes setup guide](https://docs.nvidia.com/openshell/latest/kubernetes/setup.md) as the authority for deployment and configuration behavior. ## Workflow Run diagnostics in order and stop once the root cause is clear. ### Step 1: Check CLI Reachability ```bash openshell gateway list --output json openshell gateway info openshell status ``` For a one-off endpoint check that bypasses stored gateway selection and metadata: ```bash openshell --gateway-endpoint <url> status ``` Common findings: - `No active gateway`: register one with `openshell gateway add <endpoint>`. - Connection refused: gateway process is not running, service exposure is wrong, or a port-forward/proxy is not active. - TLS/certificate errors: the endpoint scheme or trust chain is wrong, a local mTLS bundle does not match the gateway CA, or TLS termination does not match the gateway listener. - A Snap refresh restarts the gateway with its migrated mTLS config. The secure Snap gateway uses `https://127.0.0.1:17670` and requires a client bundle in the user's Snap state. Refresh replaces insecure configs without keeping a copy; follow the published Snap installation steps to re-register an old HTTP client. - `Unauthenticated` from an edge or OIDC gateway: refresh stored credentials with `openshell gateway login [name]`, then retry. Use `gateway logout` only when intentionally clearing local credentials. - A direct development endpoint with a private or self-signed certificate can be isolated with `--gateway-endpoint <url> --gateway-insecure`; do not persist or recommend insecure verification for shared gateways. ### Step 2: Identify the Compute Platform Use gateway metadata, deployment values, or the user's setup notes to identify the driver. | Platform | Primary checks | |---|---| | Docker | Gateway process logs, Docker daemon health, sandbox containers, image pulls. | | Podman | Podman socket, rootless networking, sandbox containers, image pulls. | | Kubernetes | Helm release, gateway workload, service, secrets, sandbox pods, events. | | OpenShift | Same as Kubernetes, plus SecurityContextConstraints (SCCs) and, for external access, an OpenShift `Route`. Detect OpenShift by the presence of the `route.openshift.io` API group (`oc api-resources --api-group=route.openshift.io`). | | VM | VM driver logs, rootfs availability, host virtualization support. | | Extension | External driver process, Unix socket ownership/mode, configured driver name, capability handshake, gateway logs. | ### Step 3: Check Gateway Startup Dependencies Before debugging the compute platform, inspect gateway logs for failures in dependencies initialized before the listener becomes ready. For resource-admission failures, distinguish disabled caller driver config from missing resource approval. Helm defaults `server.drivers.kubernetes.allowDriverConfig` to false and `resourceAdmission.enabled` to true. Existing PVCs, RuntimeClasses, and PriorityClasses need matching administrator-owned labels; namespace membership and read-only access do not grant approval. GPU devices and operator-selected image-pull Secrets do not need admission labels. In managed mode, inspect the configured source image-pull Secret in the gateway namespace and the generation copies (`openshell.ai/component=image-pull`) in the workspace namespace. Legacy workloads without admission provenance need recreation. Do not automatically label control-plane resources or disable enforcement as a repair. For out-of-tree compute drivers, also check that their versioned admission-policy acknowledgement matches the gateway's policy. Configure standalone driver policy through its administrator-owned `--admission-config-json` option. For out-of-tree compute drivers, confirm the selected driver name and socket agree across CLI flags or `gateway.toml`, and that the operator-owned driver is running before the gateway starts: ```bash rg -n '^version|compute_driver|socket_path|guest_tls_' /etc/openshell/gateway.toml stat /run/openshell/<driver>.sock journalctl -u <driver-service> --no-pager --lines=200 journalctl -u openshell-gateway --no-pager --lines=200 ``` Gateway configuration requires `[openshell] version = 2`, a singular `compute_driver` selector, and driver-owned settings under `[openshell.drivers.<name>]`. The gateway rejects legacy `compute_drivers` and `--drivers` selectors rather than silently migrating them. One valid, nonempty `OPENSHELL_DRIVERS` value remains a deprecated environment-only alias when the canonical selector is absent; the gateway selects that driver with a warning. Empty, invalid, comma-delimited, or conflicting values fail startup. The WebSocket tunnel for edge-proxy CLI access is off unless `enable_websocket_tunnel = true` (`server.enableWebsocketTunnel` in Helm). Homebrew and RPM package startup migrates only exact package-generated v1 defaults. If an upgraded package still reports an unsupported version, inspect the active prefix or `~/.config/openshell/gateway.toml`; an edited v1 file must follow the published schema-v2 migration steps and must not be overwritten. Guest TLS CA, certificate, and key paths are the exception to driver ownership: configure the complete bundle under `[openshell.gateway]`, and the gateway injects it only into the selected local driver. TLS-enabled Docker, Podman, and VM drivers fail startup when neither those paths nor the package-managed local bundle is available; Kubernetes projects its bundle through a Secret. Custom names use `[openshell.drivers.<name>].socket_path`. A launch-time `--compute-driver-socket` override may also use `docker`, `podman`, `kubernetes`, or `vm`; the endpoint then takes precedence over built-in construction. First-party standalone drivers require the socket parent directory to be owned by the driver's effective UID, force its mode to `0700`, create the socket with mode `0600`, and accept only peers with that same UID. Check the parent and socket separately with `stat`; a gateway running under a different UID cannot connect even when filesystem permissions or group membership would otherwise allow it. Operator-supplied drivers must provide equivalent access control appropriate to their implementation. Check gateway logs for connection errors, `GetCapabilities` failures, missing peer metadata, protocol-major mismatch, unmet required capabilities, or an unexpected advertised driver name. `openshell gateway info` reports successful startup negotiations. The advertised name is diagnostic metadata; negotiated features control optional behavior. The gateway does not create or supervise operator-supplied driver processes or sockets. For the Kubernetes Secrets credential driver, every provider credential lives in the configured `namespace`, in every workspace mode. A `PermissionDenied` error naming another namespace means the provider's credential handle points outside the configured namespace; recreate the provider. An `unknown field` startup error for `[openshell.credential_drivers.kubernetes-secrets]` means the table sets a key the driver does not accept. Confirm the gateway can reach the credential namespace: ```bash kubectl -n openshell get configmap openshell-config -o jsonpath='{.data.gateway\.toml}' | grep -A3 '^\[openshell\.credential_drivers\.kubernetes-secrets\]' kubectl auth can-i get secrets -n <credential-namespace> --as system:serviceaccount:openshell:openshell ``` For a configured Vault credential driver, inspect its endpoint and trust bundle before debugging provider resolution. Non-loopback addresses must use HTTPS, and the driver never follows redirects. A private CA bundle augments platform roots but does not disable hostname verification. With Helm, `server.credentialDrivers.vault.caConfigMapName` names a ConfigMap whose `ca.crt` key is mounted at `/etc/openshell-tls/vault-ca/ca.crt`: ```bash kubectl -n openshell get configmap openshell-config -o jsonpath='{.data.gateway\.toml}' | grep -A10 '^\[openshell\.credential_drivers\.vault\]' kubectl -n openshell get pod -l app.kubernetes.io/name=openshell -o jsonpath='{range .items[0].spec.containers[0].volumeMounts[*]}{.name}{" "}{.mountPath}{"\n"}{end}' | grep vault-ca kubectl -n openshell get configmap <vault-ca-configmap> -o jsonpath='{.data.ca\.crt}' | openssl x509 -noout -subject -issuer -dates kubectl -n openshell logs statefulset/openshell -c openshell-gateway --tail=200 ``` An HTTP service DNS address fails configuration validation. `UnknownIssuer` or an invalid CA error means the ConfigMap is missing, the `ca.crt` key is wrong, or the bundle does not contain the Vault server's issuer. A hostname mismatch means the HTTPS `address` host is absent from the server certificate SANs; keep verification enabled and issue a certificate for the service DNS name. For configured gateway interceptors, inspect `[[openshell.gateway.interceptors]]`, their Unix or network endpoints, and gateway startup logs: ```bash rg -n 'interceptors|provider_profile_sources|grpc_endpoint|tls_ca_cert_path|audience|allow_insecure_transport|binding_policy|failure_policy|gateway_jwt' /etc/openshell/gateway.toml stat /run/openshell/interceptors/<name>.sock journalctl -u <interceptor-service> --no-pager --lines=200 journalctl -u openshell-gateway --no-pager --lines=200 ``` The gateway calls each interceptor's `Describe` RPC and validates its manifest at startup. Check for missing peer metadata, protocol-major mismatch, unmet required capabilities, unreachable endpoints, invalid RPC/phase bindings, strict `allowlist` or `exact` mismatches, and `post_commit` bindings that resolve to `fail_closed`. If gateway JWT signing is enabled, authenticated network interceptors require HTTPS and a valid bearer token; check the private CA path, endpoint hostname, expected audience, issuer, `kid`, and interceptor logs for token rejection. `allow_insecure_transport = true` explicitly preserves unauthenticated plaintext behavior. If `provider_profile_sources` names an interceptor, that interceptor must advertise provider-profile capability and return a valid, duplicate-free catalog. A selected interceptor-only source is authoritative; include a `user` source explicitly when composition is intended. The `builtin` source type was removed: a config that still names it is rejected at startup. If the deployment uses supervisor middleware, follow the [supervisor middleware troubleshooting reference](references/supervisor-middleware.md) for startup, authentication, policy validation, and HTTP or WebSocket failures. For network policy validation failures, first distinguish a gateway mutation rejection from a supervisor runtime rejection. Direct policy updates, incremental merges and approvals, provider attachments, and provider-profile fanout are validated against the complete effective policy before persistence when the gateway knows the affected sandbox scope. A `FAILED_PRECONDITION` ambiguity response means no invalid revision or partial fanout was stored. Supervisor validation remains defense in depth for startup, races, and policy sources outside those mutation paths. Runtime rejection behavior is configured only in `gateway.toml`: ```toml [openshell.gateway] policy_validation_failure_mode = "fail_closed" ``` The default `fail_closed` mode deactivates the previous generation, closes pinned relays, and quarantines new egress until a valid generation loads. `retain_last_valid` explicitly keeps the previous valid policy active; without one it still fails closed. Restart the gateway after changing this field. Inspect sandbox OCSF configuration and finding events for the validation rationale, configured and effective modes, active generation, and the explicit `previous_policy_active` state. The published supervisor image uses a shell-free distroless Debian 13 base. Use container logs, engine inspection and the configured exec health probe for diagnostics; `exec ... sh`, package installation and in-container shell scripts are unavailable. Workload shells belong to the separate sandbox image. Preserve the driver-selected UID and writable runtime/log mounts when reproducing a supervisor startup failure. A `ConfigurationInvalid` readiness condition means startup admission rejected the image/effective policy or provider configuration. The supervisor remains alive while the workload stays unstarted. Inspect `openshell sandbox get` and repair the desired configuration with a complete policy replacement or provider change; do not treat a healthy container as proof that the workload is ready. If the 300-second provisioning repair window expires, the gateway records `ProvisioningTimedOut` and stops workload and supervisor compute. Inspect `provisioning` in sandbox JSON and TUI NOTES to distinguish cleanup pending from complete. Repairing configuration after expiry does not restart compute: wait for cleanup, then explicitly use `sandbox start`. Repeated rejected reports do not refresh the deadline, and the CLI wait timeout does not control it. See [policy validation and repair](https://docs.nvidia.com/openshell/latest/how-it-works/policies/overview). The isolated supervisor requests image-policy discovery through the authenticated sandbox boundary before admission. The workload boundary can remain alive without launching the workload while configuration is repaired. An unavailable boundary fails discovery within its control-request deadline. Permanent gateway errors and exhausted transient retries terminate startup; inspect those errors as connectivity, authorization, or lifecycle failures. ### Step 4: Check Docker-Backed Gateways ```bash docker info docker ps --filter name=openshell docker logs <container> --tail=200 docker run --rm --entrypoint /openshell-sandbox "${OPENSHELL_SANDBOX_RUNTIME_IMAGE:-ghcr.io/nvidia/openshell/sandbox:latest}" --version openshell status ``` For Docker GPU failures, check CDI support and NVIDIA CDI discovery separately: ```bash docker info --format '{{json .CDISpecDirs}}' docker info --format '{{json .DiscoveredDevices}}' for dir in /etc/cdi /var/run/cdi; do if [ -d "$dir" ]; then find "$dir" -maxdepth 1 -type f \( -name '*.yaml' -o -name '*.json' \) -print else echo "$dir missing" fi done systemctl is-enabled nvidia-cdi-refresh.service nvidia-cdi-refresh.path || true systemctl is-active nvidia-cdi-refresh.service nvidia-cdi-refresh.path || true systemctl status nvidia-cdi-refresh.service nvidia-cdi-refresh.path --no-pager --lines=50 journalctl -u nvidia-cdi-refresh.service --no-pager --lines=100 ``` When the NVIDIA Container Toolkit CDI refresh units are not enabled or no NVIDIA CDI spec has been generated, enable them and trigger a refresh: ```bash sudo systemctl enable --now nvidia-cdi-refresh.path sudo systemctl enable --now nvidia-cdi-refresh.service sudo systemctl restart nvidia-cdi-refresh.service docker info --format '{{json .DiscoveredDevices}}' ``` Common findings: - Docker daemon unavailable: start Docker Desktop or Docker Engine. - Gateway process stopped: inspect exit status and logs. - Sandbox image missing or pull denied: verify image reference and registry credentials. - Sandbox fails before readiness with an identity-resolution error: inspect the image's OCI `USER` and matching `/etc/passwd` and `/etc/group` entries, or explicitly set both process identity fields in policy. Numeric workload identities `1` through `4294967294` are accepted; root, the invalid identity sentinel, and missing identities are rejected. - Sandbox fails before readiness with an OCI workspace validation error: inspect the image's `WorkingDir` using the immutable image ID reported by the gateway. Empty, `/`, and explicit `/sandbox` use the managed `/sandbox` compatibility workspace. Any other workdir must be an absolute normalized directory with no symlink components; the final policy UID, primary GID, and supplementary groups must pass the kernel's effective traverse/write checks, including POSIX ACL and LSM decisions. OpenShell does not create, chown, or chmod a non-default image workdir. - Docker also rejects an image `VOLUME` that covers the workdir or one of its parents because the runtime would mask the immutable path before validation. Move the `VOLUME` below the workspace or remove the declaration. - A workdir rejected as a special filesystem or OpenShell control-path collision cannot be made valid with permissions. Move the image workdir away from kernel-backed mounts and the concrete supervisor, TLS, token, runtime, and socket paths named in the error. - Local Docker gateway setup cannot copy `openshell-sandbox` after exporting a supervisor image: the sandbox runtime and supervisor are separate artifacts. The runtime image must provide `/openshell-sandbox`; the supervisor image provides `/openshell-supervisor`. - Docker driver cannot initialize because it cannot find `openshell-sandbox`: verify the sibling binary next to `openshell-gateway`, or that the configured `sandbox_runtime_image` contains `/openshell-sandbox`. - Sandbox never registers: check gateway logs and the supervisor's gateway endpoint. - Calls to an external tool server fail while the sandbox is Ready: inspect `Tool server connections` in `openshell sandbox get <name>`. For configured MCP-over-HTTP endpoints, JSON output exposes each address together with `last_result` and `last_reported_at` in `endpoint_statuses`. Select the endpoint by host, path, and ports, then check the reported failure boundary. `last_reported_at` records gateway acceptance time and can advance when retained evidence is accepted after a reset. Results do not expire or prove current availability; `HttpResponseReceived` can still contain a tool error. If several paths share a host and port, a failure before the path is known remains in logs. Verify the actual operation when current tool availability matters. - On Docker Desktop, repeated `Policy fetch failed after 5 attempts` messages can mean host networking is disabled. Enable host networking in Docker Desktop, ensure Enhanced Container Isolation is disabled, and verify the gateway's primary endpoint is reachable from a host-networked container. - Sandbox runtime image exits before printing `openshell-sandbox --version`: verify the configured image contains a static executable at `/openshell-sandbox`. - A sandbox with explicit `protocol: tcp` endpoints fails before workload readiness: confirm the selected isolation backend advertises TCP mediation, then inspect the sandbox and supervisor logs for protected-channel setup or listener failures. A driver that cannot supply the required outer egress fence and authenticated runtime channel must reject the policy before starting the agent. - Supervisor runtime validation fails: verify `supervisor_image` contains an `/openshell-supervisor` executable from the same release as the sandbox runtime, and that the dynamic loader and shared libraries it links against are available inside that image. `docker run --rm --network none --entrypoint /openshell-supervisor <supervisor_image> --version` should print that release; a `no such file or directory` error for a binary that exists means the loader or a library is missing. The supervisor runs from its own image and does not need to be static; only `/openshell-sandbox` must be. - The sandbox fails its enforcement probe: inspect the sandbox log for the exact nested seccomp user-notification, task-memory, Landlock, loopback DNS, or socket-injection check that failed. A runtime may return `ENOSYS` for `process_vm_readv` and `process_vm_writev` while satisfying the production parent-to-workload-child task-memory probe through `/proc/<pid>/mem`; only failure of both backends is fatal. Do not add capabilities or switch to an unconfined seccomp profile; use a runtime whose default profile permits the unprivileged probe. - A GPU sandbox fails because Docker reports no discovered NVIDIA CDI devices: verify `.DiscoveredDevices` contains entries such as `nvidia.com/gpu=all`, verify `/etc/cdi` or `/var/run/cdi` contains a generated NVIDIA spec, and check that `nvidia-cdi-refresh.service` and `nvidia-cdi-refresh.path` from NVIDIA Container Toolkit are enabled and healthy. The service is a one-shot unit, so `inactive (dead)` can be normal after a successful run; use `systemctl status` and `journalctl` to distinguish success from a skipped or failed refresh. Restart `nvidia-cdi-refresh.service` to regenerate missing or stale CDI specs, then restart or reload Docker and re-check `docker info`. During a graceful gateway restart, Docker, Podman, and VM sandboxes with running intent should stop before the gateway exits and restart after it returns. Check for `Stopped sandbox during gateway shutdown` and `Started sandbox during gateway startup` in gateway logs. A sandbox explicitly stopped through the CLI remains stopped. Kubernetes sandboxes are cluster-owned and do not follow this local gateway lifecycle. Internal and external drivers follow the same rule: `GetCapabilities.gateway_manages_lifecycle` must be true for the gateway to run shutdown and startup sweeps. The gateway also drains supervisor-session ownership cleanup before exiting. If shutdown reports `Gateway supervisor session cleanup incomplete`, inspect the associated persistence errors: a stopped supervisor's owner record may remain until its lease expires and temporarily block reconnection. Successful compute stop alone does not confirm that session cleanup finished. ### Step 5: Check Podman-Backed Gateways ```bash podman info podman ps --filter name=openshell podman logs <container> --tail=200 openshell status ``` Common findings: - Podman socket unavailable: start or expose the user socket. - Rootless networking unavailable: inspect Podman network configuration. - Sandbox image missing or pull denied: verify image reference and registry credentials.
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub