- name
- linux-safe-script-execution
- description
- Performs pre-flight validation and interaction mapping for Linux automation scripts to identify availability risks before execution.
- license
- MIT
- compatibility
- opencode
- archetypes
- ["tactical","diagnostic"]
- anti_triggers
- ["brainstorming","vague ideation","long-form architecture"]
- response_profile
- {"verbosity":"low","directive_strength":"high","abstraction_level":"operational"}
- metadata
- {"version":"1.0.0","domain":"linux","triggers":"script safety, pre-flight checks, availability risk, script interaction, automation safety, service disruption, Linux operational safety","role":"implementation","scope":"implementation","output-format":"code","content-types":["code","guidance","do-dont","examples"]}
# Safe Script Execution for Linux Infrastructure Automation
Infrastructure engineer performing pre-flight validation and interaction mapping for Linux automation scripts to identify availability risks, resource conflicts, and service disruption points before execution.
## TL;DR Checklist
- [ ] Verify script has `set -euo pipefail` and explicit error handling on every command
- [ ] Map all file-system paths the script touches and confirm no data-loss operations without backup
- [ ] Check every `systemctl` call — would restarting this service cause downtime for dependent services?
- [ ] Verify no `rm -rf`, `dd`, `truncate`, or filesystem resize operations without explicit user confirmation
- [ ] Confirm all network ports the script opens are within the approved range and documented
- [ ] Check resource consumption estimates against host capacity (disk space, memory, CPU)
- [ ] Validate that the script runs under a dedicated service account, not root, unless explicitly required
---
## When to Use
Use this skill when:
- **Deploying automation scripts on production hosts** — Any script that modifies system state, restarts services, or changes configurations
- **Reviewing third-party or vendor scripts** — Evaluating scripts from external sources before running them on your infrastructure
- **Implementing CI/CD pipeline steps that touch production systems** — Ensuring automated deployment steps don't disrupt running services
- **Auditing existing runbooks and playbooks** — Identifying risky operations in documented procedures
- **Preparing to run migration or upgrade scripts** — Systems undergoing version changes where downtime or data loss is unacceptable
- **Validating infrastructure-as-code changes before apply** — Terraform, Ansible, or Salt states that modify live systems
---
## When NOT to Use
Avoid this skill for:
- **Local development or sandbox environments** — Where no availability risk exists, standard testing is sufficient
- **Read-only inspection commands** — `systemctl status`, `journalctl`, `df`, `free` do not modify system state
- **One-time setup on disposable infrastructure** — Where the entire VM/container is ephemeral and can be discarded
- **Network packet capture or passive monitoring** — Tools like `tcpdump` or `ss` that observe without modifying
Use `linux-service-integrity-operations` when the script is ready to execute and you need safe restart patterns and health monitoring during the change.
---
## Core Workflow
### 1. Parse Script for Dangerous Operations
Identify all operations that modify system state. Build an inventory of every file, service, and network endpoint the script touches.
```bash
#!/usr/bin/env bash
# preflight_audit.sh — Analyze a deployment script for availability risks
# Usage: ./preflight_audit.sh <script.sh>
set -euo pipefail
readonly DANGEROUS_PATTERNS=(
'systemctl\s+(restart|stop|reload|kill)'
'rm\s+(-rf|--no-preserve-root)'
'dd\s+of='
'truncate\s+-s\s+0'
'mkfs\.'
'lvremove|lvreduce|vgremove'
'iptables|nftables\s+-F'
'modprobe|insmod|rmmod'
'fallocate|truncate\s+-s'
'curl|wget.*\|.*bash|sh\s*-'
'chattr\s+-i'
'umount'
'swapoff'
)
readonly DANGEROUS_PATTERNS_NAMES=(
'service_restart' 'destructive_rm' 'disk_write' 'file_truncate'
'filesystem_create' 'lv_destruction' 'firewall_flush' 'kernel_module'
'disk_allocation' 'pipe_risk' 'attr_remove' 'unmount' 'swap_disable'
)
audit_script() {
local script_file="$1"
local risk_level="LOW"
local findings=()
local service_ops=()
local file_ops=()
local resource_risks=()
if [[ ! -f "$script_file" ]]; then
echo "ERROR: Script not found: $script_file" >&2
return 1
fi
echo "=== Pre-flight Audit: $(basename "$script_file") ==="
echo "File: $script_file"
echo "Size: $(wc -c < "$script_file") bytes"
echo ""
# Check for safety patterns
if ! grep -q 'set -euo pipefail' "$script_file"; then
findings+=("CRITICAL: Missing 'set -euo pipefail' — unhandled errors will cascade")
risk_level="CRITICAL"
fi
if ! grep -qE '^\s*(set|trap)' "$script_file"; then
findings+=("WARNING: No error trapping or set flags — failures may go undetected")
fi
# Check dangerous operations
for i in "${!DANGEROUS_PATTERNS[@]}"; do
if grep -qE "${DANGEROUS_PATTERNS[$i]}" "$script_file"; then
local matches
matches=$(grep -cE "${DANGEROUS_PATTERNS[$i]}" "$script_file")
findings+=("HIGH: ${DANGEROUS_PATTERNS_NAMES[$i]} detected — ${matches} occurrence(s)")
risk_level="HIGH"
fi
done
# Check for hard-coded credentials
if grep -qiE '(password|secret|key|token)\s*=\s*["\x27]' "$script_file"; then
findings+=("CRITICAL: Hard-coded credentials detected in script")
risk_level="CRITICAL"
fi
# Map systemctl operations
while IFS= read -r line; do
service_ops+=("$line")
done < <(grep -nE 'systemctl\s+(restart|stop|reload|kill|disable|enable)' "$script_file" 2>/dev/null || true)
if [[ ${#service_ops[@]} -gt 0 ]]; then
echo ""
echo "=== Service Impact Analysis ==="
for op in "${service_ops[@]}"; do
local svc_name
svc_name=$(echo "$op" | grep -oE '[a-zA-Z0-9_-]+\.service' || echo "unknown service")
echo " [${op}] → Service: $svc_name"
done
fi
echo ""
echo "=== Risk Assessment ==="
echo "Overall risk: $risk_level"
echo "Findings: ${#findings[@]}"
for finding in "${findings[@]}"; do
echo " • $finding"
done
echo ""
echo "=== Recommendations ==="
if [[ "$risk_level" == "CRITICAL" ]]; then
echo " • DO NOT run this script in production without remediation"
echo " • Address all CRITICAL findings before proceeding"
echo " • Obtain approval from on-call engineer"
elif [[ "$risk_level" == "HIGH" ]]; then
echo " • Review all HIGH findings before execution"
echo " • Schedule during maintenance window"
echo " • Ensure rollback procedure is ready"
else
echo " • Standard change process applies"
echo " • Ensure monitoring is active during execution"
fi
}
audit_script "${1:?Usage: $0 <script.sh>}"
```
**Checkpoint:** All dangerous operations are identified and catalogued. No CRITICAL findings remain unresolved.
### 2. Map Service Dependencies
Determine what other services depend on any service this script will restart or modify. Use `systemctl list-dependencies` to map the full dependency tree.
```bash
#!/usr/bin/env bash
# service_dependency_map.sh — Map full dependency impact before restarting a service
# Usage: ./service_dependency_map.sh <service.service>
set -euo pipefail
map_dependencies() {
local target_service="$1"
if ! systemctl list-unit-files "${target_service}" &>/dev/null; then
echo "ERROR: Service '${target_service}' not found" >&2
return 1
fi
echo "=== Dependency Impact Analysis: ${target_service} ==="
echo ""
# Upstream dependencies — what must be running for this service to work
echo "--- Upstream Dependencies (Required Before Restart) ---"
systemctl list-dependencies --reverse "${target_service}" --no-pager 2>/dev/null | \
sed 's/├── / /; s/└── / /; s/│ / /' | \
grep -v 'list-dependencies' || echo " (none)"
echo ""
# Downstream dependents — what breaks if this service restarts
echo "--- Downstream Dependents (At Risk During Restart) ---"
systemctl list-dependencies "${target_service}" --no-pager 2>/dev/null | \
sed 's/├── / /; s/└── / /; s/│ / /' | \
grep -v 'list-dependencies' || echo " (none)"
echo ""
# Check for network ports that would be briefly unavailable
echo "--- Network Port Impact ---"
local socket_unit
socket_unit=$(systemctl cat "${target_service}" 2>/dev/null | grep -E '^ListenStream|^ListenDatagram' || true)
if [[ -n "$socket_unit" ]]; then
echo " Socket activations detected:"
echo "$socket_unit" | while IFS= read -r line; do
echo " $line"
done
fi
# Check if this is a core infrastructure service
local critical_services=(
"systemd-journald" "systemd-logind" "dbus" "NetworkManager"
"sshd" "cron" "systemd-timesyncd" "containerd" "docker"
"polkit" "udev" "runit"
)
for crit in "${critical_services[@]}"; do
if [[ "$target_service" == "${crit}.service" || "$target_service" == "${crit}.timer" ]]; then
echo ""
echo " ⚠ CRITICAL: '${target_service}' is a core infrastructure service"
echo " ⚠ Restart may cause brief system-wide instability"
echo " ⚠ Use 'systemctl reload' instead of 'restart' where possible"
echo " ⚠ Ensure remote access is available via console/serial before proceeding"
break
fi
done
# Check for OnFailure handlers
local failure_handler
failure_handler=$(systemctl cat "${target_service}" 2>/dev/null | grep '^OnFailure=' || true)
if [[ -n "$failure_handler" ]]; then
echo ""
echo "--- Cascade Failure Risk ---"
echo " OnFailure handler: $failure_handler"
echo " If this service fails to restart, $failure_handler will be triggered"
fi
}
map_dependencies "${1:?Usage: $0 <service.service>}"
```
**Checkpoint:** Full dependency tree is mapped. No core infrastructure services are targeted for restart without console access.
### 3. Verify Resource Headroom
Confirm the host has sufficient disk space, memory, and CPU headroom for the script's expected workload.
```bash
#!/usr/bin/env bash
# resource_headroom.sh — Verify host has sufficient resources before running a script
# Usage: ./resource_headroom.sh <script.sh>
set -euo pipefail
check_resource_headroom() {
local script_file="$1"
echo "=== Resource Headroom Check ==="
echo ""
# Disk space analysis
echo "--- Disk Space ---"
local disk_usage
disk_usage=$(df -h / | awk 'NR==2 {print $5}' | tr -d '%')
local disk_free
disk_free=$(df -h / | awk 'NR==2 {print $4}')
if [[ "$disk_usage" -gt 90 ]]; then
echo " ⚠ CRITICAL: Root filesystem at ${disk_usage}% — script may fail"
echo " ⚠ Free space: $disk_free"
elif [[ "$disk_usage" -gt 80 ]]; then
echo " WARNING: Root filesystem at ${disk_usage}% — tight on space"
echo " Free space: $disk_free"
else
echo " OK: Root filesystem at ${disk_usage}% — sufficient space"
echo " Free space: $disk_free"
fi
# Check if script creates/extends any files
local estimated_disk
estimated_disk=$(grep -oE '(truncate\s+-s\s+(\d+[kmgKMG])?)|(dd\s+.*count=\d+)|(mkfs|fdisk)' "$script_file" 2>/dev/null || true)
if [[ -n "$estimated_disk" ]]; then
echo " ⚠ Script contains disk-extending operations"
echo " Operations: $estimated_disk"
fi
echo ""
# Memory analysis
echo "--- Memory ---"
local mem_total
mem_total=$(grep MemTotal /proc/meminfo | awk '{print int($2/1024)}')
local mem_available
mem_available=$(grep MemAvailable /proc/meminfo | awk '{print int($2/1024)}')
local mem_used_pct
mem_used_pct=$(awk "BEGIN {printf \"%d\", (1 - $mem_available / $mem_total) * 100}")
if [[ "$mem_used_pct" -gt 90 ]]; then
echo " ⚠ CRITICAL: Memory at ${mem_used_pct}% — ($mem_available MB available of ${mem_total} MB)"
elif [[ "$mem_used_pct" -gt 80 ]]; then
echo " WARNING: Memory at ${mem_used_pct}% — ($mem_available MB available of ${mem_total} MB)"
else
echo " OK: Memory at ${mem_used_pct}% — ($mem_available MB available of ${mem_total} MB)"
fi
echo ""
# CPU load check
echo "--- CPU Load ---"
local load_avg
load_avg=$(cat /proc/loadavg | awk '{print $1}')
local cpu_count
cpu_count=$(nproc)
local load_ratio
load_ratio=$(awk "BEGIN {printf \"%.2f\", $load_avg / $cpu_count}")
if awk "BEGIN {exit !($load_ratio > 2.0)}"; then
echo " ⚠ WARNING: Load ratio ${load_ratio}x CPU count — consider delaying execution"
else
echo " OK: Load ratio ${load_ratio}x CPU count (${load_avg} on ${cpu_count} CPUs)"
fi
echo ""
# File descriptor check
echo "--- File Descriptors ---"
local fd_limit
fd_limit=$(ulimit -n)
local fd_used
fd_used=$(ls /proc/$$/fd 2>/dev/null | wc -l)
Ver en GitHub