用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/jmagly/aiwg --skill acquire命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
WCAG accessibility analysis for color palettes including contrast ratios, compliance checking, and remediation suggestions. Use when user needs to verify colors meet accessibility standards.
Generate, analyze, compare, export, and suggest color palettes using color theory. Use when user asks about colors, palettes, color schemes, or needs help choosing colors for a project.
Research current color trends from Pantone, architecture, film, and design. Use when user asks about trending colors, popular palettes, or wants research-backed color inspiration.
基于 SOC 职业分类
正在显示 SKILL.md
| namespace | aiwg |
| name | acquire |
| platforms | ["all"] |
| description | Download media from discovered sources with format selection and progress tracking |
| commandHint | {"argumentHint":"--plan <sources.yaml> | --url <URL> [--format audio|video|best] [--output <dir>] [--parallel N]","allowedTools":"Bash, Read, Write, Glob, Grep","model":"sonnet","category":"media-curator","modelRole":"coding","modelTier":"standard"} |
Download media from discovered sources with intelligent format selection, parallel execution, and comprehensive progress tracking.
The /acquire command orchestrates media downloads from various sources (YouTube, Internet Archive, Bandcamp, direct links) using the Acquisition Manager agent. It handles:
--plan <sources.yaml>
/find-sources command.curator/sources/plan-001.yaml--url <URL>
https://youtube.com/watch?v=abc123--format <audio|video|best>
audio: Extract audio only (Opus 128K default)video: Best video up to 1080p with audiobest: Let agent decide based on content type (default)--format audio--output <directory>
./downloads/<timestamp><output>/<artist>/<era>/audio/ and /video/--output /mnt/archive/music--parallel <N>
--parallel 5--local-work <directory>
/tmp/curator-work-$$--local-work /fast-local-disk/tmp--verify-after
--verify-after--extract-audio
--extract-audio--resume <session-id>
.curator/sessions/<session-id>/state.json--resume 20260214-143022# Download single YouTube video (audio only)
/acquire --url "https://youtube.com/watch?v=abc123" --format audio
# Download single video with metadata
/acquire --url "https://youtube.com/watch?v=abc123" \
--format video \
--output /mnt/archive/concerts \
--verify-after
# Download entire source plan (generated by /find-sources)
/acquire --plan .curator/sources/plan-001.yaml
# Download to network mount (uses local working dir)
/acquire --plan .curator/sources/plan-001.yaml \
--output /mnt/network-storage/archive \
--local-work /fast-ssd/curator-work \
--parallel 3
# Download videos and auto-extract audio
/acquire --plan .curator/sources/plan-001.yaml \
--format video \
--extract-audio \
--verify-after
# Result structure:
# downloads/artist/era/video/concert.mkv
# downloads/artist/era/audio/concert.opus
# Resume after network failure
/acquire --resume 20260214-143022
# Resume with different output (move completed files)
/acquire --resume 20260214-143022 \
--output /new/location
# Validate parameters
- Check --plan exists OR --url provided
- Verify output directory is writable
- Check required tools (yt-dlp, wget, curl, ffmpeg)
- Test network mount performance if applicable
- Create session directory structure
# Session setup
SESSION_ID=$(date +%Y%m%d-%H%M%S)
SESSION_DIR=".curator/sessions/$SESSION_ID"
mkdir -p "$SESSION_DIR"/{logs,metadata}
# Initialize state file
cat > "$SESSION_DIR/state.json" <<EOF
{
"session_id": "$SESSION_ID",
"started_at": "$(date -Iseconds)",
"status": "initializing",
"parameters": {
"plan": "$PLAN_FILE",
"format": "$FORMAT",
"output": "$OUTPUT_DIR",
"parallel": $PARALLEL
},
"downloads": []
}
EOF
# If --plan provided, parse YAML
if [[ -n "$PLAN_FILE" ]]; then
# Extract URLs and metadata
SOURCES=$(yq eval '.sources[] | .url' "$PLAN_FILE")
TOTAL_SOURCES=$(echo "$SOURCES" | wc -l)
# Validate each URL is accessible
while IFS= read -r url; do
if yt-dlp --dump-json "$url" >/dev/null 2>&1; then
echo "VALID: $url"
else
echo "WARNING: Cannot access $url"
fi
done <<< "$SOURCES"
fi
# If --url provided, validate single URL
if [[ -n "$SINGLE_URL" ]]; then
if ! yt-dlp --dump-json "$SINGLE_URL" >/dev/null 2>&1; then
echo "ERROR: Cannot access URL: $SINGLE_URL"
exit 1
fi
fi
create_acquisition_structure() {
local output_base="$1"
local artist="$2"
local era="$3"
# Sanitize directory names
local safe_artist=$(echo "$artist" | sed 's/[^a-zA-Z0-9_-]/_/g')
local safe_era=$(echo "$era" | sed 's/[^a-zA-Z0-9_-]/_/g')
# Create directory tree
local audio_dir="$output_base/$safe_artist/$safe_era/audio"
local video_dir="$output_base/$safe_artist/$safe_era/video"
mkdir -p "$audio_dir/.curator"
mkdir -p "$video_dir/.curator"
mkdir -p "$output_base/$safe_artist/.curator"
# Write artist metadata
cat > "$output_base/$safe_artist/.curator/artist-info.json" <<EOF
{
"name": "$artist",
"era": "$era",
"created_at": "$(date -Iseconds)",
"session_id": "$SESSION_ID"
}
EOF
echo ":"
}
# Download orchestration with concurrency control
ACTIVE_DOWNLOADS=()
MAX_CONCURRENT=${PARALLEL:-3}
launch_download() {
local url="$1"
local output_dir="$2"
local format="$3"
local download_id="dl-$SESSION_ID-$(date +%s)"
# Wait for slot availability
while [[ ${#ACTIVE_DOWNLOADS[@]} -ge $MAX_CONCURRENT ]]; do
check_completed_downloads
sleep 2
done
# Determine target directory (audio vs video)
local target_dir
if [[ "$format" == "audio" || "$format" == "bestaudio" ]]; then
target_dir="$output_dir/audio"
else
target_dir="$output_dir/video"
fi
# Launch download in background
(
download_with_retry "$url" "$target_dir" "$format" "$download_id"
echo "COMPLETED:$download_id:$?" >>
) &
pid=$!
ACTIVE_DOWNLOADS+=()
update_state_file
}
() {
url=
target_dir=
format=
download_id=
attempt=0
max_attempts=3
[[ -lt ]];
attempt=$((attempt + ))
yt-dlp -f \
--output \
--write-info-json \
--write-thumbnail \
--no-playlist \
\
2>&1 | ;
update_state_file
0
error_type=$(classify_error )
[[ == || == ]];
pkill -P $$
2
[[ -lt ]];
wait_time=$(( * attempt))
update_state_file
1
}
() {
new_active=()
entry ;
pid=
-0 2>/dev/null;
new_active+=()
ACTIVE_DOWNLOADS=()
}
# Real-time progress monitoring
monitor_progress() {
local session_dir="$1"
while true; do
# Load current state
local state_file="$session_dir/state.json"
local total=$(jq '.downloads | length' "$state_file")
local completed=$(jq '[.downloads[] | select(.status == "completed")] | length' "$state_file")
local in_progress=$(jq '[.downloads[] | select(.status == "in_progress")] | length' "$state_file")
local failed=$(jq '[.downloads[] | select(.status == "failed")] | length' "$state_file")
# Clear screen and display status
clear
cat <<EOF
ACQUISITION PROGRESS
====================
Session: $SESSION_ID
Started: $(jq -r '.started_at' "$state_file")
Status: $completed/$total completed ($failed failed, $in_progress in progress)
Active Downloads:
EOF
# Show active download details
jq -r '.downloads[] | select(.status == "in_progress") | " [\(.progress_percent)%] \(.filename) @ \(.speed_mbps)MB/s (ETA: \(.eta_seconds)s)"' "$state_file"
# Exit if all done
if [[ $((completed + failed)) -eq $total ]]; then
echo ""
echo "All downloads completed."
5
}
monitor_progress &
MONITOR_PID=$!
if [[ "$EXTRACT_AUDIO" == "true" ]]; then
echo "Extracting audio from video files..."
find "$OUTPUT_DIR" -type d -name "video" | while read -r video_dir; do
local audio_dir="${video_dir%/video}/audio"
find "$video_dir" -type f \( -name "*.mkv" -o -name "*.mp4" -o -name "*.webm" \) | while read -r video_file; do
local basename=$(basename "$video_file" | sed 's/\.[^.]*$//')
local audio_file="$audio_dir/${basename}.opus"
echo " $video_file -> $audio_file"
ffmpeg -i "$video_file" \
-vn \
-acodec libopus \
-b:a 128K \
"$audio_file" \
2>&1 | tee -a "$SESSION_DIR/logs/audio-extraction.log"
if [[ ${PIPESTATUS[0]} -eq 0 ]]; then
echo
if [[ "$VERIFY_AFTER" == "true" ]]; then
echo "Verifying downloaded files..."
local verification_results="$SESSION_DIR/verification.json"
echo '{"verified": [], "failed": []}' > "$verification_results"
find "$OUTPUT_DIR" -type f \( -name "*.opus" -o -name "*.mkv" -o -name "*.mp4" -o -name "*.flac" \) | while read -r file; do
echo "Verifying: $file"
# Check file size
local size=$(stat -c%s "$file" 2>/dev/null || stat -f%z "$file")
if [[ $size -eq 0 ]]; then
echo "FAILED: Zero-byte file"
jq ".failed += [\"$file\"]" "$verification_results" > "$verification_results.tmp"
mv "$verification_results.tmp" "$verification_results"
continue
ffmpeg -v error -i -f null - 2>&1 | grep -q ;
jq >
jq >
verified_count=$(jq )
failed_count=$(jq )
if [[ -n "$LOCAL_WORK" && "$OUTPUT_DIR" != "$LOCAL_WORK"* ]]; then
echo "Copying from local work directory to network mount..."
# Batch copy with progress
rsync -av --progress \
--exclude='.DS_Store' \
--exclude='Thumbs.db' \
"$LOCAL_WORK/" "$OUTPUT_DIR/"
# Verify checksums
echo "Verifying network copy..."
(cd "$LOCAL_WORK" && find . -type f -exec sha256sum {} \; | sort) > /tmp/local-checksums
(cd "$OUTPUT_DIR" && find . -type f -exec sha256sum {} \; | sort) > /tmp/remote-checksums
if diff /tmp/local-checksums /tmp/remote-checksums; then
echo "Verification SUCCESS - removing local working copy"
rm -rf "$LOCAL_WORK"
else
echo "WARNING: Checksum mismatch - keeping local copy at $LOCAL_WORK"
fi
fi
Real-time console output during acquisition:
ACQUISITION PROGRESS
====================
Session: 20260214-143022
Started: 2026-02-14T14:30:22Z
Status: 5/12 completed (1 failed, 3 in progress)
Active Downloads:
[67%] concert-1.mkv @ 12.5MB/s (ETA: 245s)
[23%] album.flac @ 8.2MB/s (ETA: 680s)
[91%] podcast-ep5.opus @ 2.1MB/s (ETA: 45s)
Recent Completions:
[✓] documentary.mp4 (1.2GB, 8m 34s)
[✓] live-set.opus (85MB, 1m 12s)
Recent Failures:
[✗] unavailable-video.mkv (Video unavailable - marked for alternate source)
Total Downloaded: 4.8GB / ~12.5GB estimated
Average Speed: 8.9MB/s
Estimated Completion: 14:58 (28 minutes remaining)
| Error | Detection | Response | User Action Required |
|---|---|---|---|
| URL inaccessible | Pre-download validation fails | Skip URL, mark for manual review | Check URL, update source plan |
| Network timeout | Download stalls for 60s | Retry with exponential backoff (3x) | Check network connection |
| Rate limited | HTTP 429 response | Wait 60s, retry (2x) | Reduce parallel count |
| Disk full | Write fails with ENOSPC | STOP all downloads, escalate | Free disk space |
| Format unavailable | yt-dlp format error | Try fallback formats | Accept lower quality or skip |
| Corrupted download | ffmpeg verification fails | Delete and retry (2x) | Report to source |
| Permission denied | Write fails with EACCES | STOP, escalate | Fix permissions |
| Mount failure | Network mount timeout | Fall back to local-only mode | Check mount health |
# Per-download error log: .curator/sessions/<session-id>/logs/<download-id>.log
[2026-02-14 14:35:22] Starting download: https://youtube.com/watch?v=abc123
[2026-02-14 14:35:23] Format selected: bestvideo[height<=1080]+bestaudio
[2026-02-14 14:37:45] ERROR: HTTP Error 429: Too Many Requests
[2026-02-14 14:37:45] Classified as: rate_limited
[2026-02-14 14:37:45] Applying retry strategy: wait 60s (attempt 1/2)
[2026-02-14 14:38:45] Retrying download...
[2026-02-14 14:42:10] Download completed: concert-1.mkv (1.2GB)
# Detect stale sessions and offer recovery
detect_incomplete_sessions() {
find .curator/sessions -name "state.json" -mtime -7 | while read -r state_file; do
local status=$(jq -r '.status' "$state_file")
local session_id=$(jq -r '.session_id' "$state_file")
if [[ "$status" == "in_progress" ]]; then
local completed=$(jq '[.downloads[] | select(.status == "completed")] | length' "$state_file")
local total=$(jq '.downloads | length' "$state_file")
echo "Incomplete session detected: $session_id ($completed/$total completed)"
echo "Resume with: /acquire --resume $session_id"
fi
done
}
generate_acquisition_report() {
local session_dir="$1"
local state_file="$session_dir/state.json"
local report_file="$session_dir/report.md"
cat > "$report_file" <<EOF
# Acquisition Session Report
**Session ID**: $(jq -r '.session_id' "$state_file")
**Started**: $(jq -r '.started_at' "$state_file")
**Completed**: $(date -Iseconds)
**Duration**: $(calculate_duration "$(jq -r '.started_at' "$state_file")" "$(date -Iseconds)")
## Summary
- **Total Downloads**: $(jq '.downloads | length' "$state_file")
- **Completed**: $(jq '[.downloads[] | select(.status == "completed")] | length' "$state_file")
- **Failed**: $(jq '[.downloads[] | select(.status == "failed")] | length' "$state_file")
- **Total Size**: $(jq '[.downloads[] | select(.status == "completed") | .filesize_bytes] | add | . / 1073741824' "$state_file") GB
## Successful Downloads
$(jq -r '.downloads[] | select(.status == "completed") | "- \(.filename) (\(.filesize_bytes | tonumber / 1048576 | floor)MB)"' "$state_file")
## Failed Downloads
$(jq -r '.downloads[] | select(.status == "failed") | "- \(.url)\n Error: \(.error)"' "$state_file")
## Parameters
- **Source Plan**: $(jq -r '.parameters.plan // "N/A"' "$state_file")
- **Format Preference**: $(jq -r '.parameters.format' "$state_file")
- **Output Directory**: $(jq -r '.parameters.output' "$state_file")
- **Parallel Downloads**: $(jq -r '.parameters.parallel' "$state_file")
## File Locations
- **Session Directory**: $session_dir
- **Download Logs**: $session_dir/logs/
- **Metadata**: $session_dir/metadata/
- **Verification Results**: $session_dir/verification.json
EOF
echo "Report generated: $report_file"
cat "$report_file"
}
# Export metadata for external tools
export_metadata() {
local session_dir="$1"
local export_file="$session_dir/metadata-export.json"
jq '{
session_id: .session_id,
started_at: .started_at,
downloads: [
.downloads[] | select(.status == "completed") | {
url: .url,
filename: .filename,
format: .format,
filesize_bytes: .filesize_bytes,
duration_seconds: .duration_seconds,
checksum_sha256: .checksum_sha256
}
]
}' "$session_dir/state.json" > "$export_file"
echo "Metadata exported: $export_file"
}
For the media-curator to research-complete handoff, /acquire writes a
per-media acquisition manifest beside each downloaded file (in addition to
the session-level state.json and metadata-export.json). This manifest is the
producer side of the contract that induct-media consumes — it is the file you
pass as the first positional argument to /induct-media.
Recommended filename: <media-basename>.acquisition.json
Schema and required fields:
schema: aiwg.media.acquisition.v1title — human-readable title of the acquired mediasource_url — canonical source URL (strip trackers; normalize watch?v= form)platform — YouTube, Internet Archive, Bandcamp, direct, etc.format — audio or videomedia_path — local path to the acquired filesha256 — SHA-256 of the exact local media bytes (sha256:<hex> convention,
matching the integrity-verification skill)duration — HH:MM:SSlicense — explicit posture: platform ToS only, CC-BY-4.0,
public domain, etc. Never assert a redistributable license you cannot verify.Optional fields (recorded when known — they let induct-media backlink
profiles and skip re-derivation):
speakers — array of { "name": "Family, Given" }channel_or_venue — { "name": "..." } (e.g. a YouTube channel, podcast
feed, or conference such as NeurIPS)acquisition_id, acquired_at, session_id — provenance back to this runEmit one manifest per completed download:
write_acquisition_manifest() {
local media_path="$1" source_url="$2" platform="$3" format="$4"
local title="$5" duration="$6" license="$7"
local sha
sha="sha256:$(sha256sum "$media_path" | awk '{print $1}')"
local manifest="${media_path%.*}.acquisition.json"
jq -n \
--arg title "$title" --arg url "$source_url" --arg platform "$platform" \
--arg format "$format" --arg path "$media_path" --arg sha "$sha" \
--arg duration "$duration" --arg license "$license" \
--arg sid "$SESSION_ID" --arg at "$(date -Iseconds)" \
'{
schema: "aiwg.media.acquisition.v1",
acquisition_id: $sid, acquired_at: $at, session_id: $sid,
title: $title, source_url: $url, platform: $platform, format: $format,
media_path: $path, sha256: $sha, duration: $duration, license: $license
}' > "$manifest"
echo
}
See examples/sample.acquisition.json for a minimal manifest, and
@$AIWG_ROOT/docs/integrations/media-curator-to-research-handoff.md for the full
acquire → transcribe-media → induct-media flow.
# 1. Discover sources
/find-sources --artist "Pink Floyd" --era "1970s" --sources youtube,archive
# 2. Acquire media from discovered sources
/acquire --plan .curator/sources/plan-001.yaml \
--format video \
--extract-audio \
--verify-after \
--output /mnt/archive/pink-floyd
# 3. Extract metadata (runs automatically or manually)
/extract-metadata --source /mnt/archive/pink-floyd/The_Wall/audio
# 4. Organize final collection
/organize --source /mnt/archive/pink-floyd
--local-work when output is network mount--parallel for network mounts (recommend 2)--limit-rate to yt-dlp for shared networksaiwg.media.acquisition.v1) for the research handoff