원클릭으로
web-search
Web search (text/news/images/videos/books) and URL content extraction using DuckDuckGo (ddgs and primp).
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Web search (text/news/images/videos/books) and URL content extraction using DuckDuckGo (ddgs and primp).
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Generate speech locally using the pocket-tts CLI. Defaults to the `fantine` voice; use `george` when a second voice is needed.
Transcribe audio/video files to text using OpenAI Whisper (faster-whisper backend). Supports word-level timestamps, multiple output formats (srt, vtt, json, txt, ass), auto language detection, and GPU acceleration. Use when transcribing audio, creating subtitles/captions, or extracting text from spoken content.
Extract clean text and structured Markdown from PDF files using PyMuPDF4LLM. Handles multi-column layouts, tables, headings, lists, and images. Use when converting PDFs to readable text, extracting content for LLM ingestion, or inspecting PDF structure. Does NOT handle scanned PDFs (no OCR).
| name | web-search |
| description | Web search (text/news/images/videos/books) and URL content extraction using DuckDuckGo (ddgs and primp). |
Web search and content extraction powered by ddgs.
All commands below run from .pi/skills/web-search using uv run. The venv should already exist.
If a command fails because it's missing, create it once:
uv venv .venv --relocatable --python 3.13 --python-preference=only-managed
uv sync
uv run python web_search.py --query "your query here"
Common patterns:
uv run python web_search.py --query "AI news" --type news --timelimit w
uv run python web_search.py --query "mountains" --type images --limit 5
uv run python web_search.py --query "Python docs" --extract-top
Search types: text, news, images, videos, books.
Fetch and extract readable content from a URL. Always save to a temp file if the content is too large to read in one go.
uv run python web_search.py --extract https://example.com --extract-fmt text_markdown --output /tmp/page.md
--extract-fmt values: text_markdown (default, markdown), text_plain, text_rich, text (raw HTML), content (raw bytes).
Note:
--extractmay intermittently return HTTP 403 on some sites due to rotating browser fingerprints. Retrying sometimes helps, but some sites (e.g. Reddit) are consistently blocked. See Known Limitations & Workarounds below.
| Site / Pattern | Issue | Workaround |
|---|---|---|
GitHub (github.com) | HTML is JS-heavy; extraction returns only navigation markup, no commits/PRs/issues. | Use api.github.com endpoints (no auth needed for public repos). Examples: api.github.com/repos/{owner}/{repo}/commits, /pulls, /issues/{n}/comments. |
Consistently returns HTTP 403 / bot challenge on all endpoints (www, old, .json). | Fall back to --type text search for post titles/URLs instead of extracting. | |
| Substack | Homepage requires JavaScript; extraction returns "enable JS" message. | Use the RSS feed at /feed (e.g. localbench.substack.com/feed). |
HuggingFace (huggingface.co/models?sort=trending) | Extraction returns filter UI only, no actual model listings. | Search for "Hugging Face Trending Models Digest" or use the agents-radar GitHub issues. |
| Weather / News sites | Often login-gated or paywalled; extraction returns signup prompts. | Rely on search result snippets, or try lightweight APIs like wttr.in. |
--type news | Intermittently fails with DecodeError: Body collection error. | Fall back to --type text --timelimit d (or w/m). |
site: + --type news | Frequently causes DecodeError. | Use --type text without site: and filter URLs manually, or extract the target site directly if it works. |
Shell redirection (> /tmp/file) | Can truncate or mix stderr into output. | Prefer the built-in --output /tmp/file flag. |
Hacker News front page — extracts cleanly as a Markdown table:
uv run python web_search.py --extract https://news.ycombinator.com/ --extract-fmt text_markdown --output /tmp/hn.md
GitHub API (commits):
uv run python web_search.py --extract https://api.github.com/repos/ikawrakow/ik_llama.cpp/commits?per_page=10 --extract-fmt text_markdown --output /tmp/commits.json
Substack RSS:
uv run python web_search.py --extract https://localbench.substack.com/feed --extract-fmt text_markdown --output /tmp/feed.xml
Run with --help for all options.