- name
- dogpile
- description
- Deep research aggregator that searches Brave (Web), GitHub (Code/Issues), ArXiv (Papers), YouTube (Videos), optional Context7 library docs, and optional feed/archive/book sources. Provides a consolidated Markdown report with an ambiguity check, grounded synthesis, and Agentic Handoff.
- allowed-tools
- ["run_command","read_file"]
- triggers
- ["dogpile","research","deep search","find code","search everything"]
- metadata
- {"short-description":"Deep research aggregator (Web, Code, Papers, Videos, Feeds)"}
- provides
- ["deep-research","web-search"]
- composes
- ["memory","tau","brave-search","github-search","arxiv","ingest-youtube","context7","ingest-website","ops-darpa","fetcher","extractor","ingest-book","task-monitor","agentic-evals"]
- complies
- ["best-practices-skills","best-practices-python"]
- runtime_self_improvement
- substantial
- taxonomy
- ["research","aggregation","resilience"]
- disciplines
- ["research-retrieval"]
> STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT.
# Dogpile: Deep Research Aggregator
Orchestrate a multi-source deep search to "dogpile" on a problem from every angle.
## Analyzed Sources
1. **Tau-owned LLM lanes (🤖)**: Query ambiguity checks, query tailoring, technical overview, synthesis, and code/paper relevance evaluation belong behind Tau. Tau may call SciLLM internally; project agents should consume Tau receipts and Dogpile reports, not raw SciLLM responses.
2. **Concurrent Brave question lanes (🌐)**: Perplexity replacement. Dogpile fans out multiple bounded Brave web queries and records each result set separately.
3. **Brave Search (🌐)**: **Three-Stage Search** (Search → Evaluate → Deep Extract via /fetcher).
4. **ArXiv (📄)**: **Three-Stage Search** (Abstracts → Details → Full Paper via /fetcher + /extractor).
5. **YouTube (📺)**: **Two-Stage Search** (Brave-first video discovery with yt-dlp fallback → Detailed transcripts via `ingest-youtube` Direct/Proxy/Whisper).
6. **GitHub (🐙)**: **Three-Stage Search**:
- **Stage 1**: Search repositories and issues
- **Stage 2**: Fetch README.md and metadata for top repos, agent evaluates relevance
- **Stage 3**: Deep code search inside the selected repository
7. **Context7 library documentation (📚, opt-in)**: Current library, SDK, framework, and DSL documentation for code/API questions after a library name or Context7 library ID is known.
8. **Fetcher (📥, internal primitive)**: Fetch selected web pages, PDFs, and documents after Brave/ArXiv/user URLs identify targets; this is not a standalone search provider.
9. **Feed monitors (📰, opt-in)**: Fresh RSS feed monitor dry-runs through `consume-feed`; this is source-health/freshness evidence, not query-specific web search.
10. **DARPA operations (🛰️, opt-in specialized lane)**: DARPA programs, opportunities, BAAs, and Grants.gov searches belong to `ops-darpa` when defense R&D or funding context is explicitly relevant.
11. **Website ingestion (🧠, opt-in handoff)**: Promote selected sites or documentation URLs into `/ingest-website` when durable RAG/memory is intentionally needed.
12. **Wayback Machine (🏛️, opt-in)**: Historical snapshots for URLs.
13. **Readarr / books / Usenet (📚, opt-in)**: Local long-form source discovery when intentionally requested.
## Features
1. **Query Tailoring**: Uses Tau-owned model orchestration to generate service-specific queries optimized for each source:
- **ArXiv**: Academic/technical terms
- **Brave Questions**: Natural-language research questions formerly sent to Perplexity
- **Brave**: Documentation-style keyword queries that must fit Brave's hard limits (`<=400` chars, `<=50` words)
- **GitHub**: Code patterns, library names
- **YouTube**: Tutorial-style phrases
2. **Ambiguity Guard**: Uses Tau-owned model orchestration to analyze the query first. If ambiguous, it asks you for clarification before wasting resources.
3. **Three-Stage Deep Dive**:
- **ArXiv**: Fetches detailed metadata → Agent evaluates → Full PDF extraction via /fetcher + /extractor
- **GitHub**: Fetches README + metadata → Agent evaluates most relevant repo → Deep code search
- **Brave**: Fetches results → Agent evaluates → Full page extraction via /fetcher
- **YouTube**: Extracts full transcripts for the most relevant videos
4. **Report Assembly and Synthesis**: Consolidates successful provider results into a Markdown report and generates a compact grounded synthesis. LLM source failures are reported as degraded provider results, not as a total search failure.
5. **Textual TUI Monitor**: Real-time progress tracking of all concurrent searches via `run.sh monitor`.
6. **Resilience Features** (2025-2026 Best Practices):
- **Per-provider semaphores**: Limits concurrent requests to avoid rate limit bans
- **Exponential backoff with jitter**: Prevents thundering herd on retries (via tenacity)
- **Rate limit header parsing**: Respects Retry-After, x-ratelimit-*, and IETF RateLimit-* headers
- **Automatic retry**: Retries rate-limited requests after appropriate backoff
- **Brave query budgeting**: Compresses overlong Brave queries before dispatch instead of sending invalid 422 requests
- **Incremental result publishing**: Writes structured partial results as providers finish so the caller does not need to wait for the final report
## Tau Provider Boundary
Dogpile should have exactly one model-orchestration boundary: Tau.
- Tau owns provider/model routing. Tau may call SciLLM internally, but Dogpile project-agent workflows must not call `$scillm`, `/scillm`, `http://localhost:4001`, `/v1/chat/completions`, or `/v1/scillm/*` directly.
- Dogpile model work should be expressed as a Tau `tau.dag_contract.v1` node, Tau skill node, or Tau-executed local `command_spec` that returns receipts.
- Dogpile retrieval sources remain native: Brave Search, GitHub, ArXiv, YouTube, Context7, and opt-in feed/Wayback/Readarr use their provider APIs or skill CLIs.
- Tau/model tasks are for query tailoring, ranking, summarization, ambiguity checks, and review of retrieved evidence.
- If Tau/model synthesis fails, Dogpile records the model lane as degraded and continues with Brave, GitHub, ArXiv, YouTube, optional feed, optional Readarr, and optional Wayback results.
- Perplexity status: retired. Dogpile does not call Perplexity by default or by flag; it records a skipped/degraded source and uses concurrent Brave question searches instead.
Implementation note: older Dogpile modules still contain a direct SciLLM adapter
for query tailoring and synthesis. Treat that adapter as legacy migration work,
not as the desired project-agent contract. Do not extend direct SciLLM usage;
move model-backed Dogpile steps behind Tau when a stable Tau adapter is
available for the target workflow.
## Orchestration Boundary
Dogpile is the retrieval engine and report emitter. Tau is the orchestration
boundary for model-backed synthesis, reviewer loops, and creator/reviewer
research workflows. Dogpile should not call WebGPT/browser tools directly.
- Use `$ask` for WebGPT, browser-oracle, oracle, deep-review, parallel-review, or
credibility review workflows.
- A Tau researcher should sit above Dogpile when creator/reviewer loops are
needed, consuming Dogpile receipts and requesting follow-up Dogpile fan-outs
when the synthesis reports weak coverage.
- Dogpile itself must emit enough grounded synthesis that a project agent can
use the result without guessing from raw provider dumps.
### Battle Consumer Boundary
`$battle` is a major downstream consumer for Dogpile research. Red and Blue
Battle agents may use Dogpile reports and receipts for technique scouting,
candidate exploit families, blue-team hardening strategies, relevant GitHub
tools/rules, DARPA/AIxCC context, papers, videos, and negative evidence.
Dogpile research is design input only for Battle. It may populate a Battle
research packet, candidate-method menu, or memory lesson, but it does not prove
exploit success, patch effectiveness, or safe tool reuse. Battle must still run
payloads, scanners, repo code, generated specimens, patch builds, regression
tests, and Judge replay inside its Docker/QEMU/digital-twin evidence gates.
For Battle use, prefer bounded Dogpile searches with clear team intent:
```bash
./run.sh search "Zip Slip exploit mitigations Java archive extraction" \
--persona battle-red \
--rationale "Battle research scout needs candidate exploit families and mitigations" \
--context "Treat sources as design input only; Docker/Judge replay is required for any exploit-success claim"
```
Useful Battle-oriented Dogpile lanes:
| Battle need | Dogpile lane |
|-------------|--------------|
| Red exploit-family scouting | Brave questions, GitHub evaluated repos, ArXiv, YouTube transcripts, security feeds |
| Blue mitigations and detection | Brave, GitHub evaluated detection/rule repos, security feeds, vendor docs via Fetcher |
| DARPA/AIxCC context | `ops-darpa`, Brave `site:darpa.mil`, ArXiv |
| Source-bearing child research | Dogpile report plus Tau reviewer/researcher receipt above Dogpile |
| Untrusted tool validation | `$github-search` isolated evaluation first, then Battle Docker/Judge if adopted |
Threat-intel and security feed hits are enrichment by default. Do not treat
feed hits as automatic block decisions, alerts, or proof of compromise unless
the project workflow adds high-confidence environmental corroboration. The
default rule is: block on certainty, hunt on suspicion, enrich everything else.
Dogpile-to-Hack handoff starts with the `dogpile.hack_scan_request.v1` contract.
That packet is a hash-bound design input describing a target identity, source
packet artifact, selected source-bearing evidence, requested scan lanes, and
explicit non-claims. Hack validates it with `./skills/hack/run.sh
validate-scan-request ...`; validation does not authorize or execute a scan,
does not invoke Docker, and does not prove source truth, exploitability, patch
effectiveness, or Battle readiness. Live Hack consumption remains gated by the
separate authorization, Compose policy, sterile environment, and proof-authority
tickets.
### Fetcher Boundary
`fetcher` is part of Dogpile as a fetch/deep-extraction primitive, not as a
separate broad discovery source. Use it after Dogpile has a concrete URL from
Brave, ArXiv, user input, Wayback, a feed item, or another provider.
| Fetcher use | Activate when | Do not use as |
|-------------|---------------|---------------|
| Single-page fetch | A selected result needs full text, markdown, PDF download, SPA rendering, content verdicts, or source receipts before synthesis | A replacement for Brave or GitHub search |
| Manifest fetch | Dogpile has a bounded URL set and needs comparable extracted text across those exact sources | An arbitrary crawl of a whole site |
| PDF/document fetch | ArXiv, Brave, or user input identifies a paper, standard, report, manual, or attachment that needs extraction | A way to infer paper/code relevance without provider metadata |
Every Fetcher-backed result must preserve the URL, final URL, content verdict,
and artifact path when available. If `content_verdict` is `empty`, `thin`,
`paywall`, or `error`, Dogpile must report that degraded evidence instead of
using the result as if content was extracted. For durable site-wide learning,
handoff to `ingest-website`; for historical URL state, use Wayback.
### Optional Context7 Library Documentation Lane
Context7 is an opt-in documentation source for code-related questions. Use it
when the project agent needs current library, SDK, framework, package, or DSL
documentation and can name the relevant library or Context7 library ID.
| Context7 use | Activate when | Avoid when |
|--------------|---------------|------------|
| Library/API syntax | A code question depends on current API signatures, options, examples, config fields, or migration behavior for a named library | The question is broad web research, security news, threat intel, papers, videos, or repository discovery |
| Framework/SDK docs | Dogpile found or was given a target dependency such as React, FastAPI, ArangoDB, Lean, PyTorch, or a vendor SDK | The library is unknown; use Brave/GitHub first to identify candidates |
| Hack/Battle support | Hack or Battle needs safe usage, mitigation, hardening, or dependency behavior docs for a known target library | Runtime exploit success, patch effectiveness, tool safety, or score must be proven by Hack/Battle receipts |
Context7 requires `CONTEXT7_API_KEY`. It is skipped by default and must not gate
baseline Dogpile health. Select it explicitly:
```bash
./run.sh search "React useEffect cleanup race condition mitigation" \
--with-context7 \
--context7-library react
./run.sh search "ArangoDB AQL BM25 search syntax" \
--with-context7 \
--context7-library /arangodb/arangodb
```
Context7 evidence is source-bearing only as current documentation context bound
to a library ID and local artifact. It does not replace Brave/GitHub discovery,
Tau synthesis, Fetcher deep extraction, or Battle/Hack runtime proof.
### Optional Feed And API Source Selection
Feeds are disabled by default. Enable them only when fresh security/code
monitoring is relevant to the research question, and use them as contextual
enrichment alongside Brave, GitHub, ArXiv, and YouTube evidence.
| Feed pack/source | Excels at | Activate when | Avoid when |
|------------------|-----------|---------------|------------|
| `security_code` | Low-noise default mix for code, AppSec, vulnerability, red-team, and operational security news | The project needs compact fresh security/code context without overwhelming the report | The task is not security/code-related or only needs a direct answer from Brave/GitHub/ArXiv |
| `security_code_extended` | Adds practitioner-grade malware, cloud, exploit-development, email-threat, and policy context | The compact pack is too narrow or the question spans malware/cloud/policy tradeoffs | The task is time-constrained, broad, or likely to drown in enrichment |
| BleepingComputer | Daily incidents, ransomware, malware, breach reporting, and active exploitation | The agent needs current operational security news or incident context | Deep exploit root cause, code-level AppSec, or academic rigor is the primary need |
| Krebs on Security | Investigative cybercrime, fraud, breach infrastructure, and underground economy reporting | Attribution, criminal infrastructure, or breach-background context matters | The task needs fast CVE mechanics, tool usage, or code examples |
| SANS Internet Storm Center | Handler diaries, near-term defender awareness, and operational observations | Blue-team triage, current scanning, exploit attempts, or defender context matters | The task needs polished tutorials, broad news, or detailed exploit-development internals |
| Help Net Security | Security tooling, industry trends, and general security updates | The agent needs tool/trend awareness around a topic | The task needs high-confidence threat intel, code-level vulnerability research, or exploit mechanics |
| PortSwigger Research | Web application security, HTTP/browser attacks, payload research, and Burp ecosystem findings | Web/AppSec exploitation, testing methodology, or request/response attack classes are relevant | The topic is infrastructure, malware, cloud, or policy rather than web security |
| Google Project Zero | Deep vulnerability research, exploit chains, memory safety, root cause, and platform internals | The agent needs rigorous technical depth and vulnerability mechanics | The task needs daily news, tooling updates, or quick operational triage |
| Google Online Security Blog | Platform security, browser/ecosystem defenses, secure engineering, and policy-relevant technical context | The project needs Google/platform security direction or secure-engineering context | The question is about exploit PoCs, red-team tradecraft, or specific IOC enrichment |
| GitHub Security Blog | Supply chain, dependencies, open-source security, GitHub platform defenses, and DevSecOps | The task involves package ecosystems, CI/CD, dependency risk, or GitHub-native workflows | The task needs malware detonation, network indicators, or non-code threat reporting |
| GitHub Security Lab | CodeQL, variant analysis, code-level bug research, and open-source vulnerability writeups | The agent needs source-code vulnerability patterns or CodeQL/security-lab research | The task is not code-centric or needs operational incident news |
| SpecterOps | Active Directory, Windows internals, identity attack paths, and enterprise red-team tradecraft | AD/Windows/identity abuse or red-team methodology is in scope | The task is web AppSec, malware triage, or general news |
| Black Hills Information Security | Practical pentest methods, tooling, defensive/offensive operations, and approachable tradecraft | The agent needs practitioner technique context or operator-oriented explanation | The task needs academic depth, exact CVE status, or primary vendor documentation |
| TrustedSec | Red-team methodology, tooling, attack simulations, and practitioner writeups | The project needs offensive workflow, tool-release, or enterprise pentest context | The task requires vendor-neutral standards, legal/policy analysis, or low-noise news only |
| SentinelOne Labs | Malware reverse engineering, APT/campaign analysis, and technical malware behavior | Malware families, loader behavior, campaign infrastructure, or reversing detail matters | The task is general AppSec, dependency security, or non-malware code review |
| Malwarebytes Labs | Commodity malware, malvertising, consumer/enterprise threat landscape, and practical malware news | Broad malware awareness or user-facing threat explanation is useful | The task needs deep reverse engineering or source-code-level vulnerability analysis |
| Wiz Blog | Cloud, Kubernetes, identity, and infrastructure security research | Cloud posture, cross-tenant bugs, Kubernetes/runtime, or IAM risk is central | The topic is endpoint malware, web payload research, or on-prem AD tradecraft |
| Unit 42 | Threat research, cloud campaigns, network security, adversary reporting, and incident context | The agent needs broad vendor threat research with campaign and infrastructure detail | The task needs neutral academic literature or small, low-noise code sources |
| Offensive Security | Exploit-development education, offensive security training, Kali ecosystem, and technique walkthroughs | Learning/offensive methodology or exploit-development education is relevant | The task needs current breach reporting, official advisories, or defensive-only policy |
| Corelan Team | Windows exploit development, mitigation bypass, stack/heap exploitation, and low-level training | The project needs exploit-dev mechanics or legacy-to-modern Windows exploitation concepts | The task is cloud, policy, news, or high-level incident triage |
| Proofpoint Threat Insight | Email-borne threats, phishing, BEC, loaders, and initial-access tradecraft | Email security, phishing campaigns, or initial access matters | The task is web AppSec, AD tradecraft, or non-email infrastructure research |
| EFF Deeplinks | Security-relevant privacy, CFAA/DMCA, policy, civil liberties, and legal context | Legal/policy constraints affect security research, disclosure, or tooling decisions | The task needs direct technical exploitation details or IOC enrichment |
DARPA is a specialized government R&D lane, not a default security-news feed.
When the question is about DARPA programs, BAAs, opportunities, technical
offices, national-security R&D direction, or defense funding, route through
`ops-darpa` first:
```bash
../ops-darpa/run.sh feed programs --json
../ops-darpa/run.sh feed opportunities --json
../ops-darpa/run.sh grants "cybersecurity" --limit 10 --json
```
Use DARPA feed hits as official program/opportunity context. They do not prove
that a technology is deployed, available for use, secure, funded to your
organization, or relevant to ordinary vulnerability triage. For normal CVE,
malware, phishing, pentest tooling, or blue-team detection questions, prefer
Brave, GitHub, ArXiv, YouTube, and the `security_code` feed packs before
activating DARPA-specific scans.
Raw IoC feeds such as CISA KEV JSON, URLhaus, Spamhaus, OpenPhish, and
AlienVault OTX are not part of the default readable RSS lane. Treat them as
separate enrichment/TIP inputs that need freshness, confidence, relevance,
allowlist, and corroboration checks before any alerting or blocking decision.
The built-in `security_code` and `security_code_extended` RSS packs do not
require API keys. Raw/vendor threat-intel feeds and TIP integrations may require
API keys or access controls and must be reported as unproven when credentials
are absent.
### Feed Credential Requirements
The configured Dogpile RSS feed packs are public readable RSS sources and should
run without API keys:
| Pack | Configured sources | API key requirement |
|------|--------------------|---------------------|
| `security_code` | BleepingComputer, Krebs, SANS ISC, Help Net Security, PortSwigger, Google Project Zero, Google Online Security Blog, GitHub Security Blog, GitHub Security Lab, SpecterOps, BHIS, TrustedSec | None |
| `security_code_extended` | Extends `security_code` plus SentinelOne Labs, Malwarebytes Labs, Wiz, Unit 42, Offensive Security, Corelan, Proofpoint Threat Insight, EFF Deeplinks | None |
Do not mix these public RSS packs with optional raw/TIP/API sources:
| Source class | Examples | Credential/access status |
|--------------|----------|--------------------------|
| Public raw enrichment | CISA KEV JSON, URLhaus, Spamhaus, OpenPhish | Not part of the readable RSS lane; many public endpoints need custom parsers, TTL/confidence handling, and false-positive controls |
| Vendor/TIP APIs | VirusTotal, ANY.RUN, Hybrid Analysis, GreyNoise, Malpedia API, PhishTank API | API key or account required in the resource registry; Malpedia API is invite-only |
| Public code/resource mirrors | Malpedia GitHub, PhishTank Database GitHub | No API key; prefer public repo/data evidence before invite-only or credentialed API access when it fits the task |
| Internet/OSINT APIs | Shodan, Censys, ZoomEye, Hunter.io, SecurityTrails | API key or account required in the resource registry |
| Manual communities | BHIS Discord, TrustedSec Discord, OffSec Discord, Red Team Village, Hack The Box Discord, BloodHound Gang, and similar Discord/Slack communities | Manual user membership/invite only; not RSS, not an API-key feed, and not assumed bot-readable |
### Credentialed API References
Credentialed APIs are optional enrichment lanes. They are never part of the
default readable RSS feed packs and must not be treated as required Dogpile
health unless the project explicitly enables that paid/account-backed provider.
| API source | Documentation | Dogpile default | Activate when | Required environment |
Ver no GitHub