Audit scientific implementation, exported-model equivalence, dependency completeness, manifests, packaging, and isolated inference. Use when claims depend on code, trained artifacts, evaluators, or deployable entry points.
Audit scientific data semantics, sample and label alignment, leakage, group splits, preprocessing boundaries, and train-to-inference transforms. Use for any result based on datasets, feature matrices, tensors, repeated observations, or learned preprocessing.
Audit numeric, artifact, log, citation, and cross-report claims against inspectable evidence. Use for reports, syntheses, benchmark claims, external citations, or conflicting Expert outputs.
Coordinate iterative evidence and reliability reviews between BrainPilot's Principal Investigator and Auditor. Use when PI needs to audit its own draft, an Expert result, or a multi-agent synthesis; when Auditor receives such a review task; or when a previous…
Audit whether method discovery, comparison, representative real-data validation, collapse diagnostics, pruning, and selection evidence support claims of suitability or superiority. Use for research-method selection, empirical evaluation, benchmarking,…
Create or update the canonical Markdown inventory of task-relevant research data. Engineer must invoke this skill before creating or updating any data inventory, data contract, or dataset-coverage summary that downstream agents will use.
Research a bounded factual, documentation, API, or literature question from authoritative sources and save a self-contained Markdown report with claim-level citations. Use for reading-heavy evidence gathering, not experiment execution or data analysis.
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.