Filter, compare, and rank papers, posts, captures, threads, bookmarks, product claims, or research ideas for high-entropy mechanistic insight and underpriced leverage. Use when the user asks for alpha, high entropy, the most intriguing or insightful items,…
Estimate whether an AI model can complete a task and how long it will take, using METR-style time-horizon modeling. Use when scoping agent work, deciding if a task is within reach, planning retries/parallelism, estimating wall-clock time for SWE/MLE/math…
Audit supervised fine-tuning datasets against the behavior and task they are meant to teach. Use when inspecting SFT JSONL, chat messages, instruction-response pairs, tool or agent trajectories, code corpora, synthetic examples, revised datasets, base-model…
Use as Codex's default rhetoric backbone when writing, reviewing, naming, positioning, debating, or sharpening papers, proposals, essays, launches, social posts, narratives, and claims where the user wants provocative attention-pull, category-defining…
Trigger when: (1) the user asks for Manim, Manim Community, or ManimCE, (2) code contains `from manim import *`, or (3) the task is to build a mathematical explainer animation. Opinionated Manim Community skill for concise math scenes. Focuses on scene…
Canonical end-to-end workflow for turning a paper PDF or URL into grounded OCR artifacts and a teachable notes.md.
High-accuracy OCR refinement workflow with self-scoring, Maj@K consensus voting, and targeted repair for page-faithful transcription.
Use when the task requires automating a real browser from the terminal (navigation, form filling, snapshots, screenshots, data extraction, UI-flow debugging) via `playwright-cli` or the bundled wrapper script.