Analyze what a plan's execution run actually produced and judge it against what the plan expected. A PLAN_NAME (slug / numeric prefix / filename) resolves through the plan's exec_runs to its wkdrs/<run>/ directory; a wkdrs/<run>/ path back-resolves to its plan; no argument lists the runs on disk and asks. Inventories the §4 deliverables against disk, corroborates EXEC_LOG's step claims with artifacts, scans training / eval logs for health signals (crashes, NaN, OOM, divergence, overfitting), extracts the metrics the §5 done-criteria name and scores them against those criteria plus root §4 metrics and stated baselines, interprets what the numbers mean for the claim the plan traces to (root kill-criteria, leakage smells, single-seed limits), and appends a lightweight comparison when sibling runs of the same plan exist. Renders curves only when matplotlib is already installed (never installs anything), re-reads every cited number before it enters the report, and writes the analysis under wkdrs/<run>/. Read-only
Summarize what the experiment programme has done lately, on the time axis. No argument resumes from the previous digest's watermark and covers everything since; a PLAN_NAME covers that node's whole family — ancestors for claim context, every descendant for evidence — unbounded in time; `<N>d` or a date covers a window; `all` re-seeds from the beginning. Collects each in-scope run's newest EXPT_ANALYSIS report, tabulates verdicts and headline metrics with their provenance, derives what moved since the previous digest (new runs, changed verdicts, newly analyzed runs), gathers strategy signals and kill-criteria hits, notes which plans were created or revised in the period, and lists the gaps. A run with no analysis report is read raw for a provisional line only, tagged unverified in its own table, never scored and never quoted as a result. Writes one dated digest to wkdrs/digests/. Read-only otherwise: never edits plans, exec_status, logs, or the results ledger, and never re-runs an experiment. Use when the user
Read-only overview of the whole research flow. Scans every metds/plans/*_plan.md, rebuilds the decomposition tree from parent/prefix, reads each node's section status, children, depends_on, and exec_status (plus each run's EXEC_LOG.md for step-level progress), then renders the tree with status, a progress rollup, the single next action, and any staleness. Also checks the surrounding stages — ideas, refs, code reviews, experiment analyses, method documents — for finished work whose follow-up is missing or out of date. Never writes. Use when the user invokes $star-flow-status, or asks for the status / overview / progress of their research or plans, what to work on or execute next, what is left owing, how far a plan or its sub-plans have gotten, or to see the plan tree. Bilingual (en/zh).
Analyze what a plan's execution run actually produced and judge it against what the plan expected. A PLAN_NAME (slug / numeric prefix / filename) resolves through the plan's exec_runs to its wkdrs/<run>/ directory; a wkdrs/<run>/ path back-resolves to its plan; no argument lists the runs on disk and asks. Inventories the §4 deliverables against disk, corroborates EXEC_LOG's step claims with artifacts, scans training / eval logs for health signals (crashes, NaN, OOM, divergence, overfitting), extracts the metrics the §5 done-criteria name and scores them against those criteria plus root §4 metrics and stated baselines, interprets what the numbers mean for the claim the plan traces to (root kill-criteria, leakage smells, single-seed limits), and appends a lightweight comparison when sibling runs of the same plan exist. Renders curves only when matplotlib is already installed (never installs anything), re-reads every cited number before it enters the report, and writes the analysis under wkdrs/<run>/. Read-only
Summarize what the experiment programme has done lately, on the time axis. No argument resumes from the previous digest's watermark and covers everything since; a PLAN_NAME covers that node's whole family — ancestors for claim context, every descendant for evidence — unbounded in time; `<N>d` or a date covers a window; `all` re-seeds from the beginning. Collects each in-scope run's newest EXPT_ANALYSIS report, tabulates verdicts and headline metrics with their provenance, derives what moved since the previous digest (new runs, changed verdicts, newly analyzed runs), gathers strategy signals and kill-criteria hits, notes which plans were created or revised in the period, and lists the gaps. A run with no analysis report is read raw for a provisional line only, tagged unverified in its own table, never scored and never quoted as a result. Writes one dated digest to wkdrs/digests/. Read-only otherwise: never edits plans, exec_status, logs, or the results ledger, and never re-runs an experiment. Use when the user
Read-only overview of the whole research flow. Scans every metds/plans/*_plan.md, rebuilds the decomposition tree from parent/prefix, reads each node's section status, children, depends_on, and exec_status (plus wkdrs/<run>/EXEC_LOG.md for step-level progress), then renders the tree with status, a progress rollup, the single next action, and any staleness. Also checks the surrounding stages — ideas, refs, code reviews, experiment analyses, method documents — for finished work whose follow-up is missing or out of date. Never writes. Use when the user runs /star-flow-status, or asks for the status / overview / progress of their research or plans, what to work on or execute next, what is left owing, how far a plan or its sub-plans have gotten, or to see the plan tree. Bilingual (en/zh).
Analyze what a plan's execution run actually produced and judge it against what the plan expected. A PLAN_NAME (slug / numeric prefix / filename) resolves through the plan's exec_runs to its wkdrs/<run>/ directory; a wkdrs/<run>/ path back-resolves to its plan; no argument lists the runs on disk and asks. Inventories the §4 deliverables against disk, corroborates EXEC_LOG's step claims with artifacts, scans training / eval logs for health signals (crashes, NaN, OOM, divergence, overfitting), extracts the metrics the §5 done-criteria name and scores them against those criteria plus root §4 metrics and stated baselines, interprets what the numbers mean for the claim the plan traces to (root kill-criteria, leakage smells, single-seed limits), and appends a lightweight comparison when sibling runs of the same plan exist. Renders curves only when matplotlib is already installed (never installs anything), re-reads every cited number before it enters the report, and writes the analysis under wkdrs/<run>/. Read-only
Summarize what the experiment programme has done lately, on the time axis. No argument resumes from the previous digest's watermark and covers everything since; a PLAN_NAME covers that node's whole family — ancestors for claim context, every descendant for evidence — unbounded in time; `<N>d` or a date covers a window; `all` re-seeds from the beginning. Collects each in-scope run's newest EXPT_ANALYSIS report, tabulates verdicts and headline metrics with their provenance, derives what moved since the previous digest (new runs, changed verdicts, newly analyzed runs), gathers strategy signals and kill-criteria hits, notes which plans were created or revised in the period, and lists the gaps. A run with no analysis report is read raw for a provisional line only, tagged unverified in its own table, never scored and never quoted as a result. Writes one dated digest to wkdrs/digests/. Read-only otherwise: never edits plans, exec_status, logs, or the results ledger, and never re-runs an experiment. Use when the user