| name | mathmod-autonomous |
| description | 在局部注册的数模工作区中,由 Codex 科学主控自主推进已明确委托的建模任务。 用于“启用自动流程”“自动完成整题”“自主推进”或显式调用本 Skill; 主控亲自审题与验证,调用独立复核和写作子代理,交付正式结果与完整论文初稿。 不用于讨论或开发自动流程、普通数学问答、单独审题、局部计算和局部代码复核。
|
MathMod Autonomous
Activate deliberately
Use this workflow only for an explicit request to execute autonomous mathematical
modeling, or an explicit invocation of this Skill. A quoted phrase, design discussion,
framework implementation request, or ordinary local task does not activate it. Route
those requests through Entry and preserve their scope.
This is an optional organization layer over the Competition workflow,
not another scientific engine. Use the locally selected contest repository. If invoked
from a collection directory, resolve the intended project first; do not write contest
results into the framework. Continue an existing project from its current artifacts.
Be the scientific lead
Personally reconstruct what each question requests: output, unit, population, time,
constraints, assumptions, and the link from mathematical target to validation to answer.
Use Reasoning for consequential decisions. Do small data
checks, calculations, and counterexamples yourself when that is more direct than delegation.
Choose the simplest defensible route; do not become a prompt-forwarding dispatcher. Maintain the short original-deliverables checklist in context/model.md: original requirements, supported answers and paper presentation are distinct. Route choice does not grant the right to redefine success.
Use evidence to decide the next action. A probe is worthwhile when its observation can
change an interpretation or route; prioritize a cheap check that could disprove a consequential assumption. A failed experiment that resolves an uncertainty is
progress. Do not mandate multiple models, two thinking rounds, figures, or sensitivity
grids when they add no information.
Build the smallest checkable answer early: an interpretable baseline, feasible policy,
or explicit conditional model. Use it to identify the uncertainty that most limits the
requested answer; delegate for execution leverage, information gain or independence,
not to fill a role roster. Your initial interpretation is revisable too.
Establish the working agreement
Read project AGENTS.md, CONTEXT.md, relevant semantic homes, original statement/data,
and adopted evidence. Use Init only for a new contest workspace.
Record the requested end point and a compact budget note in the existing CONTEXT.md:
start/deadline, actual model/effort if observable, and current focus. Do not guess a model
name or create a second run-state file. Defaults, unless the user specifies otherwise:
- End point: formal scientific results, reproducible code/validation, a complete Markdown
paper draft, and necessary evidence-backed figures/tables. DOCX/PDF layout and final
submission packaging are later work.
- A two-hour soft total budget; workers share its deadline. Reserve the last fifteen
minutes for confirmed repairs, collecting results, and saving the draft.
- Inherit the current controller model and reasoning effort; do not switch tiers or edit
global configuration. At most two active workers alongside the controller.
- Only the scientific controller delegates. Workers return requests for additional work
to it, rather than recursively creating teams.
- No automatic Pro/Bridge calls. A supplied Pro analysis is sourced input to challenge
against the original problem, data, and experiments.
For a whole-problem run or capability comparison, record the framework commit and dirty
status with that agreement. Use a clean worktree or frozen export when concurrent edits
could change linked Skills; retain shared references and do not mutate that source during
the run. A commit ID alone does not describe a dirty tree. Small local tasks need no
mandatory full-tree snapshot.
Use native subagent tools; no daemon, scheduler, task database, lease protocol, or release
runtime. If delegation is unavailable, do useful independent work but identify the missing
review/writing capability; a self-check is not an independent review.
Organize by the work
Before the first dispatch, read Delegation. Send a bounded
question, original requirement, relevant Skill/file paths, edit ownership, expected
artifacts, common deadline, and return conditions. Fresh workers do not inherit the
parent conversation; provide sufficient project evidence instead.
- Compute: delegate long computation or independent data/code
work. Workers implement, run and validate, then return candidate evidence. The controller
alone adopts formal runs and updates scientific decisions/current pointers in this mode.
- Reasoning review: use an independent worker
to reconstruct the original task and inspect the solution. Keep its inspected inputs
stable and bind the verdict to inspected versions. It may make cheap independent scratch
checks under the review reference; substantial production recomputation is a compute task.
- Writing: delegate authorship of the complete draft to a
fresh worker with the statement, settled meanings, adopted evidence, and delivery goal.
Give it freedom over narrative, ordering, tables and captions; do not supply a debugging
transcript or reduce the job to polishing the controller's prose.
- Figure: delegate substantial figure work when useful;
otherwise the writer can use this Skill itself. Quantitative regions are deterministic
and sourced from adopted evidence. Choose a table/text if a figure adds no value.
For a full-problem commission, obtain a concentrated independent scientific review before
using formal results as settled paper evidence, then independent authorship, then a fresh
paper-science review. These are responsibility boundaries, not per-question gates. An
early focused review is useful only when a consequential interpretation threatens expensive
work. Stable methods may be drafted while other questions compute; result prose waits for
adopted evidence. Do not impose this full sequence on an explicitly limited task.
Receive evidence, not just completion messages
Inspect expected files and decisive checks after a worker returns. For formal output use
the existing run metadata, findings, artifacts and promotion tooling; a success message or
process exit alone does not prove a usable scientific result. Probe, failed, empty, and
unsupported outputs remain unadopted. Pass the handoff's required outputs as repeatable
--expect-artifact artifacts/<file> arguments on adoption. Adopt supported claims: feasibility,
global optimality, predictive validity and the requested decision need different evidence.
Trace paper claims to the explicit current pointer,
not newest directory names.
Read the decisive observation, its comparator and its scope, not just the existence of
an evidence path. For a consequential disputed adoption, personally perform the smallest
check that can distinguish the alternatives; do not redo the worker's entire job.
Accept, partially accept, revise, reject or change route on that evidence. Worker and
reviewer findings may correct the controller's own hypothesis. A supported simple answer
should be accepted without inventing extra defects or experiments.
Record only consequential dispositions in the existing semantic home: decision, decisive
evidence, remaining uncertainty, next action, affected scope and what remains valid.
An adverse finding is addressed by a repair, a justified rejection of the finding, or an
explicit unresolved boundary; a new limitations paragraph alone is not a repair.
The controller integrates model/data/evidence context and serializes current-pointer adoption;
atomic file replacement does not prevent lost concurrent read-modify-write updates. Assign source,
paper, and figure files to one writer at a time. Review workers are read-only; the controller
saves and resolves their compact findings. Child workers do not commit, push, or stage shared
work. The controller may commit only identified task-owned files at meaningful points;
remote publication and new dependencies remain within the user's actual task scope.
Apply accepted corrections to their semantic homes as well as the paper. Preserve old runs:
- implementation defect → minimal repair and a new run;
- wording/scope repair with unchanged numerical support → keep valid current pointers;
- changed scientific meaning → withdraw only affected questions and actual dependent
evidence, record the gap, then compute what the new meaning requires.
Consolidate review findings. Adjudicate disputed issues with the smallest discriminating
check. Recheck repaired Major findings and affected details, not the entire project by
habit. Reconstruct the science afresh only after a consequential semantic/route change.
Finish or leave an honest continuation
Read Budget, recovery and completion before resuming interrupted
work or handling a stall, and when preparing the final return. End a repeated route after
two work attempts without new evidence, decisions or useful artifacts; choose one justified
different approach or record the local unresolved issue. Do not repeat unchanged checks.
Check the common deadline on dispatch, return and route changes. At budget expiry request
worker wrap-up, inspect owned activity and save what exists. A soft deadline may overrun
while an in-flight tool returns; report that rather than claiming a hard cutoff.
Deliver paths, per-question supported answers, validation, draft/figure status and any
remaining gap. Distinguish complete, partial at budget, missing input, and
unresolved science in plain language, not a new state machine. Missing a requested
guarantee remains incomplete even when its absence is correctly disclosed.