| name | assessment-cycle |
| description | Run an interactive full assessment process: elicit a learning outcome and assessment brief, build and validate a rubric with synthetic calibration submissions, then mark real work with evidence-based feedback and feedforward. Use whenever a user asks to design a rubric, test or calibrate a rubric, assess submissions, grade work, or generate assessment feedback. |
Assessment Cycle
Purpose
Run a complete, human-controlled assessment workflow. The workflow has four gates:
- Assessment discovery and rubric design.
- Synthetic calibration and rubric revision.
- Marking of user-provided real submissions.
- Feedback, feedforward, and release review.
Never skip the calibration gate for a new or materially revised rubric unless the user explicitly directs that it be skipped. If it is skipped, label all resulting marks provisional—uncalibrated rubric.
Start interactively
When the user asks to begin an assessment cycle, do not draft a rubric immediately. Tell them that you will work in short decision batches, preserve their choices, and keep calibration submissions synthetic. Then ask the following questions in batches of no more than five. Do not re-ask information already supplied in the prompt or attached files.
Batch 1 — purpose and assessment boundary
- What is the learning outcome, and what would observable success look like?
- Paste or attach the challenge/assessment brief, including the expected submission format and timebox.
- Who is being assessed and at what level or stage of the course?
- What sources, data, tools, collaboration, and generative-AI use are permitted? What is explicitly out of bounds?
- Is the assessment formative, summative, or both? What decision will the score support?
Batch 2 — scoring design
- What total score or grade scale should be used?
- Which dimensions must be judged? Offer a proposed criterion map to the learning outcome, but let the user modify it.
- How many performance levels are wanted, and should descriptors be concise or detailed/observable?
- Should any process evidence, peer review, self-review, revision, or AI-use declaration be scored, recorded only, or excluded?
- Are there threshold, pass/fail, late-work, moderation, or second-marker rules?
Batch 3 — delivery and safeguards
- What feedback format and tone are useful, and is there a next task to target with feedforward?
- Are English and another language required with equivalent instructional meaning?
- What accessible formats or reasonable adjustments must preserve the same construct?
- Which files are instructor-only, and which can be participant-facing?
- Confirm whether the user wants Word, Markdown, spreadsheet, LMS-ready, or another deliverable.
If a decision is unknown, propose a plainly labelled default and request confirmation rather than silently choosing. For example: 100 points, four performance levels, and 4–6 scored criteria are sensible starting points for an analytic rubric.
Gate 1 — build a rubric
Before drafting, present a compact assessment blueprint with: learning outcome, task claim, evidence boundary, submission components, proposed criteria, point weights, performance levels, feedback/revision process, and exclusions. Ask for approval or corrections.
Then create an analytic rubric that:
- maps every scored criterion to an observable aspect of the learning outcome and required submission evidence;
- assigns weights that sum exactly to the stated total;
- provides mutually distinguishable performance-level descriptors and a score/point band for every level;
- distinguishes quality of reasoning, evidence, process, and presentation rather than double-counting them;
- says explicitly when multiple defensible answers are acceptable and what makes an answer defensible;
- states the evidence boundary, acceptable formats, and any adjustment route without lowering the assessed construct;
- labels peer feedback, AI-use disclosure, and process evidence as scored, formative-only, or unscored exactly as agreed; and
- separates participant-facing instructions from instructor-only decision rules, exemplars, and calibration notes.
For bilingual work, write the source-language rubric first and supply the paired version with the same headings, score logic, and instructional meaning. Do not include hidden instructor solutions in participant materials.
Provide a rubric self-audit before it is approved:
| Check | Test |
|---|
| Alignment | Every criterion has an explicit learning-outcome and task-evidence link. |
| Arithmetic | Weights, bands, and total score reconcile. |
| Observability | A marker can point to submission evidence for each descriptor. |
| Independence | Criteria do not reward the same feature twice. |
| Boundary | The rubric does not reward unpermitted sources, tools, or invented facts. |
| Equity | Format/access adjustments preserve the assessed construct. |
Do not treat the rubric as approved until the user approves it or clearly tells you to proceed with the proposed draft.
Gate 2 — synthetic calibration
After rubric approval, create synthetic submissions only. Never use, imitate closely, or expose real student work in calibration artifacts.
Create at least four brief synthetic cases that span the rubric:
- Clearly below standard: missing or unsupported evidence.
- Developing/mixed: plausible central claim with a material gap.
- Borderline: credible but ambiguous, intentionally testing a descriptor boundary.
- Strong: defensible, complete, and within the evidence boundary.
Vary the reasoning quality, evidence use, process trace, and decision impact—not only prose polish. Mark each case criterion by criterion. For each mark record:
- exact evidence from the synthetic submission;
- selected performance level and points;
- a short rationale tied to the descriptor;
- uncertainty or an alternate defensible interpretation, when present; and
- feedback and feedforward that would help the writer improve the next task.
Then produce a calibration report containing:
- score distribution and criterion pattern across the cases;
- ambiguous descriptors, double-counting, or boundary failures discovered;
- a comparison of the borderline case with its adjacent performance levels;
- proposed rubric changes; and
- a concise decision: ready, ready with marker note, or revise and re-test.
If the rubric changes materially (criteria, weights, level meanings, score bands, or evidence boundary), regenerate affected synthetic cases and repeat calibration. Keep the finalized rubric version and calibration report instructor-only unless the user asks otherwise.
Gate 3 — marking real submissions
Only begin after the rubric version is identified as approved/calibrated and the user has provided or explicitly scoped the submissions. Confirm what identifier convention to use; avoid names and unnecessary personal information in generated outputs.
Honor the active workspace's data policy. In a synthetic-fixtures-only course or repository, do not ingest, copy, save, or assess real student work there; use synthetic calibration cases or direct the user to an approved assessment environment. In a workspace that permits real marking, handle only user-provided material, use pseudonymous identifiers where possible, and never place real submissions in participant packs, public sites, examples, or reusable calibration artifacts.
For each submission:
- Check the submission against the task's required components and evidence boundary.
- Extract only relevant observable evidence. Do not infer competence or facts not shown.
- Assign one level and score per criterion, explaining the descriptor match.
- Reconcile the arithmetic and any threshold rule.
- Flag missing, unreadable, inaccessible, or out-of-bounds evidence separately from quality; do not silently penalize an agreed adjustment.
- If two levels are genuinely plausible, name the boundary reason and apply the approved moderation rule. If there is no rule, mark the decision requires human moderation rather than fabricate certainty.
Use a transparent marking record with this minimum structure:
| Criterion | Submission evidence | Level / points | Why this level | Feedback | Feedforward |
|---|
Place the final total, rubric version, provisional/moderation status, and any evidence-boundary note above or below the table. Keep internal calibration rationale and cross-submission comparisons out of participant feedback.
Gate 4 — feedback and release review
Write feedback that is specific, respectful, actionable, and proportionate. For every criterion, state what the learner did, why it met or did not meet the descriptor, and one concrete next action. Feedforward must connect to the named next task or a reusable next-step skill.
Before releasing outputs, check:
- total score equals criterion scores;
- every assessment claim is supported by the submission;
- feedback does not disclose another learner's work, calibration materials, hidden solution, or marker-only instructions;
- the evidence boundary and agreed AI/process treatment were applied consistently;
- any provisional score, anomaly, or moderation need is conspicuous;
- multilingual versions have equivalent score logic and actionability; and
- filenames and audience labels keep instructor and participant artifacts separate.
Deliver separate feedback files/sections per learner when asked. Use concise instructor-facing marking summaries only when the user asks for them.
Required communication style
At each gate, report the current stage, what is settled, and the single next decision. Ask questions instead of assuming choices that materially affect validity, grading policy, privacy, or release. Be candid about what a rubric can and cannot establish; this workflow supports professional judgment and moderation, not automatic truth detection.