| name | monitoring-and-evaluation |
| description | Use when a proposal requires a theory of change, logframe, results framework, indicators, baselines, targets, learning, or evaluation. Unlike data-management, this skill defines performance questions and evidence use rather than the full data lifecycle. |
| metadata | {"portable":true,"compatible_with":["claude-code","codex"]} |
Monitoring and Evaluation
Acknowledgement: Shared by Peter Bamuhigire, techguypeter.com, +256 784 464178.
Use When
- Use this skill when the assignment explicitly needs monitoring, evaluation, results, indicators, or learning content.
- Load it when another proposal section needs this domain expertise.
Do Not Use When
- The task only needs formatting or proposer-profile selection.
- Another supporting skill is a closer fit for the assignment.
Required Inputs
| Artefact | Source | Required? | If absent |
|---|
| Programme logic, decisions, deliverables, and reporting obligations | ToR, logframe, donor framework | required | Stop indicator design and reconstruct the results chain. |
| Baselines, targets, sources, frequency, and owners | Client records and verified evidence | conditional | Mark values TBD and specify a baseline method. |
Workflow
Stop or block the workflow when a required input, permission, or acceptance basis is missing. Recover by revising the scope, obtaining evidence, or returning the narrowest qualified draft before proceeding.
- Identify where monitoring, evaluation, or results logic matters in the assignment.
- Read the local references only where they materially improve the output.
- Convert the guidance into proposal-ready frameworks, indicators, and reporting logic.
- Integrate the result into the target section and check for consistency.
Quality Standards
- Translate M&E theory into practical proposal language, outputs, and controls.
- Keep the approach specific to the client context and implementation reality.
- Preserve compatibility with existing repository workflows and file paths.
Anti-Patterns
- Creating indicators that available data cannot measure. Fix: name source, method, frequency, and owner.
- Counting activities as outcomes. Fix: distinguish inputs, outputs, outcomes, and impact.
- Inventing baselines or targets. Fix: mark TBD and specify how they will be set.
- Using averages that hide excluded groups. Fix: disaggregate where ethical and decision-useful.
- Reporting metrics with no decision. Fix: attach thresholds and response actions.
Outputs
| Artefact | Consumer | Acceptance condition |
|---|
| M&E framework or plan | Evaluator, programme manager, donor | Links results, indicators, definitions, sources, baselines, targets, frequency, ownership, learning, and decisions. |
Evidence Produced
| Evidence | Consumer | Acceptance condition |
|---|
| Indicator reference sheet and results chain | Controlled tables | Every indicator is reproducible, owned, and tied to a management question. |
Capability and Permission Boundaries
Read and search are required; any edit or external action remains within the explicit authority and permission boundary stated below.
Default to read-only analysis. Data collection, respondent contact, dashboard publication, target approval, and evaluation certification require explicit authority and ethical controls.
Degraded Mode
Without baseline, system, respondent, or verification access, produce an evaluability assessment and data plan. Mark measures unassessed; do not convert missing evidence into positive performance.
Decision Rules
| Need | Action | Risk avoided |
|---|
| Causal pathway is unclear | Build a theory of change first | Indicator shopping |
| Donor requires formal hierarchy | Build a logframe with assumptions | Untraceable reporting |
| Operational action is required | Add threshold and response owner | Passive dashboards |
Worked Example
For an e-commerce BDS programme, link selection and assistance to adoption and sales outcomes; baseline each company; define evidence windows; separate reach from attributable change.
SaaS Health Dashboard Pattern
For SaaS implementation and SaaS product-development engagements, layer on a SaaS health dashboard alongside the standard donor-grade logframe / theory of change:
- Three growth metrics (track high): activation rate, expansion ARR, net dollar retention.
- Three drag metrics (track low): gross churn, time-to-first-value, support cost per active account.
- Adoption metrics: DAU/MAU, feature-adoption breadth, business-unit-rollout count.
- Trust metrics: SLA performance, security incidents, regulator findings.
- Value-realisation metrics: against the value-at-stake baseline (Flavour A) or against the unit-economics plan (Flavour B).
Use internal-control measurement (A/B for content and feature variants, holdout for incrementality, propensity-matched for cohort programs) rather than industry-benchmark comparisons — benchmarks signal weak measurement and rarely survive CFO scrutiny.
AI-on-SaaS Dashboard Pattern
For AI-on-SaaS engagements, layer on an AI dashboard alongside the SaaS health dashboard:
- Eval drift: change in eval score between two consecutive refreshes; threshold breach > 5 pp triggers AI Governance Forum.
- Hallucination rate per feature: production-measured rate against the stated ceiling per feature; monthly cadence; weekly alerting.
- Abstain-correctness in production: should-abstain cases where the system abstained; healthy ≥ 0.95.
- Cost-per-call and cost-per-tenant at P50 / P90 / P99: weekly operational telemetry; quarterly tenant-cost-attribution review.
- p95 latency per feature: daily; SLO-bound.
- Cross-tenant probe pass rate: per probe run; 100 % is required, anything else is an incident.
- Adoption metrics: feature adoption, override rate (5–25 % healthy), user trust score (trend up), redress events (trend stable / down).
- AI incidents per quarter: hallucination liability, prompt-injection success, data leakage, sensitive-output, sub-processor incident; target zero critical.
- Retraining cadence: monthly eval refresh; quarterly golden-set refresh; quarterly red-team; annual independent red-team.
The AI dashboard is signed at the AI Governance Forum monthly and the SteerCo quarterly; an annual board-level Responsible-AI declaration consolidates it. See ../references/ai-on-saas-metrics-glossary.md for metric definitions and ../ai-on-saas-risk-and-responsible-ai/SKILL.md for the governance framing.
AI-Agent Dashboard Pattern
For agentic engagements, layer on an agent dashboard alongside (or replacing) the AI-on-SaaS dashboard:
- Task-success rate per use case per agent — binary at task level; trend signal.
- Intervention rate — primary autonomy signal; sustained breach for two weeks triggers Agent Safety Council.
- Irreversible-action incidents — target zero per quarter; any incident is a Class 1 alert.
- Scope breaches — target zero per quarter; any breach is a Class 2 alert.
- Audit-log completeness — measured monthly; SLA ≥ 99 %.
- Kill-switch drill log — quarterly drill executed and recorded with time-to-stop.
- Abstain-correctness on should-abstain subset — ≥ 0.95 healthy.
- Override rate — supervisor changes after the fact; trend signal for drift.
- Cost per outcome at P50 / P90 / P99 — weekly operational; quarterly tenant-attribution.
- Supervisor CSAT and queue headroom — adoption signals; floor 4.0 / 5 and 20 % headroom.
- Contestability case load and redress SLA performance — affected-party signal.
- Red-team scorecard — quarterly refresh.
- Multi-agent metrics (where applicable) — loop-detection events; orchestrator pause events; inter-agent disagreement clusters.
The agent dashboard is signed at the Agent Safety Council monthly and the SteerCo quarterly; an annual Responsible-AI Agent Commitment sign-off consolidates it. See ../references/ai-agent-metrics-glossary.md for definitions and ../ai-agent-risk-and-responsible-ai/SKILL.md for governance framing.
PDCA/QC Story learning loop
For each material result, write the QC Story in proposal-ready form: problem and scope, baseline, target, evidence/root cause, countermeasure, owner, timebox, result, residual gap, standardisation, and next cycle. Pair every indicator with a decision threshold and response owner. A dashboard is not an M&E system until it changes implementation, resource, safeguarding, quality, or sequencing decisions.
Apply PDCA at the cadence promised in the proposal: plan the result and data method; do the intervention; check observed change against baseline and counter-metrics; act by standardising, revising, escalating, or stopping. Include beneficiary/user feedback and disaggregation where ethical and useful. Preserve a learning register and evidence trail so the evaluator can see how findings are reviewed and converted into corrective action.
References
- Proposal skills router for repository-wide routing and mandatory quality gates.
- Local
references/ files when detailed frameworks or examples are needed.
- ../references/saas-metrics-glossary-for-proposals.md for SaaS vocabulary (ARR, MRR, NRR, magic number, Rule of 40, PSAR).
- ../references/ai-on-saas-metrics-glossary.md for AI vocabulary (eval, hallucination, abstain, citation, cost-per-call, drift).
- ../references/ai-agent-metrics-glossary.md for agent vocabulary (task-success, intervention, irreversibility, scope-confinement, audit completeness, kill-switch drill, supervisor CSAT).
- ../references/saas-business-case-and-roi-template.md for the value-realisation review template.
- ../references/saas-customer-success-engagement-package.md for the customer-success health-scoring composite.
- ../references/saas-lifecycle-email-program-proposal-template.md for per-program measurement standards.
Donor-funded and government assignments almost always require an M&E component. Proposals that present a clear results framework with measurable indicators score significantly higher than those offering vague "monitoring activities". This skill provides the M&E structures that proposal sections draw from.
When to Read This Skill
- The ToR mentions "monitoring and evaluation", "results framework", "logical framework", "KPIs", or "theory of change"
- The assignment has a capacity building, institutional reform, or programme implementation component
- Drafting methodology sections that need to describe how progress and impact will be measured
- When the ToR asks for a standalone M&E plan or framework as a deliverable
Frameworks
Logical Framework (Logframe)
The standard tool for donor-funded projects — World Bank, AfDB, UNDP, EU, and bilateral donors all expect it:
| Level | Description | Indicator | Baseline | Target | Means of Verification | Assumptions |
|---|
| Goal | The long-term impact this assignment contributes to | [Impact indicator] | [Current] | [Target] | [Data source] | [External conditions] |
| Purpose | The specific outcome this assignment will achieve | [Outcome indicator] | [Current] | [Target] | [Data source] | [Assumptions] |
| Outputs | The deliverables and products produced | [Output indicator] | [Current] | [Target] | [Data source] | [Assumptions] |
| Activities | The work done to produce the outputs | [Process indicator] | — | — | [Activity reports] | [Assumptions] |
Rules for logframes in proposals:
- Every indicator must be SMART: Specific, Measurable, Achievable, Relevant, Time-bound
- Baselines should be stated or flagged as "to be determined during inception"
- Means of verification must be practical — the data must actually be collectible
- Assumptions must be genuine risks, not filler ("government remains stable" is too broad)
Theory of Change
A narrative and visual model explaining the causal logic of the intervention:
- Problem statement: the specific problem being addressed
- Inputs: resources the project will deploy (funding, expertise, time)
- Activities: what the project will do
- Outputs: the direct products of the activities
- Outcomes: the changes in behaviour, capacity, or systems that result from the outputs
- Impact: the long-term change the outcomes contribute to
- Assumptions: the conditions that must hold for each causal link to work
Present as a flow diagram in the proposal, with a one-page narrative explaining the causal logic and key assumptions.
Results-Based Management (RBM)
The overarching philosophy that M&E supports. When referencing RBM in proposals:
- Focus on outcomes and impact, not just activity completion
- Reporting tracks results achieved, not just activities performed
- Adaptive management — use M&E data to adjust the approach during implementation
Key Indicator Types
| Type | Measures | Example |
|---|
| Input indicators | Resources deployed | Number of consultants mobilised, budget disbursed |
| Process indicators | Activities completed | Number of workshops conducted, reports submitted |
| Output indicators | Products delivered | System deployed, staff trained, policy drafted |
| Outcome indicators | Changes achieved | Processing time reduced, adoption rate, compliance rate |
| Impact indicators | Long-term effects | Revenue increase, service coverage, poverty reduction |
Proposals should include indicators at all levels but emphasise outcomes — this is what donors care about.
Data Collection Methods
When the proposal includes primary data collection:
| Method | Best For | Sample Size Guidance |
|---|
| Surveys / questionnaires | Quantitative data from large populations | Statistically representative sample |
| Key informant interviews | Qualitative insights from experts and leaders | 10–20 per stakeholder group |
| Focus group discussions | Community perspectives, user feedback | 6–10 participants per group, 4–8 groups |
| Document review | Policy analysis, process assessment | All relevant documents |
| Direct observation | Process verification, facility assessments | Purposive sample of sites |
| Administrative data | Routine monitoring, system usage | Census of available records |
Reference Files
Load these reference files for deeper guidance when writing M&E sections:
- references/results-frameworks-and-indicators.md — results chain, problem tree method, theory of change vs logframe, CREAM indicator criteria, indicator reference sheets, process/outcome/progression indicators (ILO), baselines and targets (four-step process), disaggregation requirements, results and resources framework (UNDG), performance management cycle, sector-specific indicator frameworks
- references/evaluation-design-and-methods.md — evaluation types (formative/summative/ex-post/real-time), four evaluation approaches (Civicus), OECD-DAC criteria (six), four-step evaluation design process (Seasons), evaluation matrix, impact evaluation designs (RCT, quasi-experimental, contribution analysis, outcome harvesting, Most Significant Change), quantitative and qualitative methods, mixed methods designs, sampling (probability and non-probability), sample size, data quality assurance (five dimensions), evaluation quality standards (UNDP), evaluation report structure, management response, joint evaluations, evaluation ethics
- references/monitoring-systems-and-reporting.md — monitoring vs evaluation distinction, monitoring plan development (six steps), M&E governance and working groups (UNDG), roles and responsibilities (IFRC), M&E capacity assessment, routine and periodic data collection systems, digital data collection tools, reporting hierarchy, traffic light (RAG) system, adaptive management and learning loop, decision triggers, periodic reviews, costed M&E plan (3–5% budget rule), open data and FAIR principles, common M&E pitfalls
- references/impact-evaluation-and-economic-analysis.md — impact evaluation core concepts (ATE/ITT/ATT/LATE), theory of change for IE (seven steps, funnel of attrition), RCT designs (nine types including cluster, pipeline, encouragement, factorial), quasi-experimental methods (DiD, synthetic controls, PSM, RDD, ITS, IV), design selection decision tree, power calculations (formula, cluster adjustment, software), data collection (six survey types, electronic platforms), IE process management (timeline, budget benchmarks, team composition), IE quality checklists (12-point design, 10-point data collection), cost-benefit analysis (NPV/B-C/IRR, shadow pricing, non-market valuation), discount rate selection, risk and uncertainty in CBA, distributional analysis (incidence matrix, five equity methods), multi-criteria evaluation frameworks (Planning Balance Sheet, Goals Achievement Matrix)
Generating a Standalone Section
When the ToR asks for a dedicated M&E plan or framework, generate a document covering:
- M&E approach and methodology
- Theory of change — narrative and diagram (see
references/results-frameworks-and-indicators.md section 2)
- Logical framework matrix with CREAM-quality indicators
- Indicator reference sheets — definition, data source, frequency, responsibility, disaggregation (see
references/results-frameworks-and-indicators.md section 3)
- Monitoring plan matrix with schedule and responsibilities (see
references/monitoring-systems-and-reporting.md section 2)
- Data collection plan — methods, sampling, digital tools (see
references/evaluation-design-and-methods.md sections 4–5)
- Data quality assurance framework (see
references/evaluation-design-and-methods.md section 6)
- Reporting framework with traffic light system (see
references/monitoring-systems-and-reporting.md section 5)
- Baseline and endline assessment methodology
- Evaluation design — type, OECD-DAC criteria, evaluation matrix (see
references/evaluation-design-and-methods.md sections 1–3)
- Adaptive management framework — decision triggers and review schedule (see
references/monitoring-systems-and-reporting.md section 6)
- Costed M&E budget (see
references/monitoring-systems-and-reporting.md section 7)
- Impact evaluation design — method selection decision tree, identification strategy, counterfactual approach (see
references/impact-evaluation-and-economic-analysis.md sections 3–4)
- Power calculations and sample size — formula, cluster adjustment, minimum detectable effect (see
references/impact-evaluation-and-economic-analysis.md section 5)
- Cost-benefit or cost-effectiveness analysis — NPV, shadow pricing, distributional analysis (see
references/impact-evaluation-and-economic-analysis.md sections 9–11)
- Analytics workstream — question framing, data inventory, quality profiling, analysis method, visualization, dashboard/report handover (see
../data-management/references/data-analytics-methodology-for-proposals.md)
Follow east-african-english standards throughout.