Reading a programme evaluation: five questions funders should ask
16 JUL 2026 · 5 MIN READ · SOCIAL BRIDGE CONSULTING
Funders commission evaluations to make decisions: continue, scale, redesign, exit. Yet the reports that arrive are often read the way one reads a certificate — a glance at the executive summary, a check that the tone is broadly positive or suitably critical, and on to the annexures for the photographs. This is understandable; evaluation reports can run to two hundred pages, and programme officers are stretched. But it hands the real decision to whoever wrote the executive summary.
An evaluation is not a verdict. It is an argument built on choices — about design, sampling, comparison, measurement and interpretation — and the strength of its conclusions depends entirely on the strength of those choices. You do not need statistical training to interrogate them. You need five questions, asked stubbornly, of every evaluation that crosses your desk.
1. Compared to what?
Every claim of effect is a comparison, stated or smuggled. "Institutional deliveries rose during the project period" compares the endline to the baseline — but so did everything else in those years: other schemes, road construction, a new medical college, secular trends visible across the whole state. The question to ask is: what would have happened in these villages without the programme, and how does the evaluation know?
Strong designs answer with a genuine comparison — control areas, matched non-participants, phased rollout. Weaker designs answer with before-after alone, which is sometimes the only feasible option but should be labelled as such. What you are checking for is not methodological purity; it is honesty. An evaluation that says plainly "we cannot separate the programme's contribution from the district-wide trend, but the following pattern suggests..." is more useful than one that claims attribution it cannot support. Be especially wary when improvements in programme areas mirror improvements everywhere else, and the report does not mention it.
2. Who was actually asked — and who was left out?
Turn to the methodology section and stay there longer than feels natural. How were respondents selected? If the implementing agency assembled the groups, arranged the village visits, or accompanied the interviewers, the findings tilt toward the programme's friends — not through anyone's dishonesty, but through the ordinary physics of who shows up when the sarpanch is told visitors are coming.
Ask specifically about three groups that evaluations quietly omit: dropouts (people who joined the programme and left — often the most informative respondents in the entire study), non-participants in programme villages (why did they stay out?), and the hardest-to-reach hamlets, which fieldwork skips exactly when time runs short, and which are usually where the programme's claims are weakest. If the report cannot say what its non-response and replacement rates were, treat every percentage in it as softer than it looks.
3. Do the numbers and the narratives agree — and what happens where they don't?
Most evaluations in our sector are mixed-methods: a survey plus interviews and focus groups. Read them against each other. When the survey says uptake is high and the focus groups are full of complaints about access, the report owes you a reconciliation, not a diplomatic silence. The points of friction between numbers and narratives are frequently where the real findings live — the entitlement is technically delivered but practically degraded; the training happened but changed nothing downstream.
Also check a humbler thing: do the numbers agree with each other? Tables that do not sum, denominators that shift between chapters, percentages of unstated bases — these are small tells with large implications about how carefully the analysis was done.
4. Does the report explain the why, or only the how much?
A funder's decision is rarely "did it work?" in the abstract; it is "should we do this again, elsewhere, at scale?" That decision needs mechanism, not just magnitude. An evaluation that reports outcomes without explaining the pathway — which components did the work, under what local conditions, with what dependence on particular people or on an unusually supportive district administration — cannot support a scaling decision, however positive its findings.
Ask where the programme worked least well, and whether the report explains the variation. Variation is the most honest teacher in evaluation: a programme that succeeded in blocks with strong panchayat engagement and stalled without it has just told you the true cost of replication. Reports that present a single averaged result across wildly different contexts are averaging away exactly the information you paid for.
5. Whose voice is missing from the findings?
Finally, read the report for its silences. Are frontline workers — the ASHAs, anganwadi workers, cluster coordinators who actually carried the programme — quoted on what made their work harder or easier, or do they appear only as counts of trainings attended? Are women's responses reported separately where the intervention touched household decisions, or folded into household-level answers a male head provided? Did anyone ask the block- and district-level officials who will inherit the programme after the project staff leave?
An evaluation can be technically competent and still be written entirely from the vantage of the project office. Those reports systematically overstate sustainability, because sustainability lives precisely in the voices they left out.
None of these questions requires you to reject imperfect evaluations — in real field conditions, all evaluations are imperfect. The discipline is proportioning your confidence to the evidence: knowing which findings can bear the weight of a scaling decision, which support only a cautious continuation, and which are, on inspection, decorated anecdote. Funders who ask these five questions consistently also get better evaluations over time, because evaluators write for the readers they expect. Commissioning terms of reference that demand comparison logic, sampling transparency and disaggregated voice — the standards our own evaluation practice works to — is the other half of the same discipline.
If you have an evaluation on your desk and want a second pair of eyes on what it can and cannot support, we are easy to reach.