GuideSocial Sector
How to build an impact measurement framework, and the data system that runs it
An impact measurement framework links what your organization does to the changes it expects in people's lives, names an indicator for each link, and says how and when each indicator is collected. Build it in this order: a theory of change with explicit assumptions, a short indicator set drawn from shared catalogs where possible, a collection plan that respects respondents' time, and a data system that keeps case, survey and activity records joined. Evaluation depth should match the decisions the evidence will inform.
On this page
- Outputs, outcomes and impact, and why contribution beats attribution
- From theory of change to indicators: where the assumptions sit
- Choosing indicators against the five dimensions of impact
- Seven moves from a blank page to a working measurement system
- A hypothetical youth employment program picks its indicators
- Qualitative evidence at scale with AI-assisted coding
- An evaluation ladder: matching rigor to the decision at stake
- How impact frameworks drift into vanity metrics
- Questions and answers
- Sources
Outputs, outcomes and impact, and why contribution beats attribution
Most program reports count outputs: sessions run, food parcels handed out, people trained. Outputs are necessary for management, but they say nothing about whether anyone is better off. Outcomes are the changes in knowledge, behavior, income, health or safety that follow, and impact is the share of that change that would not have happened without you.
That last clause is where many frameworks overreach. Claiming attribution, that your program alone caused a change, usually needs a comparison group. Most organizations can make a credible contribution claim instead: a plausible causal story, evidence at each link and an honest account of other influences. The Impact Frontiers five dimensions of impact, What, Who, How Much, Contribution and Risk, give a shared vocabulary for that story1.
From theory of change to indicators: where the assumptions sit
Each arrow is an assumption. Your indicators should tell you whether the arrows hold, not just whether the boxes were filled.
- Inputs
Staff, funding, partners and data you commit to the program.
- Activities
What you do: mentoring, cash transfers, clinics, training sessions.
- Outputs
Countable products of activity, tracked in case or activity records.
- Short-term outcomes
Early changes such as skills gained or a debt plan agreed, measured by survey or assessment.
- Long-term outcomes
Sustained change such as stable work or housing, measured at follow-up.
- Contribution to impact
The change you can credibly link to your work, set against what would have happened anyway.
Choosing indicators against the five dimensions of impact
Start with shared catalogs and add custom indicators only where your theory of change needs them.
| Dimension | Question it answers | Indicator sources to try first | Typical data source |
|---|---|---|---|
| What | Which outcome changes, and does it matter to the people affected? | IRIS+ core metric sets by theme2; SDG indicator framework3 | Outcome survey or validated assessment |
| Who | Who experiences the change, and how underserved were they? | Baseline demographic and need fields defined in your intake form | Case records and registration data |
| How Much | How many people, how deep a change, for how long? | Custom depth scales plus follow-up rates | Repeat surveys at fixed intervals |
| Contribution | Is the change better than what would have happened anyway? | Comparison with a baseline, a waiting list or published local rates | Evaluation design, not routine monitoring |
| Risk | What could make the impact differ from expectations? | Drop-out, unintended effects and complaints | Feedback channels and exit interviews |
Dimension names follow Impact Frontiers. The indicator sources are starting points; most frameworks end up mixing catalog and custom indicators.
Seven moves from a blank page to a working measurement system
Draft the theory of change with the people who deliver it
Run a workshop with frontline staff and, ideally, participants. Write each causal step as a sentence and list the assumptions under it. A logframe can follow if a funder requires one.
Cap the indicator set
Aim for a handful of outcome indicators per program, each tied to an assumption you are unsure about. Map each to an IRIS+ metric or SDG indicator where the definition genuinely matches.
Design collection around respondent burden
Decide baseline timing, follow-up intervals and sampling. Use offline-capable mobile forms such as ODK4 or KoboToolbox where connectivity is poor, and drop any question nobody will analyze.
Join the records
Give every participant a stable pseudonymous ID so case records, survey responses and activity logs can be linked without names. Store them in one small warehouse, even a managed database with scheduled imports.
Add quality checks at entry and weekly
Validate ranges and skip logic in forms, then run weekly checks for duplicates, impossible dates and enumerators whose answers look too uniform.
Build a few decision dashboards
One view per decision-maker: managers see take-up and drop-out, trustees see outcomes by group, fundraisers see the indicators their funders track.
Review the framework every program cycle
Retire indicators that never changed a decision and add ones the last cycle showed you lacked.
A hypothetical youth employment program picks its indicators
Qualitative evidence at scale with AI-assisted coding
Open-text answers, interview transcripts and hotline messages often hold the best evidence of how change happens, and they are the first thing dropped when staff are busy. Language models can apply a codebook to thousands of responses and surface themes in several languages, which makes qualitative evidence routine rather than occasional.
Keep people in charge of meaning. Analysts write and own the codebook, the model suggests codes with a short quote as justification, and a person double-codes a random sample each round to check agreement before results reach a report. Remove names and identifying details before any text leaves your systems, and follow the safeguards in our beneficiary data checklist. The same evaluation discipline applies here as for any model; see how to build LLM evaluation sets.
An evaluation ladder: matching rigor to the decision at stake
| Approach | What it can show | When it is proportionate | Main cost |
|---|---|---|---|
| Routine monitoring | Whether delivery and outcomes move in the expected direction | Always; it runs the program | Staff time and data discipline |
| Pre-post with follow-up | Change over time for participants, without a comparison | Early programs and small charities | Follow-up attrition |
| Quasi-experimental design | Contribution, using waiting lists, matching or phased rollout | Before scaling or seeking larger grants | Analytic skill and a credible comparison group |
| Randomized controlled trial | Attribution with the strongest confidence | When a scale-up or policy decision depends on it | Budget, ethics review and long timelines |
Most organizations live on the first two rungs and commission the third for specific decisions.
How impact frameworks drift into vanity metrics
Reach reported as impact
Early signalHeadline figures are people reached, with no outcome attached.
MitigationPair every reach figure with one outcome indicator and its follow-up rate.
Indicators chosen for availability
Early signalEvery indicator comes from data you already had.
MitigationCheck each assumption in the theory of change has at least one indicator testing it.
Survivorship in follow-up
Early signalOnly participants still in touch answer the follow-up survey.
MitigationReport response rates and compare responders with non-responders at baseline.
Funder-by-funder frameworks
Early signalEach grant adds its own indicators and nothing is ever retired.
MitigationKeep one master framework and map funder indicators onto it; see grant reporting automation.
Questions and answers
How much of a program budget should go on measurement?
There is no single right share. Size measurement to the decisions it informs: routine monitoring should be built into delivery costs, while a comparison-group evaluation is a separate line justified by a scale-up or major funding decision. If data is collected that no one uses, cut it before cutting analysis. Many funders will cover reasonable evaluation costs when they are written into the proposal.
Can a small charity build a credible impact framework?
Yes. A small charity needs a clear theory of change, three to six outcome indicators, a baseline and one follow-up point, collected with free mobile forms and kept in a single spreadsheet or small database. Credibility comes from honest definitions and response rates, not from sophisticated methods. Comparison-group evaluations can be commissioned later, often with university partners.
Should we align our indicators with the Sustainable Development Goals?
Align where the definition genuinely matches, because some funders and impact investors report against the SDGs. The official SDG indicators are designed for national statistics, so most organizations map their own outcome indicators to an SDG target rather than collecting the national indicator itself. Record the mapping in your indicator sheet so it is applied consistently.
What is the difference between a theory of change and a logframe?
A theory of change explains how and why change is expected to happen, including assumptions and context. A logframe is a compact table, often required by institutional donors, listing objectives, indicators, means of verification and assumptions. Build the theory of change first, then derive the logframe from it so both tell the same story.
Sources
- Five Dimensions of Impact — Impact Frontiers · checked 10 October 2026
- IRIS+ System — Global Impact Investing Network (GIIN) · checked 10 October 2026
- SDG Indicators: global indicator framework — United Nations Statistics Division · checked 10 October 2026
- ODK: collect data anywhere — ODK · checked 10 October 2026