Designing Indicators That Survive a Donor Audit
Most reported indicators cannot withstand a data quality assessment — not because the data is wrong, but because the definition was never fixed. The discipline that prevents it.
Author
Sibasi M&E Team
Category
Indicators
Read Time
04 Mins read
Published Date
18 Mar, 2026

Inside the page
Share this
Subscribe for newsletter
An indicator fails an audit long before the auditor arrives. It fails at the moment somebody writes “number of beneficiaries reached” into a logframe and moves on.
Six months later, the field team counts anyone who attended a session. The partner counts households, not people. The last quarter’s figure counted the same person at three sessions; this quarter’s deduplicated. Nobody did anything dishonest. The number is simply not one number, and when a data quality assessment asks how it was constructed, there is no answer that holds for the whole reporting period.
This is the most common finding in donor data quality assessments across the programmes we have supported, and it is entirely preventable.
Fix the definition before you collect anything
The instrument for this in USAID practice is the Performance Indicator Reference Sheet (PIRS) — one page per indicator that fixes, before collection starts:
- The precise definition, including what counts and what explicitly does not.
- The unit of measure, and whether it counts people, events, households or something else.
- The disaggregation required, named in advance.
- The data source and the collection method.
- The frequency of collection and of reporting.
- Known data limitations — stated up front rather than discovered at audit.
- The baseline and the targets by period.
- Who is responsible for the number.
The name does not matter. FCDO partners call it an indicator methodology note; some organisations call it an indicator protocol. The discipline is what matters: one authoritative definition per indicator, written before collection, and changed only through a documented process.
If you take one thing from this piece: an indicator without a written definition is not an indicator. It is a phrase that different people will operationalise differently, and the divergence will be invisible until someone compares two reporting periods.
Baselines: measure them, or say you did not
Three failure modes, in descending order of how often we see them.
No baseline at all, with a target expressed as a percentage change. This is unassessable. If you do not know the starting value, “a 30% increase” is not a target.
Baseline set to zero because the project had not started. Legitimate for indicators counting project deliverables. Wrong for outcome indicators, which measure a condition in the population that existed before you arrived and will not have been zero.
Baseline collected on a different instrument from the follow-up. A baseline survey with a differently worded question, a different sampling frame or a different recall period does not baseline the indicator you are now reporting. This is the one that most often destroys an endline comparison, and it is discovered too late to fix.
Where a baseline genuinely cannot be measured, say so explicitly in the indicator definition and state what will be used instead. A documented limitation is an acceptable audit finding. An undocumented one is not.
Disaggregation is a design decision, not a reporting one
You can only disaggregate by something you captured. This sounds obvious and is routinely ignored: teams report a total, are later asked for a sex or disability breakdown, and discover the field was never on the form.
Decide the disaggregation dimensions at design and put them in the indicator definition. The standard set — sex, age band, location, disability status — should be the default question, and any decision to omit one should be a decision rather than an oversight.
There is a second reason beyond compliance. The OECD-DAC effectiveness criterion explicitly includes differential results across groups, and the equity E in a value for money assessment cannot be evidenced at all without disaggregated data. An aggregate number that hit its target while excluding a group is not a success, and without the breakdown nobody can tell the difference.
Be careful with age bands. Bands that do not match the funder’s reporting categories cannot be re-cut afterwards — you cannot split a 15–24 band into 15–19 and 20–24 once collection is done. Capture date of birth or exact age where you can, and band at analysis.
Means of verification must be a source you actually have
The means of verification column asks where the evidence comes from. Two tests it must pass:
Does the source exist, in the form you need? Naming a national health information system as the MoV for an indicator that system does not disaggregate the way your funder requires is a design failure that surfaces at the first report.
Can you afford it, at the stated frequency? A quarterly household survey named as the MoV for six indicators is a budget line, not a checkbox. If the budget does not carry it, the indicator will be reported from something else, undocumented — which is precisely the divergence an audit finds.
What a data quality assessment actually tests
The standard dimensions, and what each one means in practice:
- Validity — does the indicator measure what it claims to? Attendance is not learning.
- Reliability — would the same method produce the same result if repeated? This is what a stable written definition protects.
- Timeliness — is the data available in time to inform decisions, or only in time to report?
- Precision — is the margin of error small enough for the use? A number derived from a small purposive sample should not be presented as a population estimate.
- Integrity — is the data protected against manipulation, deliberate or otherwise? Who can change a submitted value, and is that change recorded?
Integrity is where paper-and-spreadsheet systems fail hardest. If the reported figure lives in a spreadsheet that six people can edit and no version history exists, there is no answer to “who changed this and when”, and the auditor will write that down.
The practical checklist
Before collection starts on any indicator:
- A written definition exists, naming what counts and what does not.
- Unit of measure is stated.
- Disaggregation dimensions are named and the corresponding questions are on the form.
- The data source is named and access to it is confirmed.
- Collection frequency is stated and budgeted.
- A baseline value exists, or its absence is documented with the reason.
- Targets are set by period, not only at end of project.
- One named person is responsible for the number.
- Known limitations are written down before anyone asks.
Nine lines per indicator. It is a morning’s work for a typical logframe, and it is the difference between an audit that confirms your reporting and one that qualifies it.
Monival holds indicators as organisation-level records — baseline, target, collection frequency, data source and means of verification on the indicator itself, with actuals recorded by period. Indicators can be fed directly from form questions as a count, sum or distinct count, disaggregated by other questions on the same form, so the breakdown is a property of the design rather than something reconstructed later.