
The OECD-DAC Evaluation Criteria: Six Criteria and Two Principles
The six DAC criteria as revised in 2019, what each one actually asks, the two principles governing their use, and why they are an evaluation instrument rather than a planning one.
Definition
The OECD-DAC evaluation criteria are six standards used to assess development interventions. They were established by the OECD Development Assistance Committee and are the closest thing the sector has to a common evaluation language: a terms of reference written anywhere in the world can say “evaluate against the DAC criteria” and be understood.
They were substantially revised in 2019. The revision — published as Better Criteria for Better Evaluation — added a sixth criterion, coherence, sharpened the definitions of the existing five, and, importantly, attached two principles governing how the set should be used. The revised criteria were adopted on 10 December 2019.
The six criteria
Relevance — is the intervention doing the right things? The extent to which the intervention’s objectives and design respond to the needs, policies and priorities of beneficiaries, of the country and of partner institutions — and continue to do so if circumstances change. That last clause is often overlooked: relevance is assessed against the context as it developed, not only as it was at design.
Coherence — how well does the intervention fit? The compatibility of the intervention with other interventions in the same country, sector or institution. It has two halves: internal coherence (with the organisation’s own other work and its policy commitments) and external coherence (with other actors’ interventions — complementarity, harmonisation, or unhelpful duplication).
Effectiveness — is the intervention achieving its objectives? The extent to which the intervention achieved, or is expected to achieve, its objectives and results, including any differential results across groups. The disaggregation clause is part of the criterion, not an optional extra: an intervention that hit its aggregate target while excluding a group is not straightforwardly effective.
Efficiency — how well are resources being used? The extent to which the intervention delivers, or is likely to deliver, results in an economic and timely way. Efficiency covers both cost and time; a project that delivered everything two years late has an efficiency finding whatever its unit costs look like.
Impact — what difference does the intervention make? The extent to which the intervention has generated or is expected to generate significant higher-level effects — positive or negative, intended or unintended, direct or indirect. The unintended and negative clauses are the part evaluators most often under-deliver on.
Sustainability — will the benefits last? The extent to which the net benefits of the intervention continue, or are likely to continue. Net benefits — the question is about what remains after the intervention ends, not whether the organisation that delivered it survives.
Coherence: what the 2019 addition changed
Coherence was added because the sector had moved. Interventions increasingly operate in crowded fields, and a project can be relevant, effective, efficient, impactful and sustainable while duplicating a neighbouring programme, contradicting the funder’s own policy position, or fragmenting a government system it should have been strengthening. None of the original five criteria caught that. Coherence does.
In practice it is the criterion most likely to be assessed thinly, because it requires the evaluator to look outside the intervention at what other actors were doing — which costs time that terms of reference rarely budget.
The two principles for use
The 2019 revision was explicit that the criteria are not a template, and stated two principles:
Principle one — apply them thoughtfully, to support high-quality and useful evaluation. The criteria must be contextualised: understood in relation to the specific intervention, the specific evaluation and the stakeholders who will use its findings. What “relevance” means for a humanitarian response in an acute crisis is not what it means for a ten-year institutional reform programme.
Principle two — how the criteria are used depends on the purpose of the evaluation. They should not be applied mechanistically. Which criteria are covered, and in what depth, should follow from the nature of the intervention and the needs of the evaluation’s users, and the evaluation questions and design should be tailored accordingly.
Both principles push against the same failure mode: six chapters of equal length, one per criterion, regardless of what the evaluation was commissioned to find out.
They evaluate; they do not plan
This is the most consequential misunderstanding about the criteria, and it appears often enough in real programme documents to be worth stating flatly.
The DAC criteria structure judgement after the fact. They never structure a plan.
You cannot design a programme “against the DAC criteria”. There is no relevance row to fill in, no coherence objective to deliver. The criteria are the lenses through which an evaluator, at mid-term or completion, forms a judgement about an intervention that was designed using something else — a theory of change, a logframe, a results framework.
What good design can do is anticipate them. If you know an evaluator will ask about differential results across groups, you disaggregate your indicators from the start. If you know coherence will be assessed, you document the mapping of other actors you did at design. That is designing so as to be evaluable — which is not at all the same as using the criteria as a planning framework.
Artefacts it produces
- Evaluation terms of reference organised around the criteria.
- An evaluation matrix — the working instrument, mapping each evaluation question to its criterion, its indicators or judgement basis, its data sources and its methods.
- The evaluation report, with findings, conclusions and recommendations traceable to the criteria.
- A management response and the recommendation tracker that follows it, which is where evaluations either change something or do not.
How it relates to the other frameworks
- The logframe and the results framework are what the criteria are applied to. Effectiveness is largely a question about whether the logframe’s outcome row happened.
- A theory of change makes an evaluation against these criteria far more useful, because it tells the evaluator which link to interrogate when a result did not materialise.
- Outcome harvesting is a way of generating the evidence base for effectiveness and impact when results were not pre-specified.
- Value for money overlaps with efficiency but is broader — the 4Es include equity, which the DAC criteria pick up inside effectiveness rather than as a separate heading.
Common mistakes
- Using them as a design framework. Covered above.
- Six equal chapters regardless of purpose. Directly contrary to principle two.
- Treating impact as “did the goal happen”. Impact includes unintended and negative effects. An evaluation that reports only against intended higher-level results has answered part of the question.
- Assessing sustainability as organisational survival. The question is whether the net benefits continue, which may not depend on the implementer continuing at all.
- Coherence as a paragraph. If the evaluator has not looked at what other actors were doing in the same sector and geography, coherence has not been assessed.
- Efficiency reduced to a spend rate. Budget execution is not efficiency. Efficiency is results per resource, and it includes timeliness.
- Skipping the management response. The criteria produce judgements; judgements that nobody is required to respond to produce nothing.
How Monival supports this
Monival is a monitoring system, not an evaluation one, and the distinction is worth being precise about: the criteria are applied by evaluators, usually independent ones, and no software applies them for you.
What Monival does is make an evaluation against them cheaper and better evidenced. Indicators carry an explicit means of verification, a data source and a collection frequency, so an evaluator asking “where does this number come from” has an answer in the system rather than in someone’s memory. Actuals are recorded by period against baselines and targets, which is the evidence base for effectiveness. Indicators fed from field data can be disaggregated by other questions on the same form, which is what the effectiveness criterion’s differential-results clause requires. Result nodes carry their assumptions, which is where an evaluator looks first when a result did not occur.
Submissions carry their collection metadata — timestamp, GPS where captured, the enumerator and the form version — and an audit log records what changed. That is the audit trail an evaluation data-quality assessment asks for, and it is the difference between a defensible finding and a contested one.
Monival does not produce evaluation reports, apply the criteria, or generate judgements. Evaluation design and delivery is a Sibasi consulting engagement.


