Impact evaluation
Module F — difference-in-differences impact evaluation. Does SEEP actually work?
The four numbers below answer: did SEEP students improve more than similar students who didn't join the programme? Net effect = SEEP gain minus control gain (a positive number means SEEP worked). Treatment n / Control n = how many students are in each group. Significant subjects = how many subjects show a statistically reliable improvement (p < 0.05).
Net effect of +5.18 points across 228,960 SEEP students vs 28,840 control students, with all 12 subjects statistically significant. This is as clean a result as impact evaluations get — SEEP students improved meaningfully more than their matched peers, and the signal is consistent across every subject tested.
This is the result to show the Minister and the press. Publish the detailed diff-in-diff table below as evidence, maintain current teacher and content delivery, and focus managerial energy on the weakest divisions (see the division heatmap on /overview).
SEEP students gained +6.7 points on average, versus only +1.5 for the natural-control group — a +5.2-point gap that can be attributed to the programme.
This card lines up three groups side-by-side. "SEEP treatment" are students who attended SEEP sessions. "Natural control" are students eligible (scored ≤40%) who did not attend, which gives us a clean comparison. Each cell shows how many students were in the group, their average pre-test score, their average post-test score, and the point change. The programme's real impact is the DIFFERENCE between the two gains.
The SEEP gap is a strong 5.2 points — this is a large, textbook-clean programme effect. Students who attended SEEP improved dramatically more than students who didn't. This is the number your Minister should see.
This is your clean headline comparison. Use it in stakeholder decks alongside the per-subject diff-in-diff below.
Across all 12 subjects, the average treatment gain is +6.2 points vs +1.5 for control — a 4.7-point programme effect per subject.
Each pair of bars is one subject. The green bar is how much SEEP students improved. The orange bar is how much the control group improved over the same period. The gap between them is the pure SEEP effect — the rise that cannot be explained by 'kids just getting older and better anyway'.
Green dominates orange in every single subject. There is no subject where SEEP failed to beat the control — the programme is uniformly effective across the curriculum.
| Subject | Treatment | Control | Net | 95% CI | p-value | d | n (t/c) | Sig? |
|---|---|---|---|---|---|---|---|---|
| History (F4-5) | +9.3 | +1.5 | +7.85 | [7.7, 8.0] | <0.001 | 1.07 | 31,483/3,984 | Yes |
| History | +8.7 | +1.5 | +7.14 | [7.0, 7.3] | <0.001 | 1.03 | 17,024/2,148 | Yes |
| English (F4-5) | +7.8 | +1.5 | +6.30 | [6.2, 6.4] | <0.001 | 1.03 | 31,483/3,984 | Yes |
| English | +7.4 | +1.5 | +5.93 | [5.8, 6.1] | <0.001 | 1.01 | 17,024/2,148 | Yes |
| Biology | +7.2 | +1.5 | +5.62 | [5.4, 5.8] | <0.001 | 0.99 | 8,733/1,078 | Yes |
| Science (F4-5 Arts) | +6.5 | +1.5 | +5.07 | [5.0, 5.2] | <0.001 | 0.98 | 31,483/3,984 | Yes |
| Science | +5.9 | +1.5 | +4.40 | [4.3, 4.5] | <0.001 | 0.94 | 17,024/2,148 | Yes |
| Mathematics (F4-5 Arts) | +5.3 | +1.5 | +3.80 | [3.7, 3.9] | <0.001 | 0.91 | 31,483/3,984 | Yes |
| Chemistry | +4.6 | +1.5 | +3.13 | [3.0, 3.3] | <0.001 | 0.84 | 8,733/1,078 | Yes |
| Mathematics | +4.3 | +1.5 | +2.84 | [2.7, 2.9] | <0.001 | 0.81 | 17,024/2,148 | Yes |
| Mathematics (F4-5 Sci) | +4.0 | +1.5 | +2.57 | [2.4, 2.7] | <0.001 | 0.79 | 8,733/1,078 | Yes |
| Additional Mathematics | +3.4 | +1.5 | +1.93 | [1.8, 2.1] | <0.001 | 0.69 | 8,733/1,078 | Yes |
Each row is one subject. Treatment and Control are the average points gained pre → post for SEEP students and non-SEEP students respectively. Net is the difference, i.e. the pure SEEP effect for that subject. The 95% CI is the range of plausible values for the true effect — if it excludes 0, the result is reliable. p-valueunder 0.05 (or <0.001) means the effect is unlikely to be chance. dis Cohen's effect size (0.2 small, 0.5 medium, 0.8+ large). n (t/c) is how many SEEP vs non-SEEP students contributed to that row.
A green “Yes” in the Sig? column means that subject clears the statistical bar — it is safe to tell stakeholders “SEEP worked for this subject.” Every row being “Yes” is the best possible pattern and means the programme is improving outcomes across the entire curriculum, not just one or two lucky subjects.