This page is written like a statistical analysis report: a plain-language summary first, for anyone who wants the finding without the statistics, then the full method and evidence behind it, laid out so it can be checked rather than just read. Every number on this page is generated automatically from real monitored rain and river gauge data - nothing here is hand-picked or adjusted after the fact.
This analysis finds two different things depending on the size of the rain event. For large rain events - the ones that actually matter for flood risk - the Bucklebury gauge has been rising significantly less since the NFM works were built. Our best estimate is a reduction of about 7.1cm at a large rain event, and this result is statistically significant - in plain terms, the data are strong enough that this is very unlikely to just be chance or noise in the measurements. Based on the data collected so far, the true reduction is most likely somewhere between about 3.8cm and 10.3cm - in other words, we can't pin down the exact size of the improvement, but we can be confident there is one. For small-to-typical rain events - which don't put anyone at risk of flooding - the picture is much less clear: our best estimate is a small reduction of about 1.1cm, but there isn't yet enough data to be confident that's a real effect rather than normal variation - the true figure could plausibly be anywhere from a moderate improvement to a small worsening, so we're not reporting this one as a finding either way.
It may not be possible to attribute this to the NFM works alone. Other changes have also been made in the valley upstream of Bucklebury since 31 May 2023, such as alterations to Bucklebury Common's tree and vegetation cover, upstream of Elmwood - additional leaky dams were built there with the intention of offsetting that reduced cover. Both this result and the "before" baseline it's measured against also only cover the period since the Bucklebury Flood Alleviation Channel was built in 2011 - see Methods for why.
Key points:
The Natural Flood Management (NFM) interventions upstream of Bucklebury in the Elmwood and Redhill Copse areas - referred to throughout this page as the NFM works - comprise two schemes that sit on streams feeding into the River Pang upstream of the Bucklebury gauge: the Elmwood NFM scheme (24 leaky dams, log barriers and "dragon's teeth" on the Elmwood Stream, built 18 May - 25 September 2020), then the Redhill Copse scheme (13 leaky dams on Osgood's Gully, built 1-31 May 2023). Both sets of NFM interventions are on streams that flow into the River Pang upstream of the EA's Bucklebury gauge, so this page tests whether the NFM works, taken together, have changed how the river responds to rain - not either scheme in isolation, and not whether either scheme performed as designed in isolation.
PVFF's 2021 report (PVFF NFM Final Report - Statistical Analysis v1.0) first tested this question with a similar regression-based method, but had only 16 usable post-scheme rain events at the time and explicitly could not draw a conclusion either way. This page re-runs an updated version of that method (see Methods for what changed and why) against every qualifying rain event recorded since, generated automatically from the underlying rain and river gauge data rather than computed by hand.
Research question: have the NFM works - and whatever else has changed in the catchment over the same period - measurably changed how much the Bucklebury gauge rises in response to a given rain event, relative to before either scheme existed?
Note on independence: PVFF built and supports the NFM works. This page reports what the data show, without adjustment, regardless of whether the result matches what PVFF hoped they would achieve - see the result below.
| Scheme | Stream | Affects gauge | Built |
|---|---|---|---|
| Elmwood NFM - 24 leaky dams, log barriers & "dragon's teeth" | Elmwood Stream | Bucklebury | 18 May 2020 (began) / 25 Sep 2020 (complete) |
| Redhill Copse NFM - 13 leaky dams | Osgood's Gully | Bucklebury (same gauge as Elmwood) | 1 May 2023 (began) / 31 May 2023 (complete) |
Both schemes feed the Bucklebury gauge's catchment, so this page compares events from before Elmwood began (restricted to after the 2011 flood alleviation channel - see Methods 2.4) against events from after Redhill Copse was complete - the period reflecting both schemes together - skipping over the intermediate "Elmwood only" window entirely so it isn't muddled into either side of the comparison. The Bucklebury Waven Field scheme (2025) manages village surface water and isn't expected to affect the river gauge, so isn't tested here.
A "rain event" is a spell of rain preceded and followed by at least 8 dry hours, with total rain below 5mm dropped as noise (no cap on event length). For each event we record its length, total rainfall, peak 15-minute intensity, and rain in the preceding 7 days (antecedent wetness) - the same rules as PVFF's 2021 report. Full detail: overview page.
How these compare to recognised practice: separating a continuous rain record into discrete events using a minimum dry-gap threshold is a standard technique in hydrology, usually called the "minimum inter-event time" (MIT/IETD) method - there is no single agreed threshold value in the literature (studies use anywhere from a few hours to over a day depending on climate and purpose), so an 8-hour gap is a defensible, stated choice within that range rather than an invented one (see Discussion for how sensitive the result is to that choice). Event duration and peak short-interval intensity are standard rainfall-event characteristics (peak intensity measures such as maximum 15/30-minute rate are the basis of established rainfall-erosivity metrics). Our 7-day antecedent rainfall sum is a simplified relative of the classical Antecedent Precipitation Index (API), which typically weights recent rain more heavily via a decay function rather than summing it flatly - a genuine simplification, not a claim of equivalence.
A frost covariate was tried and dropped: PVFF's 2021 report flagged several unexplained high outliers that followed cold snaps, so a "hard frost in the days before the event" flag was added and tested at several lookback windows. Every version failed for a different reason: a full week let clearly-thawed frost count (and extrapolated wildly onto a genuine multi-day cold snap in January 2024 that Bucklebury's "before" data had nothing comparable to - see Discussion), while shorter, more physically defensible windows left zero frost-affected events in Bucklebury's qualifying "before" sample at all, regardless of length - its few frost-adjacent events never reached the minimum gauge level to qualify. Cold-snap events remain a known, unmodelled source of residual variance in this analysis.
Above a critical level of approximately 74cm, the river overflows upstream of the Bucklebury gauge and the raw reading under-represents the true level - this is why the original 2021 report excluded every event peaking above 70cm entirely. This analysis corrects for that with a linear adjustment above the critical level, calibrated against one independent point (the 2014 event's citizen-measured village depth against the simultaneous raw gauge reading). This lets exceptional events be included in the dataset generally, rather than silently dropped, though not every such event ends up used in every comparison on this page (see 2.4 below for one that doesn't) - see limitations for what this correction does and doesn't establish.
Events are dropped where the West Berkshire Groundwater Scheme (EA drought-relief pumping into the Pang above Bucklebury) was running at any point in the event's response window, since this raises the gauge reading independent of rainfall.
The Bucklebury Flood Alleviation Channel was built in 2011, changing how the catchment responds to rain independently of any NFM work. This was checked directly: fitting a single model across the whole pre-Elmwood record (2002-2020) with a pre/post-2011 indicator shows the rain-response slope is significantly steeper after the channel was built (interaction p = 0.009), and that same single model systematically under-predicts pre-2011 events and over-predicts post-2011 ones. Left uncorrected, the "before" baseline used on this page would partly reflect an unrelated hydraulic structure rather than only the catchment condition the NFM works were then built into. The "before" period used throughout this page is therefore restricted to events from 1 May 2011 onwards - after the channel was in place - so the only material difference between the "before" and "after" samples is the NFM works themselves. This is the same principle the existing Bucklebury flood-modelling work already applies when comparing pre- and post-channel behaviour.
This restriction removes the two largest events on record, 19 and 21 July 2007 (86.8mm and 6.4mm of rain, adjusted rises of 102.2cm and 41.6cm), from the "before" sample - both predate the channel by several years. No event of that scale has occurred since 2011 in either the "before" or "after" period, so the large-rain-event result reported below reflects the range of large rain events actually observed since 2011, not a 2007-scale extreme; see Discussion.
A multi-linear regression predicts the gauge rise from the event characteristics above plus month of year. The primary test pools every pre-Elmwood and post-Redhill-Copse event into one regression with a "post-NFM" indicator (and its interaction with total rainfall, so the effect isn't forced to be a flat shift), giving a direct significance test of the combined effect - rather than an informal count of events that came in higher or lower than a separately-fitted prediction.
Three rain gauges feed this catchment (Yattendon, Thatcham, and a blend of the two). The gauge used is picked once, using only pre-Elmwood ("before") data, by 5-fold cross-validated R² (a fairer measure of real predictive skill than in-sample R², since it's tested on data the model wasn't fitted on) - not by whichever gauge happens to produce the most favourable "after" result. The gauge choice cannot have been influenced by the outcome being tested, since it's fixed before any post-scheme event is looked at.
If none of the three gauges reaches a cross-validated R² of 50% on the "before" data, the before-model isn't trustworthy enough to build an after-comparison on, and none is reported - rather than show a comparison against a prediction we don't trust. This same threshold, applied the same way, is why the Tidmarsh hypothesis reports no result at all: it is not a bar set to fit Bucklebury's data.
We tested this model against all 3 rain gauges (Yattendon, Thatcham, and a blend of the two) on pre-scheme events, and picked whichever fit best by cross-validated R² (tested on held-out events, a fairer measure of real predictive skill than plain in-sample R² - see below) - Yattendon was the best fit, and everything below uses only that gauge. R² is the share of event-to-event variation in the rise that the model explains from rainfall, event length, intensity, antecedent wetness and month (0% = no better than guessing the average, 100% = perfect).
| Rain gauge | Pre-scheme events | R² (in-sample) | R² (cross-validated) |
|---|---|---|---|
| Yattendon (used) | 167 | 75% | 68% |
| Thatcham | 161 | 75% | 66% |
| Blended | 154 | 76% | 65% |
Cross-validated R² is normally a bit lower than plain, in-sample R² - that's expected: plain R² can look better than a model really is because it's measured on the same data used to fit it, whereas cross-validation tests each event using a model that never saw it. We use the more honest, cross-validated figure to pick the gauge, not the in-sample one.
Fitted on the pre-Elmwood events, since that's the baseline the "after" period is compared against. The bars show each factor's standardized effect - scaled so they're directly comparable to each other regardless of their original units (mm of rain vs. hours vs. mm/15min) - so a longer bar means that factor moves the predicted rise more. "Sure" / "unsure" reflects whether that factor's effect is confident enough to rely on.
| Factor | Relative effect | Effect size | Confidence |
|---|---|---|---|
| Total rain (mm) | +0.83 | sure | |
| Rain in preceding 7 days (mm) | +0.22 | sure | |
| Peak rain rate (mm/15min) | -0.20 | sure | |
| Length of event (hrs) | -0.13 | sure |
Every month is measured relative to January (the "baseline" row, fixed at zero by definition).
| Month | Relative effect | Effect size | Confidence |
|---|---|---|---|
| Jan (baseline) | +0.00 | baseline | |
| Feb | -0.30 | sure | |
| Mar | -0.42 | sure | |
| Apr | -0.68 | sure | |
| May | -1.10 | sure | |
| Jun | -0.96 | sure | |
| Jul | -0.97 | sure | |
| Aug | -1.31 | sure | |
| Sep | -1.02 | unsure | |
| Oct | +0.14 | unsure | |
| Nov | +0.14 | unsure | |
| Dec | -0.28 | unsure |
The combined effect interacts significantly with rain amount (coefficient -0.3718, 95% CI [-0.5554, -0.1882], p < 0.001) - it is not a flat shift, so a single average figure across all event sizes would be misleading. The table below reports the estimated effect, its own 95% confidence interval, and its own significance test, separately at a typical event and at a large rain event - large rain events matter most here, since they are the ones with real flood risk; small events are broadly the ones that matter least.
| Event size | Rain total | Estimated effect | 95% CI | Result |
|---|---|---|---|---|
| Typical (median) | 11.5mm | -1.08cm | [-2.77cm, 0.60cm] | not statistically significant (p = 0.207) |
| Large rain event (90th percentile) - flood-relevant | 27.6mm | -7.07cm | [-10.33cm, -3.81cm] | statistically significant (p < 0.001) |
At large rain events - the ones that carry real flood risk - the confidence interval sits entirely below zero: a statistically significant reduction in river rise, not a borderline or ambiguous result. At typical events, by contrast, the confidence interval spans both a small improvement and a small worsening - "not statistically significant" there means the data cannot currently distinguish a small effect from none at all, out of 111 post-NFM events in total. See Discussion for what can and can't be inferred about why the effect differs by event size.
How much data is behind each end of the size range: events of 27.6mm or more (the large-rain-event threshold above) are a small minority throughout - 7 of 167 before events and 13 of 111 after events. Events below that threshold make up the rest - 160 before, 98 after. The large-rain-event estimate above uses every event's rain total via the regression, not just these 7/13 events on their own, but this is the plainest way to see how thin that end of the range is.
"Large" means large rain, not necessarily a high river: a large rain event is defined by how much rain fell, not by how much the gauge rose - deliberately, since the river's own response is exactly what might be different because of the NFM works, and stratifying by an outcome that the thing being tested could itself change would risk biasing the comparison. Rain total is fixed by the weather, not by the catchment, so splitting on it doesn't have that problem. But the two aren't the same thing: of the 13 large-rain-events after the NFM works, only 6 also had a top-10% rise - some large-rain events produced only a small rise (typically when the ground was dry enough to absorb most of it), and some smaller-rain events produced a big one. Read the large-rain-event result as being about how much less the river rises for a given large amount of rain, not as a direct claim about flood peaks specifically.
For completeness and comparability with a simpler before/after test, the same pooled regression also reports a single average effect across the whole range of post-scheme event sizes (equivalent to the effect at a rain total of 0mm, an extrapolation no real event matches exactly): statistically significant +3.19cm (95% CI [+0.37cm, +6.02cm], p = 0.027, n = 278 pooled before+after events). This figure is dominated by the much larger number of small-to-typical events in the sample, so it should not be read as "the" answer to whether large, flood-relevant rain events have been affected - 3.2 above answers that question directly.
Model: rise ~ post_nfm + post_nfm:sum_rain_mm + length_hrs + sum_rain_mm + max_rate_mm_15min + preceding_7d_rain_mm + C(month),
ordinary least squares. post_nfm is 1 for events recorded after both
schemes were complete, 0 for events before Elmwood began.
Same model, applied to post-both-schemes events individually rather than pooled into one regression - useful as a sanity check on 3.2, but doesn't itself give a significance test of the combined effect: statistically significant (paired t-test, p < 0.001, Wilcoxon signed-rank p < 0.001) - on average, the river is rising less than the pre-scheme model predicts.
This analysis finds a statistically significant decrease in how much the Bucklebury gauge rises since the NFM works were completed, at the large rain events that carry real flood risk (-7.07cm at a large rain event, 95% CI [-10.33cm, -3.81cm], p < 0.001). At small-to-typical rain events - which do not carry real flood risk - the estimated effect is not statistically significant (-1.08cm, 95% CI [-2.77cm, 0.60cm], p = 0.207): the data cannot currently distinguish this from no change at all. It may not be possible to attribute either result to the NFM works alone - other catchment changes over the same period (Section 1) may be contributing, and this method cannot separate their individual effects; nor does the "before" baseline used here include an event on the scale of the largest on record, 19 July 2007 (see Discussion). What can be said with confidence is that the combined set of changes in the catchment since the NFM works began has produced a statistically significant reduction in river rise at the event sizes the works were intended to help with most, with no evidence either way at smaller, lower-risk events.
Every number on this page is generated automatically, not computed or transcribed by hand, and can be regenerated from source at any time:
See also the NFM effectiveness overview for shared methodology and limitations, and the Tidmarsh / Bourne page for the other hypothesis tested with the same method.