Have the Natural Flood Management (NFM) interventions upstream of Bucklebury in the Elmwood and Redhill Copse areas - the NFM works - and other upstream changes, made a difference to the river's response to rain? Last generated: 24 August 2026, 07:11

One of two pages testing our Natural Flood Management (NFM) work - see also the Tidmarsh / Bourne page

This page is written like a statistical analysis report: a plain-language summary first, for anyone who wants the finding without the statistics, then the full method and evidence behind it, laid out so it can be checked rather than just read. Every number on this page is generated automatically from real monitored rain and river gauge data - nothing here is hand-picked or adjusted after the fact.

In plain terms

For councillors, EA officers and residents who don't need the statistics - the full analysis follows below for anyone who wants to check it

This analysis finds two different things depending on the size of the rain event. For large rain events - the ones that actually matter for flood risk - the Bucklebury gauge has been rising significantly less since the NFM works were built. Our best estimate is a reduction of about 7.1cm at a large rain event, and this result is statistically significant - in plain terms, the data are strong enough that this is very unlikely to just be chance or noise in the measurements. Based on the data collected so far, the true reduction is most likely somewhere between about 3.8cm and 10.3cm - in other words, we can't pin down the exact size of the improvement, but we can be confident there is one. For small-to-typical rain events - which don't put anyone at risk of flooding - the picture is much less clear: our best estimate is a small reduction of about 1.1cm, but there isn't yet enough data to be confident that's a real effect rather than normal variation - the true figure could plausibly be anywhere from a moderate improvement to a small worsening, so we're not reporting this one as a finding either way.

It may not be possible to attribute this to the NFM works alone. Other changes have also been made in the valley upstream of Bucklebury since 31 May 2023, such as alterations to Bucklebury Common's tree and vegetation cover, upstream of Elmwood - additional leaky dams were built there with the intention of offsetting that reduced cover. Both this result and the "before" baseline it's measured against also only cover the period since the Bucklebury Flood Alleviation Channel was built in 2011 - see Methods for why.


Key points:

  • What we did: compared 167 rain events recorded before the Elmwood scheme began (restricted to after the 2011 flood alleviation channel - see Methods) against 111 recorded after both Elmwood and Redhill Copse were complete, using a statistical model of how the Bucklebury gauge normally responds to rain - and checked whether that response differs at typical vs large rain event sizes, since a single average across all sizes would hide exactly this kind of difference.
  • What we found: a statistically significant decrease in river rise at large, flood-relevant rain events, and no statistically significant effect either way at typical events (see above for both).
  • What it doesn't mean: "not statistically significant" at typical events does not mean the NFM works have made no difference there - it means the data cannot currently distinguish a small effect from none at all. It also doesn't mean the large-rain-event finding can be pinned on the NFM works alone - other catchment changes over the same period (above) could be contributing.
  • How sure are we, and how would you check it: the large-rain-event result is a statistically significant finding on real monitored data, not a model projection; the typical-event result is genuinely inconclusive, not a hidden null finding. Every figure below - the data, the method, the exact numbers - is laid out so it can be independently checked; see Data, code & reproducibility.

1. Background Back to top

The Natural Flood Management (NFM) interventions upstream of Bucklebury in the Elmwood and Redhill Copse areas - referred to throughout this page as the NFM works - comprise two schemes that sit on streams feeding into the River Pang upstream of the Bucklebury gauge: the Elmwood NFM scheme (24 leaky dams, log barriers and "dragon's teeth" on the Elmwood Stream, built 18 May - 25 September 2020), then the Redhill Copse scheme (13 leaky dams on Osgood's Gully, built 1-31 May 2023). Both sets of NFM interventions are on streams that flow into the River Pang upstream of the EA's Bucklebury gauge, so this page tests whether the NFM works, taken together, have changed how the river responds to rain - not either scheme in isolation, and not whether either scheme performed as designed in isolation.

PVFF's 2021 report (PVFF NFM Final Report - Statistical Analysis v1.0) first tested this question with a similar regression-based method, but had only 16 usable post-scheme rain events at the time and explicitly could not draw a conclusion either way. This page re-runs an updated version of that method (see Methods for what changed and why) against every qualifying rain event recorded since, generated automatically from the underlying rain and river gauge data rather than computed by hand.

Research question: have the NFM works - and whatever else has changed in the catchment over the same period - measurably changed how much the Bucklebury gauge rises in response to a given rain event, relative to before either scheme existed?

Note on independence: PVFF built and supports the NFM works. This page reports what the data show, without adjustment, regardless of whether the result matches what PVFF hoped they would achieve - see the result below.

SchemeStreamAffects gaugeBuilt
Elmwood NFM - 24 leaky dams, log barriers & "dragon's teeth" Elmwood Stream Bucklebury 18 May 2020 (began) / 25 Sep 2020 (complete)
Redhill Copse NFM - 13 leaky dams Osgood's Gully Bucklebury (same gauge as Elmwood) 1 May 2023 (began) / 31 May 2023 (complete)

Both schemes feed the Bucklebury gauge's catchment, so this page compares events from before Elmwood began (restricted to after the 2011 flood alleviation channel - see Methods 2.4) against events from after Redhill Copse was complete - the period reflecting both schemes together - skipping over the intermediate "Elmwood only" window entirely so it isn't muddled into either side of the comparison. The Bucklebury Waven Field scheme (2025) manages village surface water and isn't expected to affect the river gauge, so isn't tested here.

2. Data & methods Back to top

2.1 Rain events

A "rain event" is a spell of rain preceded and followed by at least 8 dry hours, with total rain below 5mm dropped as noise (no cap on event length). For each event we record its length, total rainfall, peak 15-minute intensity, and rain in the preceding 7 days (antecedent wetness) - the same rules as PVFF's 2021 report. Full detail: overview page.

How these compare to recognised practice: separating a continuous rain record into discrete events using a minimum dry-gap threshold is a standard technique in hydrology, usually called the "minimum inter-event time" (MIT/IETD) method - there is no single agreed threshold value in the literature (studies use anywhere from a few hours to over a day depending on climate and purpose), so an 8-hour gap is a defensible, stated choice within that range rather than an invented one (see Discussion for how sensitive the result is to that choice). Event duration and peak short-interval intensity are standard rainfall-event characteristics (peak intensity measures such as maximum 15/30-minute rate are the basis of established rainfall-erosivity metrics). Our 7-day antecedent rainfall sum is a simplified relative of the classical Antecedent Precipitation Index (API), which typically weights recent rain more heavily via a decay function rather than summing it flatly - a genuine simplification, not a claim of equivalence.

A frost covariate was tried and dropped: PVFF's 2021 report flagged several unexplained high outliers that followed cold snaps, so a "hard frost in the days before the event" flag was added and tested at several lookback windows. Every version failed for a different reason: a full week let clearly-thawed frost count (and extrapolated wildly onto a genuine multi-day cold snap in January 2024 that Bucklebury's "before" data had nothing comparable to - see Discussion), while shorter, more physically defensible windows left zero frost-affected events in Bucklebury's qualifying "before" sample at all, regardless of length - its few frost-adjacent events never reached the minimum gauge level to qualify. Cold-snap events remain a known, unmodelled source of residual variance in this analysis.

2.2 Bucklebury gauge overflow correction

Above a critical level of approximately 74cm, the river overflows upstream of the Bucklebury gauge and the raw reading under-represents the true level - this is why the original 2021 report excluded every event peaking above 70cm entirely. This analysis corrects for that with a linear adjustment above the critical level, calibrated against one independent point (the 2014 event's citizen-measured village depth against the simultaneous raw gauge reading). This lets exceptional events be included in the dataset generally, rather than silently dropped, though not every such event ends up used in every comparison on this page (see 2.4 below for one that doesn't) - see limitations for what this correction does and doesn't establish.

2.3 Exclusions

Events are dropped where the West Berkshire Groundwater Scheme (EA drought-relief pumping into the Pang above Bucklebury) was running at any point in the event's response window, since this raises the gauge reading independent of rainfall.

2.4 Before-period baseline

The Bucklebury Flood Alleviation Channel was built in 2011, changing how the catchment responds to rain independently of any NFM work. This was checked directly: fitting a single model across the whole pre-Elmwood record (2002-2020) with a pre/post-2011 indicator shows the rain-response slope is significantly steeper after the channel was built (interaction p = 0.009), and that same single model systematically under-predicts pre-2011 events and over-predicts post-2011 ones. Left uncorrected, the "before" baseline used on this page would partly reflect an unrelated hydraulic structure rather than only the catchment condition the NFM works were then built into. The "before" period used throughout this page is therefore restricted to events from 1 May 2011 onwards - after the channel was in place - so the only material difference between the "before" and "after" samples is the NFM works themselves. This is the same principle the existing Bucklebury flood-modelling work already applies when comparing pre- and post-channel behaviour.

This restriction removes the two largest events on record, 19 and 21 July 2007 (86.8mm and 6.4mm of rain, adjusted rises of 102.2cm and 41.6cm), from the "before" sample - both predate the channel by several years. No event of that scale has occurred since 2011 in either the "before" or "after" period, so the large-rain-event result reported below reflects the range of large rain events actually observed since 2011, not a 2007-scale extreme; see Discussion.

2.5 Statistical model

A multi-linear regression predicts the gauge rise from the event characteristics above plus month of year. The primary test pools every pre-Elmwood and post-Redhill-Copse event into one regression with a "post-NFM" indicator (and its interaction with total rainfall, so the effect isn't forced to be a flat shift), giving a direct significance test of the combined effect - rather than an informal count of events that came in higher or lower than a separately-fitted prediction.

2.6 Rain gauge selection

Three rain gauges feed this catchment (Yattendon, Thatcham, and a blend of the two). The gauge used is picked once, using only pre-Elmwood ("before") data, by 5-fold cross-validated R² (a fairer measure of real predictive skill than in-sample R², since it's tested on data the model wasn't fitted on) - not by whichever gauge happens to produce the most favourable "after" result. The gauge choice cannot have been influenced by the outcome being tested, since it's fixed before any post-scheme event is looked at.

If none of the three gauges reaches a cross-validated R² of 50% on the "before" data, the before-model isn't trustworthy enough to build an after-comparison on, and none is reported - rather than show a comparison against a prediction we don't trust. This same threshold, applied the same way, is why the Tidmarsh hypothesis reports no result at all: it is not a bar set to fit Bucklebury's data.

How good is the "before" model?

We tested this model against all 3 rain gauges (Yattendon, Thatcham, and a blend of the two) on pre-scheme events, and picked whichever fit best by cross-validated R² (tested on held-out events, a fairer measure of real predictive skill than plain in-sample R² - see below) - Yattendon was the best fit, and everything below uses only that gauge. R² is the share of event-to-event variation in the rise that the model explains from rainfall, event length, intensity, antecedent wetness and month (0% = no better than guessing the average, 100% = perfect).

Rain gaugePre-scheme eventsR² (in-sample)R² (cross-validated)
Yattendon (used)16775%68%
Thatcham16175%66%
Blended15476%65%

Cross-validated R² is normally a bit lower than plain, in-sample R² - that's expected: plain R² can look better than a model really is because it's measured on the same data used to fit it, whereas cross-validation tests each event using a model that never saw it. We use the more honest, cross-validated figure to pick the gauge, not the in-sample one.

3. Results Back to top

3.1 What matters most in predicting the rise

Fitted on the pre-Elmwood events, since that's the baseline the "after" period is compared against. The bars show each factor's standardized effect - scaled so they're directly comparable to each other regardless of their original units (mm of rain vs. hours vs. mm/15min) - so a longer bar means that factor moves the predicted rise more. "Sure" / "unsure" reflects whether that factor's effect is confident enough to rely on.

FactorRelative effectEffect sizeConfidence
Total rain (mm)
+0.83sure
Rain in preceding 7 days (mm)
+0.22sure
Peak rain rate (mm/15min)
-0.20sure
Length of event (hrs)
-0.13sure
Also in the model: month-by-month seasonal adjustment - click to show all 12 months

Every month is measured relative to January (the "baseline" row, fixed at zero by definition).

MonthRelative effectEffect sizeConfidence
Jan (baseline)
+0.00baseline
Feb
-0.30sure
Mar
-0.42sure
Apr
-0.68sure
May
-1.10sure
Jun
-0.96sure
Jul
-0.97sure
Aug
-1.31sure
Sep
-1.02unsure
Oct
+0.14unsure
Nov
+0.14unsure
Dec
-0.28unsure

3.2 Primary result: does the effect depend on event size?

The combined effect interacts significantly with rain amount (coefficient -0.3718, 95% CI [-0.5554, -0.1882], p < 0.001) - it is not a flat shift, so a single average figure across all event sizes would be misleading. The table below reports the estimated effect, its own 95% confidence interval, and its own significance test, separately at a typical event and at a large rain event - large rain events matter most here, since they are the ones with real flood risk; small events are broadly the ones that matter least.

Event sizeRain totalEstimated effect95% CIResult
Typical (median) 11.5mm -1.08cm [-2.77cm, 0.60cm] not statistically significant (p = 0.207)
Large rain event (90th percentile) - flood-relevant 27.6mm -7.07cm [-10.33cm, -3.81cm] statistically significant (p < 0.001)

At large rain events - the ones that carry real flood risk - the confidence interval sits entirely below zero: a statistically significant reduction in river rise, not a borderline or ambiguous result. At typical events, by contrast, the confidence interval spans both a small improvement and a small worsening - "not statistically significant" there means the data cannot currently distinguish a small effect from none at all, out of 111 post-NFM events in total. See Discussion for what can and can't be inferred about why the effect differs by event size.

How much data is behind each end of the size range: events of 27.6mm or more (the large-rain-event threshold above) are a small minority throughout - 7 of 167 before events and 13 of 111 after events. Events below that threshold make up the rest - 160 before, 98 after. The large-rain-event estimate above uses every event's rain total via the regression, not just these 7/13 events on their own, but this is the plainest way to see how thin that end of the range is.

"Large" means large rain, not necessarily a high river: a large rain event is defined by how much rain fell, not by how much the gauge rose - deliberately, since the river's own response is exactly what might be different because of the NFM works, and stratifying by an outcome that the thing being tested could itself change would risk biasing the comparison. Rain total is fixed by the weather, not by the catchment, so splitting on it doesn't have that problem. But the two aren't the same thing: of the 13 large-rain-events after the NFM works, only 6 also had a top-10% rise - some large-rain events produced only a small rise (typically when the ground was dry enough to absorb most of it), and some smaller-rain events produced a big one. Read the large-rain-event result as being about how much less the river rises for a given large amount of rain, not as a direct claim about flood peaks specifically.

3.3 The pooled model's flat-average estimate

For completeness and comparability with a simpler before/after test, the same pooled regression also reports a single average effect across the whole range of post-scheme event sizes (equivalent to the effect at a rain total of 0mm, an extrapolation no real event matches exactly): statistically significant +3.19cm (95% CI [+0.37cm, +6.02cm], p = 0.027, n = 278 pooled before+after events). This figure is dominated by the much larger number of small-to-typical events in the sample, so it should not be read as "the" answer to whether large, flood-relevant rain events have been affected - 3.2 above answers that question directly.

Model: rise ~ post_nfm + post_nfm:sum_rain_mm + length_hrs + sum_rain_mm + max_rate_mm_15min + preceding_7d_rain_mm + C(month), ordinary least squares. post_nfm is 1 for events recorded after both schemes were complete, 0 for events before Elmwood began.

111
post-both-schemes rain events tested
events recorded after Redhill Copse was complete
-3.41cm
average actual rise minus predicted rise
45 events higher than predicted, 66 lower

3.4 For reference: plain predicted-vs-actual comparison

Same model, applied to post-both-schemes events individually rather than pooled into one regression - useful as a sanity check on 3.2, but doesn't itself give a significance test of the combined effect: statistically significant (paired t-test, p < 0.001, Wilcoxon signed-rank p < 0.001) - on average, the river is rising less than the pre-scheme model predicts.

4. Discussion & limitations Back to top

  • Attribution: this is a before/after comparison of the one river that exists, using its own monitored record across time. Any other change in the catchment over the same ~5-year before/after span could contribute to the result shown here alongside the NFM works. A known example: alterations to Bucklebury Common's tree and vegetation cover, upstream of Elmwood, since 31 May 2023 - additional leaky dams were built there with the intention of offsetting that reduced cover. It may not be possible to attribute the result above to the NFM works alone - only to the combined set of changes in the catchment over that period.
  • A possible physical mechanism (not tested directly by this analysis): the interaction in 3.2 means the estimated benefit is concentrated at large rain events and not statistically distinguishable from zero at typical ones - not a flat "always better" shift. This is the opposite of the pattern reported for leaky dams specifically by van Leeuwen, Klaar, Smith et al. (2024) in the Journal of Hydrology, who found leaky dams' effectiveness at reducing peak flow diminishing as event size grew, once the storage behind them was exceeded. One possible explanation for the pattern found here is the reverse logic: at small events there may be too little flow engaged with the leaky dams and floodplain roughness features to produce a measurable difference either way, while larger flows give more opportunity for the structures to intercept, slow and re-route water across the floodplain before it reaches the gauge - but this analysis cannot establish which explanation, if either, is correct, and the two findings should be read as a genuine open question rather than a resolved mechanism. It's also worth noting that the "before" baseline (restricted to after the 2011 flood channel - see Methods 2.4) has not included an event on the scale of 19 July 2007, the largest on record; the large-rain-event result above reflects the range of large rain events actually observed since 2011, and a future event of that magnitude would be a genuine test of whether the pattern holds.
  • Statistical vs practical significance: the large-rain-event effect, which matters most for flood risk, is statistically significant (its 95% confidence interval excludes zero); the typical-event effect is not - its confidence interval is wide enough to include no effect at all. Neither result is the same as certainty about the cause - see attribution above.
  • Event-separation threshold: the 8-hour dry gap used to mark separate rain events (Methods 2.1) is a rain-only rule - it doesn't guarantee the river has fully settled from a previous event before the next one is treated as starting from a clean baseline. Checked directly: across all events, the median time for a rise to fall to just 10% remaining is 48 hours regardless of event size, so no gap threshold in the range that keeps events analytically distinct (a handful of hours to about a day) fully avoids this. Restricted to the events actually used in this comparison, it's a modest effect - about 30% show more than 10% of the previous rise still present at their own start, similarly in the before (32%, mean 3% excess) and after (30%, mean 6% excess) periods, so if anything this makes the improvement found harder to detect, not easier. Re-running the full methodology at gap thresholds from 6 to 24 hours (rather than only 8) shows the large-rain-event effect stays negative - an improvement - at every value tested, with point estimates from -1.4cm to -9.2cm; it clears the 5% significance threshold at 5 of the 7 values tried (6h, 8h, 10h, 18h, 24h) but not at 12h or 15h, where cross-validated gauge selection happens to pick a different rain gauge in a correspondingly smaller sample. The 8-hour figure itself was fixed before this reanalysis began, matching PVFF's 2021 report, not chosen because it produces this result.
  • Model assumptions: ordinary least squares assumes independent, similarly-variable residuals; rain-event responses may not strictly satisfy this (for example, variability in rise plausibly increases with event size), so the reported p-value and confidence interval should be read as indicative rather than exact. The interaction term partly addresses this by letting the effect vary with rain amount rather than forcing a flat shift.
  • Gauge overflow correction: the Bucklebury level correction above ~74cm rests on a single independent calibration point (the 2014 event), not a full rating-curve re-survey - treat it as the best currently-available correction, not a precisely-instrumented one. See PVFF's recorded modelling assumptions for the full parameter set.
  • Incomplete exclusion record: the West Berkshire Groundwater Scheme exclusion only covers one known pumping episode (24 Oct-14 Nov 2022) on record - it may not catch every pumping period across the full dataset.
  • Thatcham data quality: Thatcham's weather station feed has known problems (occasional obviously-wrong single readings, filtered out; uncertain historical time-conversion consistency) - see the overview page for detail. This didn't end up being the gauge used for this result (see Methods 2.6 and the gauge-fit table above).
  • Measurement basis: river levels are radar-gauge readings accurate to 1mm, but record water level, not flow volume - "rise in cm" is what this analysis works with throughout, the same basis the original 2021 report used.

5. Conclusion Back to top

This analysis finds a statistically significant decrease in how much the Bucklebury gauge rises since the NFM works were completed, at the large rain events that carry real flood risk (-7.07cm at a large rain event, 95% CI [-10.33cm, -3.81cm], p < 0.001). At small-to-typical rain events - which do not carry real flood risk - the estimated effect is not statistically significant (-1.08cm, 95% CI [-2.77cm, 0.60cm], p = 0.207): the data cannot currently distinguish this from no change at all. It may not be possible to attribute either result to the NFM works alone - other catchment changes over the same period (Section 1) may be contributing, and this method cannot separate their individual effects; nor does the "before" baseline used here include an event on the scale of the largest on record, 19 July 2007 (see Discussion). What can be said with confidence is that the combined set of changes in the catchment since the NFM works began has produced a statistically significant reduction in river rise at the event sizes the works were intended to help with most, with no evidence either way at smaller, lower-risk events.

Data, code & reproducibility Back to top

Every number on this page is generated automatically, not computed or transcribed by hand, and can be regenerated from source at any time:

  • Source data: real 15-minute river level and rain gauge readings from EA and PVFF telemetry, plus the Bucklebury overflow correction and the West Berkshire Groundwater Scheme record described above.
  • Event extraction and statistical analysis: run from source data every time this page is regenerated, with the full model coefficients and per-gauge fit table shown above.
  • This page's last run: 24 August 2026, 07:11. This page updates automatically from whatever data exists at that time - nothing here needs to be redone by hand.

See also the NFM effectiveness overview for shared methodology and limitations, and the Tidmarsh / Bourne page for the other hypothesis tested with the same method.