Count model comparisons
Count model alternatives and tests
Exploratory alternatives to the published count model. Historical identifiers below refer to the benchmark (M2), an annual random intercept (M3a), and a recent annual baseline with live updates (M3b). The benchmark here extrapolates its year spline; the published model's held-year rule is assessed separately in its description and validation page.
In these retrospective tests, M2 extrapolates its historical smooth-year trend. The live forecast holds that term at 2023 for later seasons. M3a adds a random annual deviation but has no information about an unseen year's direction. M3b shares M2's fitted weather response, removes the year smooth from the forecast, uses recent adjusted annual levels for the new-year baseline, and updates that baseline as counts arrive.
In 2014–2023 forward tests, M3b's one-year deviance was 504.0 versus 522.7 for M2. At two and three years ahead it was 518.4 versus 509.7 and 520.3 versus 518.9. These historical ERA5 results are exploratory; none assesses issued-weather errors.
M2 + prior-morning rain · Rolling next-year tests
Every tested weather input can be requested from the Open-Meteo ECMWF feed or derived from its hourly values. M2 uses eight daily weather inputs and has no cloud-base term. Only 2014–2023 seasons are evaluated here; earlier seasons can still enter each model's training set. Every test season is fitted only on earlier years and all candidates predict the same operated dates using historical ERA5 weather.
| Candidate | Dates | Deviance ↓ | Mean absolute error ↓ | Predicted mean |
|---|---|---|---|---|
| M2 | 142 | 522.72 | 483.09 | 810.28 |
| M2 + prior-morning rain | 142 | 496.44 | 463.57 | 740.70 |
On recent dates, M2 has deviance 522.7 and mean absolute error 483 birds per date. Adding prior-morning rain gives 496.4 and 464. These are paired comparisons on the same held-out dates.
| Candidate | Test period | Seasons improved | Deviance difference | 95% lower | 95% upper |
|---|---|---|---|---|---|
| Prior-morning rain | 2014–2023 | 6 | -26.28 | -64.60 | 12.76 |
Each bar is prior-morning rain minus M2 for one fully held-out season. Negative bars favor rain. The paired season-bootstrap interval conditions on this already explored candidate.
Earlier cloud, overnight and hourly CNN screens are summarized in the contender notes .
Operational audit: a live query for Ngulia returned temperature, humidity, precipitation, pressure, cloud and wind, including yesterday when requested with past_days=1. The API accepted a cloud_base field name but returned only null values, so cloud base was excluded from all active count models. See ECMWF API documentation .
Historical issued forecasts are for operational evaluation, not required to train these models: the fits above use ERA5 and observed counts. Open-Meteo's archived ECMWF individual runs start in March 2024, after the latest count labels in this dataset, so they cannot reconstruct the 2014–2023 forecasts as issued. The scheduled website workflow archives its selected M2 forecast; later count observations are still needed to score issued-weather accuracy. See Single Runs API availability .
Decision: the full model, with its year term held at the last training season, remains the published count model. Prior-morning rain improves the retrospective average, but its season-level interval includes no gain. Revisit it after obtaining new observed counts and comparable issued forecasts. Older years remain useful for fitting, but are not part of the weather-model choice score here.
M3a · Smooth year plus annual random intercept
M3a retains M2's weather, layout and smooth year trend, then adds a partially pooled annual intercept. The fitted annual log-scale SD is 0.38. That intercept may combine effort, population variation and other annual changes.
Annual deviations relative to M3a's smooth trend. Pointwise intervals exclude variance-component uncertainty.
Weather coefficients in M2 and M3a
Linear coefficients are conditional log catch associations per unit. M3a's annual intercept does not make every weather effect exclusively within-year.
| Model | Weather input | Log multiplier |
|---|---|---|
| M2 | total_cloud_cover_mean | 0.71 |
| M2 | relative_humidity_mean_pct | 0.07 |
| M2 | wind_u_10m_mean_ms | -0.60 |
| M2 | wind_v_10m_mean_ms | -0.02 |
| M3a | total_cloud_cover_mean | 0.62 |
| M3a | relative_humidity_mean_pct | 0.11 |
| M3a | wind_u_10m_mean_ms | -0.60 |
| M3a | wind_v_10m_mean_ms | -0.02 |
M3a annual totals · unseen versus known years
Solid model line: whole seasons held out, so the annual intercept is unknown. Dotted line: chronological date blocks held out while other dates from the same year estimate its intercept. Both are evaluated on the same included dates; the known-year test is retrospective.
M3b · Shared weather response and recent annual baseline
- Fit M2 to historical dates so its calendar, moon and weather terms use the long record. The historical smooth year term helps prevent annual level changes from being assigned to weather.
- For each historical season, remove the fitted year-smooth term and compare observed totals with the resulting weather/calendar expectations. This estimates an adjusted annual catch level.
- Use the most recent 10 annual levels, giving their weights a five-year half-life. This defines a constant pre-season level for future years rather than extending the spline's slope. The weights apply only to the annual baseline, not to the weather fit.
- Treat remaining annual variation as a Gaussian distribution on log level. New observed counts update that distribution through the negative-binomial likelihood; the weather response remains fixed.
The year smooth is therefore still used when estimating weather effects from historical data. It is not a future population-trend forecast in M3b. A stable recent annual baseline is an explicit assumption, and the smooth alone cannot prove that weather associations come only from within-year variation.
Marker size shows recency weight. The dashed baseline is the weighted expected annual multiplier 0.69 from 2014–2023; the uncertainty SD on log scale is 0.29. This does not assert that future effort or population size will be constant.
M3b annual totals · one-year-ahead forecast
Each season 1997–2023 is forecast using only earlier seasons. Annual totals sum predictions over the same observed operated dates, using retrospective ERA5 weather.
Validation · Different questions need different splits
Whole-season folds test transfer to seasons unseen by training. Within-year blocks test prediction when some dates from that year are already known. Rolling origins train only through the previous season and test one, two or three years ahead; they are the relevant check for 2025+ use.
Whole-season folds · M0 through M3a
| Model | Deviance ↓ | Annual mean log error ↓ | Centered daily log error ↓ | Predictive log loss ↓ | 80% coverage |
|---|---|---|---|---|---|
| M0 | 516.97 | 0.53 | 1.56 | 7.05 | 77.4% |
| M1 | 451.15 | 0.48 | 1.42 | 6.99 | 79.2% |
| M2 | 430.32 | 0.42 | 1.42 | 6.99 | 77.1% |
| M3a | 442.18 | 0.45 | 1.41 | 6.99 | 76.9% |
| Compared with M2 | Deviance difference | 95% lower | 95% upper |
|---|---|---|---|
| M3a | 11.86 | -0.00 | 26.31 |
Paired season-bootstrap intervals condition on the already explored models. Negative differences favor the candidate.
Known-year date blocks · M2 and M3a
| Model | Deviance ↓ | Within-season rank ↑ | Annual mean log error ↓ | Centered daily log error ↓ |
|---|---|---|---|---|
| M2 | 429.51 | 0.46 | 0.37 | 1.44 |
| M3a | 420.81 | 0.42 | 0.20 | 1.47 |
M3a improves the known-year level but slightly worsens centered daily error and rank. The date-block fit can use later dates from that season, so it does not simulate live updates.
Rolling forecasts · M2 versus M3b
| Model | Years ahead | Test period | Dates | Deviance ↓ | Annual mean log error ↓ | Predicted mean | Observed mean | Predictive log loss ↓ | 80% coverage |
|---|---|---|---|---|---|---|---|---|---|
| M2 | 1 | 1997–2013 | 421 | 562.44 | 0.38 | 708.31 | 694.44 | 7.35 | 78.4% |
| M2 | 1 | 2014–2023 | 142 | 522.72 | 0.46 | 810.28 | 614.43 | 7.35 | 85.2% |
| M2 | 2 | 1997–2013 | 391 | 571.32 | 0.39 | 715.22 | 703.84 | 7.37 | 79.8% |
| M2 | 2 | 2014–2023 | 142 | 509.70 | 0.45 | 832.69 | 614.43 | 7.34 | 87.3% |
| M2 | 3 | 1997–2013 | 361 | 587.25 | 0.39 | 728.41 | 706.73 | 7.38 | 78.4% |
| M2 | 3 | 2014–2023 | 142 | 518.94 | 0.46 | 865.89 | 614.43 | 7.34 | 86.6% |
| M3b | 1 | 1997–2013 | 421 | 560.07 | 0.37 | 715.26 | 694.44 | 7.35 | 79.3% |
| M3b | 1 | 2014–2023 | 142 | 504.01 | 0.44 | 799.22 | 614.43 | 7.35 | 88.0% |
| M3b | 2 | 1997–2013 | 391 | 566.23 | 0.39 | 719.82 | 703.84 | 7.37 | 80.3% |
| M3b | 2 | 2014–2023 | 142 | 518.37 | 0.46 | 828.21 | 614.43 | 7.35 | 87.3% |
| M3b | 3 | 1997–2013 | 361 | 580.01 | 0.38 | 728.24 | 706.73 | 7.38 | 79.8% |
| M3b | 3 | 2014–2023 | 142 | 520.26 | 0.47 | 859.06 | 614.43 | 7.35 | 88.0% |
| Years ahead | Test period | M3b − M2 deviance | 95% lower | 95% upper |
|---|---|---|---|---|
| 1 | 1997–2013 | -2.37 | -7.74 | 2.75 |
| 1 | 2014–2023 | -18.71 | -41.76 | 4.04 |
| 2 | 1997–2013 | -5.09 | -13.25 | 2.95 |
| 2 | 2014–2023 | 8.67 | -21.31 | 38.51 |
| 3 | 1997–2013 | -7.24 | -19.12 | 3.98 |
| 3 | 2014–2023 | 1.32 | -34.25 | 28.24 |
Every paired 95% season-bootstrap interval spans zero. The M3b year-level rule was proposed after examining historical count results, so these estimates are exploratory. All tests use ERA5 weather rather than archived issued forecasts.
Live annual updates · One new count at a time
Before the season, M3b forecasts with its recent annual baseline. For each completed operated date, enter the comparable non-swallow count and matched weather. Its negative-binomial likelihood updates only the current annual level; no whole-model refit is needed. Predict the remaining dates with the posterior average annual multiplier. Missing dates are never entered as zero.
| Model | Known dates | Remaining test dates | Seasons | Deviance ↓ | Annual mean log error ↓ | Predicted mean | Observed mean | 80% coverage |
|---|---|---|---|---|---|---|---|---|
| M2 | 0 | 563 | 27 | 552.42 | 0.41 | 734.03 | 674.26 | 80.1% |
| M2 | 5 | 428 | 27 | 525.14 | 0.51 | 748.15 | 727.45 | 82.0% |
| M2 | 10 | 293 | 26 | 512.46 | 0.58 | 674.89 | 665.07 | 81.9% |
| M3b prior | 0 | 563 | 27 | 545.93 | 0.40 | 736.44 | 674.26 | 81.5% |
| M3b prior | 5 | 428 | 27 | 518.29 | 0.51 | 751.43 | 727.45 | 83.4% |
| M3b prior | 10 | 293 | 26 | 507.65 | 0.57 | 680.88 | 665.07 | 83.3% |
| M3b updated | 0 | 563 | 27 | 545.93 | 0.40 | 736.44 | 674.26 | 81.5% |
| M3b updated | 5 | 428 | 27 | 523.60 | 0.50 | 679.43 | 727.45 | 80.8% |
| M3b updated | 10 | 293 | 26 | 488.24 | 0.51 | 631.45 | 665.07 | 82.3% |
All three lines are evaluated on identical subsequent dates at each known-date count. Moving right changes the test cohort; compare lines vertically, not the slope of one line alone. The full day-by-day trajectory is saved in count_m3b_updated_scores.csv.
The M2 specification is published for count forecasts with front-bush layout. This comparison remains retrospective: it uses historically selected operated dates, incomplete historical zero coverage and reanalysis weather. The live forecast holds M2's year term at its 2023 value for later seasons; the held-year rule is now assessed in the separate model description and validation report. Issued-weather accuracy remains unmeasured.
Rebuild: run scripts under count/scripts/alternatives/: 08_compare_m0_m2.R, 10_verify_count_annual.R, 12_compare_m3b.R, 14_verify_count_m3b.R, 15_compare_count_weather.R, 17_verify_count_weather.R, and count/scripts/build_alternatives_report.R with Rscript. Keep lib/ beside this HTML file.