Count model comparisons

Count model alternatives and tests

Exploratory alternatives to the published count model. Historical identifiers below refer to the benchmark (M2), an annual random intercept (M3a), and a recent annual baseline with live updates (M3b). The benchmark here extrapolates its year spline; the published model's held-year rule is assessed separately in its description and validation page.

In these retrospective tests, M2 extrapolates its historical smooth-year trend. The live forecast holds that term at 2023 for later seasons. M3a adds a random annual deviation but has no information about an unseen year's direction. M3b shares M2's fitted weather response, removes the year smooth from the forecast, uses recent adjusted annual levels for the new-year baseline, and updates that baseline as counts arrive.

In 2014–2023 forward tests, M3b's one-year deviance was 504.0 versus 522.7 for M2. At two and three years ahead it was 518.4 versus 509.7 and 520.3 versus 518.9. These historical ERA5 results are exploratory; none assesses issued-weather errors.

M2 + prior-morning rain · Rolling next-year tests

Every tested weather input can be requested from the Open-Meteo ECMWF feed or derived from its hourly values. M2 uses eight daily weather inputs and has no cloud-base term. Only 2014–2023 seasons are evaluated here; earlier seasons can still enter each model's training set. Every test season is fitted only on earlier years and all candidates predict the same operated dates using historical ERA5 weather.

Candidate Dates Deviance ↓ Mean absolute error ↓ Predicted mean
M2 142 522.72 483.09 810.28
M2 + prior-morning rain 142 496.44 463.57 740.70

On recent dates, M2 has deviance 522.7 and mean absolute error 483 birds per date. Adding prior-morning rain gives 496.4 and 464. These are paired comparisons on the same held-out dates.

Candidate Test period Seasons improved Deviance difference 95% lower 95% upper
Prior-morning rain 2014–2023 6 -26.28 -64.60 12.76

Each bar is prior-morning rain minus M2 for one fully held-out season. Negative bars favor rain. The paired season-bootstrap interval conditions on this already explored candidate.

Earlier cloud, overnight and hourly CNN screens are summarized in the contender notes .

Operational audit: a live query for Ngulia returned temperature, humidity, precipitation, pressure, cloud and wind, including yesterday when requested with past_days=1. The API accepted a cloud_base field name but returned only null values, so cloud base was excluded from all active count models. See ECMWF API documentation .

Historical issued forecasts are for operational evaluation, not required to train these models: the fits above use ERA5 and observed counts. Open-Meteo's archived ECMWF individual runs start in March 2024, after the latest count labels in this dataset, so they cannot reconstruct the 2014–2023 forecasts as issued. The scheduled website workflow archives its selected M2 forecast; later count observations are still needed to score issued-weather accuracy. See Single Runs API availability .

Decision: the full model, with its year term held at the last training season, remains the published count model. Prior-morning rain improves the retrospective average, but its season-level interval includes no gain. Revisit it after obtaining new observed counts and comparable issued forecasts. Older years remain useful for fitting, but are not part of the weather-model choice score here.

M3a · Smooth year plus annual random intercept

η3a=η2+by,by∼N(0,σyear2)

M3a retains M2's weather, layout and smooth year trend, then adds a partially pooled annual intercept. The fitted annual log-scale SD is 0.38. That intercept may combine effort, population variation and other annual changes.

Annual deviations relative to M3a's smooth trend. Pointwise intervals exclude variance-component uncertainty.

Weather coefficients in M2 and M3a

Linear coefficients are conditional log catch associations per unit. M3a's annual intercept does not make every weather effect exclusively within-year.

Model Weather input Log multiplier
M2 total_cloud_cover_mean 0.71
M2 relative_humidity_mean_pct 0.07
M2 wind_u_10m_mean_ms -0.60
M2 wind_v_10m_mean_ms -0.02
M3a total_cloud_cover_mean 0.62
M3a relative_humidity_mean_pct 0.11
M3a wind_u_10m_mean_ms -0.60
M3a wind_v_10m_mean_ms -0.02

M3a annual totals · unseen versus known years

Solid model line: whole seasons held out, so the annual intercept is unknown. Dotted line: chronological date blocks held out while other dates from the same year estimate its intercept. Both are evaluated on the same included dates; the known-year test is retrospective.

M3b · Shared weather response and recent annual baseline

η3b=η2−g(year)+acurrent year

The year smooth is therefore still used when estimating weather effects from historical data. It is not a future population-trend forecast in M3b. A stable recent annual baseline is an explicit assumption, and the smooth alone cannot prove that weather associations come only from within-year variation.

Marker size shows recency weight. The dashed baseline is the weighted expected annual multiplier 0.69 from 2014–2023; the uncertainty SD on log scale is 0.29. This does not assert that future effort or population size will be constant.

M3b annual totals · one-year-ahead forecast

Each season 1997–2023 is forecast using only earlier seasons. Annual totals sum predictions over the same observed operated dates, using retrospective ERA5 weather.

Validation · Different questions need different splits

Whole-season folds test transfer to seasons unseen by training. Within-year blocks test prediction when some dates from that year are already known. Rolling origins train only through the previous season and test one, two or three years ahead; they are the relevant check for 2025+ use.

Whole-season folds · M0 through M3a

Model Deviance ↓ Annual mean log error ↓ Centered daily log error ↓ Predictive log loss ↓ 80% coverage
M0 516.97 0.53 1.56 7.05 77.4%
M1 451.15 0.48 1.42 6.99 79.2%
M2 430.32 0.42 1.42 6.99 77.1%
M3a 442.18 0.45 1.41 6.99 76.9%
Compared with M2 Deviance difference 95% lower 95% upper
M3a 11.86 -0.00 26.31

Paired season-bootstrap intervals condition on the already explored models. Negative differences favor the candidate.

Known-year date blocks · M2 and M3a

Model Deviance ↓ Within-season rank ↑ Annual mean log error ↓ Centered daily log error ↓
M2 429.51 0.46 0.37 1.44
M3a 420.81 0.42 0.20 1.47

M3a improves the known-year level but slightly worsens centered daily error and rank. The date-block fit can use later dates from that season, so it does not simulate live updates.

Rolling forecasts · M2 versus M3b

Model Years ahead Test period Dates Deviance ↓ Annual mean log error ↓ Predicted mean Observed mean Predictive log loss ↓ 80% coverage
M2 1 1997–2013 421 562.44 0.38 708.31 694.44 7.35 78.4%
M2 1 2014–2023 142 522.72 0.46 810.28 614.43 7.35 85.2%
M2 2 1997–2013 391 571.32 0.39 715.22 703.84 7.37 79.8%
M2 2 2014–2023 142 509.70 0.45 832.69 614.43 7.34 87.3%
M2 3 1997–2013 361 587.25 0.39 728.41 706.73 7.38 78.4%
M2 3 2014–2023 142 518.94 0.46 865.89 614.43 7.34 86.6%
M3b 1 1997–2013 421 560.07 0.37 715.26 694.44 7.35 79.3%
M3b 1 2014–2023 142 504.01 0.44 799.22 614.43 7.35 88.0%
M3b 2 1997–2013 391 566.23 0.39 719.82 703.84 7.37 80.3%
M3b 2 2014–2023 142 518.37 0.46 828.21 614.43 7.35 87.3%
M3b 3 1997–2013 361 580.01 0.38 728.24 706.73 7.38 79.8%
M3b 3 2014–2023 142 520.26 0.47 859.06 614.43 7.35 88.0%
Years ahead Test period M3b − M2 deviance 95% lower 95% upper
1 1997–2013 -2.37 -7.74 2.75
1 2014–2023 -18.71 -41.76 4.04
2 1997–2013 -5.09 -13.25 2.95
2 2014–2023 8.67 -21.31 38.51
3 1997–2013 -7.24 -19.12 3.98
3 2014–2023 1.32 -34.25 28.24

Every paired 95% season-bootstrap interval spans zero. The M3b year-level rule was proposed after examining historical count results, so these estimates are exploratory. All tests use ERA5 weather rather than archived issued forecasts.

Live annual updates · One new count at a time

Before the season, M3b forecasts with its recent annual baseline. For each completed operated date, enter the comparable non-swallow count and matched weather. Its negative-binomial likelihood updates only the current annual level; no whole-model refit is needed. Predict the remaining dates with the posterior average annual multiplier. Missing dates are never entered as zero.

Model Known dates Remaining test dates Seasons Deviance ↓ Annual mean log error ↓ Predicted mean Observed mean 80% coverage
M2 0 563 27 552.42 0.41 734.03 674.26 80.1%
M2 5 428 27 525.14 0.51 748.15 727.45 82.0%
M2 10 293 26 512.46 0.58 674.89 665.07 81.9%
M3b prior 0 563 27 545.93 0.40 736.44 674.26 81.5%
M3b prior 5 428 27 518.29 0.51 751.43 727.45 83.4%
M3b prior 10 293 26 507.65 0.57 680.88 665.07 83.3%
M3b updated 0 563 27 545.93 0.40 736.44 674.26 81.5%
M3b updated 5 428 27 523.60 0.50 679.43 727.45 80.8%
M3b updated 10 293 26 488.24 0.51 631.45 665.07 82.3%

All three lines are evaluated on identical subsequent dates at each known-date count. Moving right changes the test cohort; compare lines vertically, not the slope of one line alone. The full day-by-day trajectory is saved in count_m3b_updated_scores.csv.

The M2 specification is published for count forecasts with front-bush layout. This comparison remains retrospective: it uses historically selected operated dates, incomplete historical zero coverage and reanalysis weather. The live forecast holds M2's year term at its 2023 value for later seasons; the held-year rule is now assessed in the separate model description and validation report. Issued-weather accuracy remains unmeasured.

Rebuild: run scripts under count/scripts/alternatives/: 08_compare_m0_m2.R, 10_verify_count_annual.R, 12_compare_m3b.R, 14_verify_count_m3b.R, 15_compare_count_weather.R, 17_verify_count_weather.R, and count/scripts/build_alternatives_report.R with Rscript. Keep lib/ beside this HTML file.