Mist model
Will there be mist?
The mist forecast follows the weather through the night to estimate the chance of mist being recorded at Ngulia. Here is what that number means, how the model learns, and how we have checked it.
The model estimates a 7-in-10 chance of any recorded mist under those weather conditions. Light or patchy mist counts too. The number does not tell us how thick the mist will be, how long it will last, or how many birds will be caught.
1. Learn the weather patterns of misty nights
The model learns from 1,184 dates with field observations across 43 seasons, from 1971 to 2013. Each observation says whether mist was recorded. About 62% of these dates had some mist; this describes the recorded sample, rather than every night at Ngulia.
For each ringing date, it reads twelve hourly weather records, from 21:00 the previous evening to 08:00 that morning, in Ngulia local time. The inputs are cloud cover, humidity, east–west and north–south wind, temperature, and the gap between temperature and dew point.
A small convolutional neural network (CNN) learns patterns across neighbouring hours and combines them into one probability. It retains the sequence of weather changes rather than reducing the whole night to averages. The forecast averages three trained networks to give its final estimate.
A little more about the model
The input has 12 hours and six weather channels. Dew-point depression is the temperature minus dew point, derived consistently from temperature and humidity. All inputs are standardised using training data.
Each network has two short convolution layers and 449 fitted parameters. Training treats any recorded mist as present, including light, sustained and unspecified mist. Nights without an observation or with incomplete weather are excluded. Bird counts, bush layout and calendar or moon terms do not enter this model.
The 21:00–08:00 window was retained because longer histories brought little benefit or performed worse. The full-data ensemble supplies the live forecast; the checks below use separate predictions for seasons left out of training.
2. Check it on seasons left out of training
A model can look convincing on the observations it learned from. To test it, we divide the seasons into five groups, train on four groups and predict the remaining group. Repeating this gives a test prediction for every included date. A season's observations never appear in both its training and test data.
We compare three approaches on exactly the same dates:
- Historical frequency: use the fraction of training dates with mist, with no weather input.
- Average weather: use a simple logistic model of the six weather inputs averaged over the night.
- Hourly model: use the weather sequence. This is the CNN used by the website.
The figure uses the Brier score: a measure of error in the predicted probabilities. Lower is better, and zero would be perfect. A confident forecast that turns out wrong receives a larger penalty.
The hourly model scores 0.1675, compared with 0.1845 for average weather: a 9.2% reduction in probability error. It performs better in all 5 season groups. This supports using it over the simpler comparison models, while leaving room for mistakes.
The season groups mix earlier and later years, so a model can learn from years later than its test season. This checks predictions for omitted seasons. A chronological test using only the past, like the count report, still needs to be done.
Detailed scores and validation method
| Approach | Brier ↓ | Log loss ↓ | AUC ↑ |
|---|---|---|---|
| Historical frequency | 0.2371 | 0.6671 | 0.4541 |
| Average weather | 0.1845 | 0.5515 | 0.7630 |
| Hourly model (used live) | 0.1675 | 0.5045 | 0.8103 |
Log loss is another measure of probability error, with a strong penalty for confident mistakes. AUC measures how well the model ranks mist dates above no-mist dates; 0.5 is chance ranking and 1 is perfect. AUC is not the percentage of forecasts that are correct.
The hourly-minus-average-weather Brier difference is -0.0169, with a paired season-bootstrap 95% interval of [-0.0249, -0.0098]. This resamples whole seasons to retain dependence between dates. It describes uncertainty in these saved comparisons, not uncertainty from the earlier model search.
The five season groups are fixed. Scaling, fitting and early stopping use training data only; test observations are excluded. Pooled scores give each date equal weight. The model was selected after exploratory comparisons on this historical record, so these scores are not an untouched final test.
The hourly CNN also allows more flexible relationships than the average-weather model. Its improvement cannot be attributed to hour ordering alone.
3. Do the probabilities match how often mist occurs?
If dates forecast around 70% have mist about seven times out of ten, that probability is useful. The figure below groups the hourly model's test forecasts by their predicted chance and compares them with how often mist was actually recorded.
Dots near the dashed diagonal indicate agreement. Below it, the forecast chance was too high; above it, it was too low. Hover over a dot to see how many dates support it.
The probabilities do not match perfectly. Small groups of dates can also give unstable comparisons. The plot is a check on average behaviour across similar forecasts, rather than a guarantee for any single night.
Light or patchy mist is harder to distinguish than sustained mist in these historical checks. Both still count as mist for the forecast. The model predicts what observers recorded, and the records do not provide a precise measurement of visibility or mist intensity.
Each dot summarises a 10-percentage-point forecast group. All forecasts shown are from seasons left out of fitting; no uncertainty bars are shown.
4. What remains to check in a new season?
The live forecast uses issued Open-Meteo ECMWF hourly weather, including the preceding evening, and the saved three-network model. It reports a chance of mist separately from the expected bird count.
The historical checks above use reconstructed ERA5 weather. They do not measure the accuracy of the issued weather forecast or how the model performs in today's conditions. The current dataset has no observed mist labels after 2013, so recent performance remains unmeasured.
For a new season, we need to record whether mist occurs and match each observation with the archived forecast as it was issued. That will let us check probability error and whether a forecast such as 70% still corresponds to mist on roughly seven out of ten comparable dates.
Reproduce the report and analysis
To rebuild this page from saved results, run from the repository root:
Rscript mist/scripts/03_build_report.R python3 scripts/publish_reports.py
To repeat data preparation, model training and validation, with the local ERA5 archive available:
python3 -m venv .venv-mist .venv-mist/bin/python -m pip install -r mist/requirements.txt MIST_PYTHON=.venv-mist/bin/python bash mist/scripts/run.sh
The saved configuration, season groups, inputs and run manifests retain the research audit. Model training and validation remain separate from the daily forecast refresh.
Sources: curated field observations, the Ngulia ERA5 hourly archive and fixed season groups. The live forecast uses the saved model weights; this page does not retrain the CNN.