Written to be checked rather than believed. Every figure carries the date it was measured and the endpoint that produced it, and where we cannot demonstrate something it is stated beside the things we can — not after them.
Compiled 14/09/2026 · measured, not estimated · prepared for the Bureau of Meteorology
Flood Sentinel is decision-support. It is not an official warning service and it does not replace one. Every forecast surface carries that wording, and the map carries it on the map itself:
Flood Sentinel provides decision-support forecasts. It is not an official warning service — always follow Bureau of Meteorology flood warnings and your state emergency service (SES) directions.
We do not render or word anything as a Bureau warning, and we do not borrow Bureau warning terminology to describe our own output. If you find anywhere in the product that reads otherwise, that is a defect and we would like to be told.
Three surfaces are public. Start there, because nothing we issue you can be checked against anything.
Two endpoints need no credentials at all. They are marked security: [] in the
OpenAPI document itself — that marking is the authority, not this sentence.
curl https://api.floodsentinel.com.au/bom/v1/health
curl https://api.floodsentinel.com.au/bom/v1/coverage
This is the part we would most like scrutinised, because the chart is drawn to be falsifiable rather than flattering.
Counted from the forecasts table in production — forecasts actually served, not registry entries — over a six-hour window on 14/09/2026. A cell counts as ML-backed only where the served method is a model rather than persistence or climatology.
| Lead time | Stations ML-backed | Elsewhere |
|---|---|---|
| 1 h | 13 of 13 | — |
| 3, 6, 12, 24 h | 12 of 13 | labelled climatology |
| 48 h | 9 of 13 | labelled climatology |
| 72 h | 8 of 13 | labelled climatology |
| 96, 120, 168 h | 0 of 13 | labelled climatology only |
A station without a gated model at a lead time is not left blank — it is labelled. We would rather show a labelled climatology curve than an unlabelled model-shaped one.
From GET /bom/v1/coverage on 14/09/2026. “Reporting” means fresh
within 7 days and not future-dated; the endpoint returns its own basis string
stating exactly that, so the threshold travels with the number.
| Stations held | Reporting, last 7 days | |
|---|---|---|
| River level | 5,249 | 2,681 |
| Rain gauges | 19,851 | 2,974 |
| Total | 23,676 | 4,429 |
Archive depth: earliest record 1832-01-01, current to the hour. Two cautions we would rather state than have you infer: the held totals are an archive, not a live network — the right denominator for any liveness claim is the last column — and we do not offer discharge or storage as live products.
A candidate replaces a production model only by passing a cutover gate on held-out 2020, 2021 and 2022 events: Gate A, CRPS at a two-day lead against the incumbent; and Gate B, threshold-conditional CRPS on above-major-flood rows only.
A standalone R² never justifies a cutover. We hold that line because we were burned by the alternative: a model graded “A” in our own registry logged an NSE of −53.8 — worse than persistence.
Conditioned on event windows — the rises that give an operator time to act — the served ensemble beats persistence broadly at mid and long leads.
Gate B has never passed. Across four model generations, including the currently-served promotes, no candidate has demonstrated skill over persistence conditioned on above-major-flood rows, and modelled peaks under-predict observed ones.
The size of that under-prediction was itself mis-stated internally until 26/08/2026. Read at the ensemble’s p90 rather than its median trajectory, on major-flood rows of the current generation, the worst-case peak miss is:
| Station | Worst-case miss (p90) |
|---|---|
| North Richmond | −19.7% |
| Windsor | −22.4% |
| Wallacia | −15.7% |
We publish this because a forecast you cannot calibrate is worse than one you can.
Bureau HCS/FWN live feeds and the WDO deep-history crawl, plus state agency sources (WaterNSW, Manly Hydraulics, VIC WMIS, QLD WMIP). Every reading carries its source, so provenance is a query rather than an assertion. Bureau attribution is rendered on the surfaces that use Bureau data.
Forcing is taken from a multi-model NWP ensemble — ECMWF IFS and AIFS (including the 51-member AIFS ensemble), GFS, ICON and UKMO — and every issued run is archived, so models train on what was actually forecast at the time rather than on hindsight. Inter-model spread feeds the published uncertainty band.
A harvester for the Bureau’s public Weather API (ACCESS-C3 hourly, ACCESS-G3 daily) exists in our codebase but was never scheduled and has produced no data. We found that on 14/09/2026 and are correcting it. Until it is running and measured we do not claim ACCESS as a source — disclosed here rather than left to be discovered by the agency whose API it is.
The pages carrying verification evidence — the performance scorecard and the event archive — currently sit behind sign-in, so they need an evaluation account. We are glad to issue one; tell us how many people need access. We would rather you checked the gate section against those pages than took this document’s word for any of it.
Figures measured 14/09/2026. Coverage and lead-time counts have a shelf life and can be re-run on request.