Evidence · neighborhood-resolution downscaling · gate publication ds-2026.07-rf4

Where the downscaling model is validated to work, and where real stations can correct it

The API's neighborhood-resolution correction is a quantile regression forest that adjusts grid-scale air temperature toward per-polygon values, trained on weather-station observations and gated by held-out cross-validation in every Köppen climate zone it serves.

This page reports two facts about it. First, the zones where each serving band's correction is validated to beat the raw grid; the historical band, the near-term band, and all seven forecast lead days each carry their own independently published gate, so forecast and near-term data are corrected today, the same as the historical record. Second, the worldwide set of 3,977 stations reporting on the METAR surface-observation network (3,377 of them four-letter ICAO airport stations, the rest coastal and marine observing platforms) that can contribute a station blend or a direct in-polygon observation. The model corrects air temperature only; humidity, wind, and solar radiation remain at grid scale, so heat index, UTCI, and WBGT improve only through their temperature term.

19
Köppen zones under the base CV gate (tmax 19 pass, tmin 15)
8
band gates published: near-term + forecast days 1–7
3,977
METAR stations worldwide that can contribute an observation
~100–300 m
resolution where covariates are rich, toward ~1 km where coarse
Tier 1 · always wins
metar_station

A real weather-station observation inside the polygon itself. Overrides every modeled value outright.

Tier 2
era5_land_station_blended

The model correction refined by a distance-weighted kriging blend against real nearby weather stations, under its own gate.

Tier 3
era5_land_downscaled

The QRF model correction, applied only where the zone, target, and band pass their cross-validation gate.

Tier 4 · last resort
era5_land_grid

The raw ERA5-Land / Open-Meteo grid value, served whenever no validated correction applies. Always disclosed as such.

Every gate figure on this page (the map, the forecast-decay chart, the matrix) is the ds-2026.07-rf4 publication. The model serving today is ds-2026.07-rf5; its published gates carry the same base cross-validation result (19 of 19 zones on daytime max, 15 of 19 on nighttime min, the same four failures), and refreshing these figures to it is tracked on issue #582.

Four tiers, strictly ordered; a lower-ranked tier never overwrites a higher-ranked one, and every row carries a data_source stamp naming which tier produced it.

01 · Validation map

Pass and fail, zone by zone, with every real station drawn on top

Zone colors show where the selected band's correction is validated; ink dots are the 3,977 METAR stations that can ever contribute a Tier 2 blend or a Tier 1 override. A gray zone is not an error: it means the gate did not pass there, so those polygons serve the raw grid value, disclosed as such. Client project locations are deliberately not drawn here; deployment is a separate fact, reported in the table further down.

Band
Target
passes: correction beats the raw grid passes via flat per-zone bias adjustment only evaluated, gate not passed: serves grid value no published gate: serves grid value METAR station (3,970 of 3,977 drawn; none sampled or clustered)

Zones are drawn from the same Köppen–Geiger classification raster the pipeline assigns polygons against (Beck et al. 2018, the ~3 km raster bundled with kgcpy), aggregated to a 0.5° display grid for this page; cells that are mostly ocean are omitted, as are latitudes above 84°N and below 60°S (Antarctica's ice-cap zone carries no gate, and the 7 Antarctic research-station METARs below that edge are the only stations not drawn). The forecast view shows day 1 as the representative state; days 2–7 are in the chart and matrix below. In the near-term and forecast bands, a published gate lists only the zones that passed review, so the gray "evaluated, not passed" category appears only in the historical view, where the four tmin failures (Am, As, BWh, Cwa) are explicit.

02 · Held-out validation

Checked block by block against 756 of Seoul's own sensors

The downscaling code is public, and no model change ships without validation against held-out ground truth. Seoul's 756-sensor municipal network gave the first city-scale check of this kind: 42,537 sensor-days across 756 stations and 315 neighborhoods, scored on 63 neighborhoods the model never trained on.

Seoul held-out validation: RMSE against 756 real sensors Bar chart comparing daytime max-temperature RMSE against Seoul's S-DoT sensor network, held-out neighborhoods only, for the raw ERA5 weather grid, HeatReady's global model, and a model trained on Seoul's own sensors, on all days and on the hottest quarter of days. All days (n=65) Raw weather grid (ERA5) 3.57°C HeatReady global model 1.95°C Model trained on Seoul’s own sensors 1.13°C Hottest 25% of days (city mean high ≥ 34°C) Raw weather grid (ERA5) 4.37°C HeatReady global model 2.50°C Model trained on Seoul’s own sensors 1.26°C
Daytime max-temperature RMSE against Seoul's 756-sensor S-DoT municipal network, held-out neighborhoods the model never trained on, lower is better. The Seoul-trained model's error grows the least on the hottest days, when it matters most (research/seoul-local-sensor-validation/EDA.md §7, independently re-confirmed with a second day-selection method).

The code is at heat-ready-downscaling. A public registry and scoring path for outside contributions has been designed and approved, though building has not started; nothing contributed will reach served data without clearing the same gate this result did.

03 · One city at full depth

Boston and Cambridge, every tier represented, tract by tract

ERA5-Land's raw grid runs roughly 9 to 28 km: one temperature for a canopy-heavy block and an impervious-heavy block a few streets apart. A statistical model, cross-validated per climate zone, corrects each polygon toward its own terrain, canopy, built fraction, surface-temperature structure, and population. A real weather station, where one sits inside a polygon, overrides the model outright. Four tiers, strictly ordered: a worse tier never overwrites a better one.

One place at full depth: Boston & Cambridge, tract by tract

Every view below is the same 591 census tracts on real stored days: first what the raw ~9 km reanalysis grid can say, then the corrected field the API actually serves, then the correction itself, the data tier behind every value, and the tract-level covariates the model reads. Hover any tract for its full record.

The served field
Why it differs

04 · Forecast decay

Validated coverage narrows as lead time grows

Each forecast lead day is gated independently; a correction validated for tomorrow is never assumed to hold a week out. The solid lines count zones whose gate passes at each lead; the dashed lines count how many of those passes rest on per-polygon spatial skill rather than a flat per-zone bias adjustment.

Tmax: zones passing Tmax: of those, with spatial skill Tmin: zones passing Tmin: of those, with spatial skill

On day-ahead forecasts, 11 zones pass for tmax and 9 for tmin. By day 7, tmax narrows to 5 passing zones and tmin holds at 8. The kind of improvement also changes with lead time: at day 1, 7 of the 11 passing tmax zones beat the grid through per-polygon spatial signal; by day 7 only Am does, and the other four passes rest on the flat bias adjustment, a real, validated improvement of a different kind, disclosed per zone in the matrix below. The near-term band (not shown here; it has no lead axis) passes 10 zones on tmax and 5 on tmin, every one of them with spatial skill.

05 · The full gate matrix

Every zone, every band, every target, including the failures

One cell per zone, band, and target. Amber cells pass only through the flat per-zone bias adjustment; the distinction is disclosed rather than folded into a single "pass". Gray cells were evaluated and did not pass, so those polygon-days serve the raw grid value, the gate working as designed.

swipe to see all 20 columns →
passes, per-polygon spatial skill passes, bias adjustment only evaluated, gate not passed: serves grid no published gate for this band dashed outline: margin is thin, see note below

The ERA5 columns are the base cross-validation gate; the spatial-vs-bias distinction is published for the band gates only, so base-gate passes render solid. The BLEND columns are the Tier 2 station-blend gate (near-term band only today): 8 of the 22 evaluated zone labels pass on tmax and 8 on tmin. The blend gate additionally evaluated the tundra zone ET and two coarse latitude-band fallback labels (tropical, temperate); none of the three passed, and they are omitted from the rows above because no other gate covers them.

Cfa/tmin (dashed cells above): passes every band shown here, but a re-score against the live model (ds-2026.07-rf5) using this program's stricter debiased-CV methodology found it clears the 3% auto-enable margin by only +0.6%, and fails the raw, undebiased qrf_beats_grid check outright. The gate itself hasn't changed and still serves the correction, since nothing here has actually regressed below the published bar: this is a disclosure that the margin behind that bar is thinner than every other zone's, not a change to what's served. See the crowdsourced program's submission report that found this (by_target.tmin.Cfa in claimed_report.json).

06 · What the database holds

Every cell here is a real count against the live database

The matrix below counts, cell by cell, how many of each project's areas actually carry each of 14 data families, across eight live projects. The temporal strip and the decadal archive answer a different question: how far back each family's record actually runs. The reconciliation ribbon shows the same handoff, provisional to confirmed to forecast, that keeps every project's numbers current. Each of the 14 families traces to one named public source; the full source list is in the documentation's source reference.

Fourteen data families, eight cities

Every cell is a real count against the live database. Solid is essentially every area covered. A tint drops off from there, and a hollow cell means only a few areas have the value at all. A faded cell hasn't been extracted here yet. The small dot marks a family where most areas are physically smaller than the satellite pixel the underlying dataset was built on, so the value is present but coarser than it looks.

swipe to see all 14 columns →
full coverage (≥98%) partial (25–98%) sparse (<25%) not extracted here resolution-marginal: source pixel larger than half these areas

Heat, baseline, and forecast columns start from ERA5-Land and Open-Meteo's ~9–28 km reanalysis grids. A statistical model corrects each polygon's grid temperature toward its own neighborhood, using terrain, canopy, built fraction, surface-temperature structure, and population, validated separately per climate zone; where a real airport weather station sits inside a polygon, its observed reading overrides the model. The corrected resolution runs roughly 100–300 m where the covariates are richest, coarser elsewhere (see how the correction is validated). Mexico City's row is the most instructive: population, age, and nightlights all show full coverage but a resolution-marginal dot, because 1,182 colonias are individually smaller than several of the source rasters they're drawn from: the data is real, the dot says exactly why it's coarser than it looks.

COVERAGE OVER TIME, NOT JUST SPACE

The matrix above answers "how much"; this answers "since when"

A date range and a percentage are two different shapes of fact, so they get two segments per row instead of one misleading average: on the left, every individual Landsat overpass this city's persistent-heat composite draws on, spanning back years at a handful of dates a month. On the right, the daily heat pipeline's dense, unbroken recent window, with the trailing air-quality readings marked on top. The dashed break between them is real: "years" and "days" don't belong on one linear axis.

SATELLITE OVERPASSES DAILY PIPELINE · LAST ~30D + 6D FORECAST
swipe to see the daily pipeline →

Ember ticks: one satellite scene date. Ember bar: daily ERA5/Open-Meteo coverage. Haze mark: the air-quality table's first-to-last date on record. Lagos's satellite record is visibly the sparsest of the eight, a real limit on cloud cover and overpass frequency this product hasn't smoothed over.

How far back "1940–now" actually goes

MARICOPA COUNTY, AZ · ONE JUNE PER DECADE
101.9°F
1946
102.4°F
1956
106.1°F
1966
106.6°F
1976
108.4°F
1986
105.4°F
1996
105.0°F
2006
111.2°F
2016
109.5°F
2026

Back to 1946: any June since then, per polygon, through the same code that ran this morning. That's the archive's real reach: not a trend claim (nine single hottest-days a decade apart is too little to call a slope), just proof of depth.

Why HI/T2m only, and where these bars come from

Each bar is that June's single hottest day, by heat index, pulled directly from the CDS archive and run through the same computation code the live API uses, not sampled from the daily pipeline, which only holds a rolling recent window. Heat index and 2m air temperature only: these are ERA5-native variables with an unbroken historical record back to 1940. UTCI/WBGT are not shown here because they depend on wind and solar fields sourced from a separate near-real-time archive that doesn't extend this far back. Computed once; historical values don't change.

ERA5 final provisional (Open-Meteo), confirmed within ~6 days forecast (Open-Meteo, 7-day outlook)
ARIZONA
ANDALUSIA
BANGLADESH
OAXACA HIGHLANDS
LAGOS
June 14July 3 · todayJuly 5

Every project, every day, three sources: hollow provisional cells become solid when ERA5 confirms them, usually within 6 days, and violet forecast cells become provisional as each day arrives: the pipeline's reconciliation, visible.

07 · Deployment status

Ten public projects, tier by tier, over the last 30 days and the forecast week

The map above answers where the model is validated; this table answers what the ten public projects are actually served, counted over every polygon-day the pipeline currently holds: the last 30 days plus the days already written for the forecast week, 36 polygon-days per polygon in this pull and 37 for Bangladesh. These are different facts, and a validated zone does not by itself put any specific project above the raw grid tier on a given day.

ProjectKöppen zonePolygons GridDownscaledBlendedStation Above grid

Polygon-day counts by data_source tier over the window the fleet endpoint reports, read live from the downscaling-fleet-status endpoint on 2026-09-02 22:02 UTC (research/site-review-2026-09/fleet-status-2026-09-02.json). That window is every day from 30 days ago onward with no upper bound, so it includes the days already written for the forecast week: 36 polygon-days per polygon here, 37 for Bangladesh. "Above grid" is the share of each project's polygon-days served from any tier above the raw grid. A trailing window, not an all-time total: a project's older history was written under whichever gates were published at the time, so this reports what the fleet serves now. Census tracts on this site are clipped to land with US Census TIGER water boundaries; New York is served as two projects and shown as one city on the landing page.

165,556
polygon-days served across the ten projects in this window
93.9%
of those served above the raw grid tier
16,578
station-blended (Tier 2) rows, now across six projects
660
direct station-observation rows; 433 of them in the Phoenix project

Tier 1 rows are rare by construction; a station must fall inside a polygon's own boundary, which is common for county-scale polygons (Phoenix, 433 of the 660) and rare for neighborhood-scale ones. The station-blend tier began serving on 2026-07-22 in one project and now reaches six. Shares still differ across projects for structural reasons: each project's zone mix determines which gates apply to it, and its polygon geometry determines whether a station can ever fall inside. Two projects stand out and are worth naming rather than smoothing over: Bangladesh and Phoenix served no model-corrected rows at all in this window, even though both their zones pass the base gate. Whatever is holding those two back is not the gate, and it is being investigated (issue #589).

DELHI · 446 polygons
MANHATTAN & BROOKLYN · 1,114 polygons

Each tile's own color scale is relative to that city, the real neighborhood-to-neighborhood texture the correction adds within one city, not a shared cross-city danger scale. Boston & Cambridge has its own full-depth figure above.

Three more projects, tier by tier

Three projects that each show a different part of the ladder, over the same window:

🇺🇸

Boston & Cambridge

591 census tracts · Cfa/Dfa/Dfb climate zones

The project that proved the station-adjusted tier: 2,822 of its 21,276 polygon-days in this window are blended against nearby airport weather stations, on top of 18,104 model-corrected, 290 still on the raw grid, and 60 real in-polygon station readings from Logan and Hanscom.

🇮🇳

Delhi

446 polygons · BSh climate zone

Every one of its 16,056 polygon-days in this window is served above the raw grid: 14,036 model-corrected and 2,020 station-blended, none on the raw grid at all.

🇺🇸

Manhattan & Brooklyn

1,114 census tracts · Cfa climate zone · second-largest project on the API, after Mexico City's 1,182 colonias

The heaviest user of the station blend: 6,626 of its 40,104 polygon-days in this window are blended, 33,116 model-corrected, 30 carry a real in-polygon reading, and only 332 are still on the raw grid.

Not every zone clears validation. A project in an excluded zone still returns a usable, disclosed grid value instead of a wrong correction, per the validation map and gate matrix above.

08 · Precedence on the ground

Two real days in one project, and the whole ladder between them

The four-tier ladder at the top of this page is the rule; this map is the rule executing on real stored rows. Boston-Cambridge is the one public project where every tier has served so far, and its 591 census tracts are drawn here for two consecutive real days, each tract colored by the data_source stamp actually stored for that tract and day.

Day
Tier 1 · metar_station Tier 2 · era5_land_station_blended Tier 3 · era5_land_downscaled Tier 4 · era5_land_grid

These two days are shown because together they exhibit all four tiers, and no single day can: precedence is an UPSERT guard, so when the station blend cleared its gate it replaced the model correction on the identical 467 tracts, a better tier replacing a worse one, never the reverse. The 122 raw-grid tracts and the two station-override tracts are unchanged across both days. The overrides are the tracts with a reporting METAR station physically inside their boundary: Logan International (KBOS) in Census Tract 9813, Suffolk, and Hanscom Field (KBED) in Census Tract 3593.01, Middlesex. Hover any tract for its stored values; blended tracts also report how many stations contributed, the nearest one's distance, and the station weight actually applied. Geometry is the project's own published tract boundaries; tier assignments are verbatim from the same 2026-07-23 snapshot as the table above.

09 · Method

What the model is, and how it earns each pass

Model

Quantile regression forest on anomalies

The model predicts each polygon's temperature anomaly relative to its own grid cell, never the absolute temperature, from static covariates the API already holds: warm-season land-surface temperature, tree canopy, land-cover fractions, built-settlement fraction, population density, and terrain from a digital elevation model.

Validation

Leave-region-out cross-validation

Folds hold out whole countries or admin regions, never random rows, because random splits leak through spatially correlated neighboring station-days. Per zone and per target, the correction must beat the raw grid on held-out error; a zone that fails ships the grid value instead.

Uncertainty

Conformal intervals, calibrated per zone

Reported 95% intervals are split-conformal calibrated per Köppen zone on held-out stations, with empirical coverage checked against the 0.95 target during cross-validation rather than assumed from the forest's own quantiles.

Applicability

Out-of-distribution flagging

Polygons whose covariates sit outside the training distribution are flagged (out_of_distribution) in the response rather than silently corrected, and the interval widens accordingly.

Resolution

A floor, not a fixed number

Effective resolution is roughly 100 to 300 meters where the LST, canopy, and WorldCover covariates are present, degrading toward roughly 1 kilometer where only coarse covariates exist. The floor is a model-card value that moves with the covariates, never a flat promise.

Precedent

Assembled from established parts

Each component (QRF, anomaly regression, spatial cross-validation, conformal calibration) is independently established in the literature, but no published study validates this exact combination. The closest precedent, comparing random-forest and kriging approaches in a tropical setting, reported mixed results; no decisive advantage over alternatives is claimed here, only the per-zone gates above.

10 · Model drift and retraining

A standing policy, executed manually today

Retraining follows a calendar backstop plus threshold triggers. The calendar half: the model is re-fit at least annually, folding in a year of new station-days to catch slow concept drift. The threshold half: the deployed version is re-scored against freshly pulled held-out station-days whenever a new region onboards (exactly when out-of-applicability rates are most likely to spike) or a zone's held-out error degrades meaningfully against its original gate. A covariate refresh alone runs as a stateless re-join against the existing model, without triggering a retrain.

Execution today is a manual training build on a temporary EC2 instance, with each new gate reviewed before it is published; there is no live, automated continuous-training pipeline, and this page does not claim one. The current model version, ds-2026.07-rf5, is stamped on every corrected row it serves.

Current model version
ds-2026.07-rf5
Calendar backstop
At least annual, refit on a year of new station-days
Threshold triggers
New-region onboarding, or a zone's held-out error degrading against its gate
Covariate refresh alone
Stateless re-join, not a retrain trigger
Execution today
Manual training build, each gate reviewed before publishing
11 · Limitations

What this correction is not

  • LST is skin temperature, not air temperature. The satellite thermal covariate enters the model as a structural descriptor of where a neighborhood sits in its city's persistent heat pattern, never as a day-to-day predictor of air temperature.
  • Only air temperature is corrected. Humidity, wind, and solar radiation are served at grid scale, so heat index, UTCI, and WBGT improve only through their temperature term; their moisture, wind, and radiation terms carry grid-scale error unchanged.
  • Conformal coverage is calibrated per zone, with a known theoretical strain. Split-conformal calibration assumes exchangeability, which spatial autocorrelation undermines; per-zone calibration partially mitigates this by not pooling unlike regimes, a documented limitation, still open.
  • Tmin fails its gate in specific zones, by design. In the historical band, nighttime minimum temperature did not beat the grid in Am, As, BWh, or Cwa; polygons in those zones serve the grid tmin, disclosed as era5_land_grid, while tmax in the same zones is corrected. The near-term and forecast bands validate fewer zones still, as the matrix above shows.
  • Some passes are bias adjustments, not spatial corrections. An amber cell in the matrix beats the grid through a flat per-zone offset; it does not resolve neighborhood-to-neighborhood contrast in that zone and band.
  • Station coverage is uneven. The map's station layer is dense across North America and Europe and sparse across much of central Africa, the Amazon interior, and Central Asia; where no station exists, Tiers 1 and 2 can never apply and the model correction is the ceiling.

A failed gate serves the raw grid value, labeled as such. The alternative, serving an unvalidated correction everywhere, would look better on a map and be worse in every decision that relies on it.