Satellite-gauge-radar precipitation merging for hydrological forcing
No single precipitation source captures rain accurately across terrain, climate zones and storm types. Merging GPM IMERG, ground radar and gauge networks with quantified uncertainty is now the operational standard for hydrological forcing.
Sensors
- GPM IMERG (NASA/JAXA): Global passive-microwave precipitation estimate at 0.1° (~10 km) every 30 minutes. Final-run latency is roughly 3.5 months after gauge adjustment; Early and Late runs available within hours but without gauge correction. Systematically underestimates orographic enhancement and solid precipitation poleward of about 60°.
- TRMM 3B42 (historical, NASA/JAXA): Predecessor to IMERG, 0.25° resolution, 3-hourly, covering 1998–2019. Still the backbone of long climatological bias-correction baselines in the tropics and subtropics.
- Meteosat SEVIRI (EUMETSAT): Geostationary imager over Europe, Africa and the Indian Ocean. Rapid-Scan mode delivers imagery every 5 minutes at roughly 3 km in the visible and thermal infrared. Cloud-top temperature is used as a rain-rate proxy in cloud-motion-vector and convective-initiation algorithms, filling gaps between microwave overpasses.
- GOES-R ABI (NOAA): Western Hemisphere geostationary imager, 16 spectral bands, full-disc every 10 minutes and mesoscale sectors every 30–60 seconds. ABI-derived rain rates feed into NOAA's Multi-Sensor Precipitation Estimator and provide the high-temporal-frequency anchor for merging over the Americas.
- National and regional ground-radar composites: C-band and S-band networks (e.g., NEXRAD in the US, OPERA in Europe) provide 1–4 km spatial detail at 5–15 minute intervals over their coverage footprints, typically within 200–250 km of each antenna. Coverage drops sharply over oceans, complex terrain and most of the developing world.
Why every source fails alone
Rain gauges measure a point. A standard 200 mm orifice samples roughly 0.03 m² of the Earth's surface. In complex terrain, gauge density rarely exceeds one per 200 km², which means a single gauge is asked to represent rainfall fields that vary by a factor of three or more across a single ridge. Gauge undercatch in wind-driven rain or snow can reach 20–50% without correction.
Ground radar converts reflectivity to rain rate via the Z-R relationship, which is empirical and highly sensitive to drop-size distribution. Beam overshooting in mountainous terrain, partial beam blockage and anomalous propagation all introduce structured errors that are hard to separate from real precipitation signals. Radar coverage over oceans, the tropics and much of the developing world is sparse or absent.
Satellite passive-microwave retrievals, the core of GPM IMERG, infer precipitation from the scattering and emission signatures of hydrometeors at 10–89 GHz. Over land, the algorithm relies primarily on scattering at high frequencies, which responds better to ice aloft than to warm rain processes. The documented consequence is a systematic underestimate of orographic precipitation, where rain forms and falls quickly at low altitudes, and of snowfall at high latitudes, where the emission signal from a snow surface masks the precipitation signal. Published validation studies in the Himalayas, Andes and Scandinavian highlands routinely report IMERG biases of 20–60% against dense gauge networks.
The merging problem: statistics in service of physics
Merging is not averaging. The goal is to produce a gridded precipitation field that is unbiased in the mean, has realistic spatial structure, and carries honest pixel-level uncertainty estimates that can propagate into downstream hydrological models.
Quantile mapping is the most widely applied bias-correction method. It aligns the cumulative distribution function of the satellite product to that of a reference, typically a dense gauge network or a reanalysis, separately for each grid cell and season. It corrects systematic biases in the frequency distribution of rain rates but cannot redistribute rain spatially, so it does not fix the orographic underestimate in cells where the reference network is itself sparse.
Bayesian model combination weights multiple precipitation products according to their posterior error covariance, estimated from withheld gauge observations. The approach is principled: weights vary by location, season and rain-rate regime, and the output includes a full posterior distribution rather than a single best-guess field. The Ensemble Kalman Filter and its variants extend this to sequential data assimilation, updating the precipitation field as new radar or gauge observations arrive during a storm.
Machine-learning approaches, including random forests and convolutional neural networks trained on terrain, atmospheric reanalysis fields and historical gauge records, have shown skill in correcting the orographic bias specifically. They learn the statistical relationship between topographic exposure, wind direction and the ratio of gauge-observed to IMERG-estimated rainfall, then apply that relationship in real time. The honest caveat is that these models generalise poorly to storm regimes outside their training distribution, which matters most for extreme events, exactly the ones that drive flood risk.
Forcing a distributed hydrological model: where the errors go
A distributed model, whether a grid-based land-surface scheme or a semi-distributed catchment model such as VIC, mHM or SWAT, partitions precipitation into runoff, infiltration and evapotranspiration at each grid cell or sub-catchment. Precipitation bias enters the water balance directly: a 30% underestimate of mean annual rainfall produces a roughly proportional underestimate of mean annual runoff, with non-linear amplification during peak events because the soil moisture state at storm onset determines how much rainfall becomes fast runoff.
Uncertainty quantification matters here in a practical sense. If the merged product provides an ensemble of precipitation fields rather than a single deterministic field, the hydrological model can be run as an ensemble, producing a probability distribution over discharge at the forecast point. That distribution is what a flood manager actually needs to make a decision about gate operations or evacuation orders. Providing a single deterministic precipitation field and then asking a hydrologist to guess at the uncertainty is not a workflow, it is a guess dressed as a workflow.
Latency is a real constraint. GPM IMERG Early run is available within 4–6 hours of observation time, which is adequate for medium-range flood forecasting but not for flash-flood warning in small catchments with response times under 3 hours. In those basins, geostationary infrared estimates from SEVIRI or ABI, despite their lower accuracy, are the only satellite-based option, and they must be merged with any available radar data in near real time.
High-latitude and orographic regimes: the hardest cases
The GPM Core Observatory covers latitudes up to 65°N/S. Beyond that, passive-microwave coverage comes from the constellation members of the GPM network, with reduced revisit frequency. At high latitudes, frozen precipitation dominates for much of the year, and IMERG's solid-precipitation retrieval is acknowledged by the algorithm developers to be substantially less reliable than its liquid-precipitation estimates. Published comparisons against Norwegian and Finnish gauge networks show IMERG underestimating winter precipitation by 40–70% in some catchments.
In orographic regimes, the spatial resolution mismatch is as important as the algorithmic bias. A 10 km IMERG pixel placed across a ridge that receives 2,000 mm/year on the windward face and 600 mm/year on the leeward face will return something close to the spatial average, which is wrong for both sub-pixel catchments. Downscaling using a climatological precipitation-elevation relationship, or a physically based orographic precipitation model, is necessary before the merged product can force a hydrological model at sub-10 km resolution.
Operational products and their honest limits
Several global merged products are available openly. IMERG Final is gauge-adjusted and reprocessed but arrives too late for real-time operations. The CHIRPS product from USGS/UCSB combines infrared cold-cloud duration with gauge data at 0.05° resolution and has a daily latency of about 3 weeks for the final version, with a preliminary version available in roughly 2 days. GSMaP from JAXA offers a near-real-time product at 0.1°/hourly. Each has documented regional biases; none is universally superior.
The practical recommendation for an operational hydrological service is to run at least two merged products in parallel, use withheld gauge observations to track their relative performance by season and catchment, and weight them accordingly. This is not methodological indecision. It is the correct response to the fact that no single product dominates across all regimes.
Satellize applies these merging workflows operationally using open constellation data, with bias-correction models calibrated against available gauge networks. The approach is the same one used in the Tonga crop-estimation programme, where precipitation forcing quality directly determines the skill of the soil-moisture and yield estimates that the programme delivers.
What an honest merged product looks like
A well-specified merged precipitation product should state its spatial resolution explicitly, note that effective resolution (the scale at which it carries genuine information) is coarser than the grid spacing, quantify bias against an independent gauge set held out from the merging procedure, provide pixel-level uncertainty as a standard deviation or ensemble spread, and document performance separately for convective events, stratiform events, orographic regimes and solid precipitation. Products that do not provide these diagnostics are not more accurate; they are less honest.
For hydrological model forcing, the deliverable is not a map. It is a set of gridded time series at the model's spatial and temporal resolution, accompanied by an uncertainty ensemble that the model can consume directly. The difference between a precipitation product and a hydrological forcing product is that second step, which requires understanding both the sensor physics and the model's sensitivity to input error.
Typical figures
| Spatial resolution (IMERG) | 0.1° (~10 km at equator); effective resolution coarser due to sensor footprint averaging |
| Temporal resolution (IMERG) | 30 minutes (Early/Late/Final runs); Final run gauge-adjusted |
| Latency (IMERG Early run) | 4–6 hours after observation time |
| Latency (IMERG Final run) | ~3.5 months (reprocessed with gauge adjustment) |
| Geostationary temporal resolution (SEVIRI / ABI) | 5–15 minutes (SEVIRI standard/rapid-scan); 10 minutes full-disc for ABI |
| Ground radar spatial resolution | 1–4 km within ~200–250 km of antenna; coverage absent over oceans and most of developing world |
| GPM Core Observatory latitude coverage | 65°N to 65°S; higher latitudes covered by constellation members at reduced revisit |
| TRMM 3B42 archive depth | 1998–2019 at 0.25°, 3-hourly; used for climatological bias-correction baselines |
| Typical merged product uncertainty | ±20–60% in complex terrain; ±10–25% over flat, well-gauged regions (published validation ranges) |
| Delivery formats | NetCDF-4 gridded time series, GeoTIFF per timestep, ensemble members for probabilistic forcing |
Analytics Satellize can run
| Bias-corrected gridded precipitation (daily/sub-daily) | Quantile mapping of IMERG against available gauge network, stratified by season and elevation band | NetCDF forcing files at client model resolution, with correction factors documented per grid cell |
| Probabilistic precipitation ensemble | Bayesian model combination of IMERG, GSMaP and geostationary IR estimates, weighted by posterior error covariance from withheld gauges | 50-member ensemble precipitation grids for direct ingestion into ensemble hydrological models |
| Orographic downscaling layer | Climatological precipitation-elevation regression combined with terrain exposure index derived from a digital elevation model | Monthly correction surfaces at 1 km resolution, applied to coarse merged product before model forcing |
| Near-real-time merged precipitation for flash-flood basins | Kalman-filter fusion of SEVIRI/ABI infrared rain-rate estimates with available radar composites; gauge nudging where telemetry permits | Hourly gridded precipitation GeoTIFF updated every 30 minutes, with latency flag per pixel |
| Precipitation product performance monitoring | Continuous verification against withheld gauge observations: bias, RMSE, probability of detection and false-alarm ratio by season, elevation band and rain-rate class | Monthly performance dashboard report; automated alert when product bias exceeds agreed threshold |
| Historical reanalysis forcing for model calibration | TRMM 3B42 and IMERG Final merged and bias-corrected over the full 1998-to-present archive | Multi-decade gridded precipitation archive in client hydrological model format, with uncertainty time series |
Who does the work
We can get this done for you. Satellize runs its own analyst desk and a strong science team. You do not buy a data feed and work out what it means; our people source the imagery, run the analysis described on this page, and hand you the answer with its confidence limits stated. Discuss this requirement.