Crop yield estimation using SAR and optical fusion
Combining Sentinel-1 SAR backscatter with Sentinel-2 vegetation indices lets analysts predict end-of-season crop yield at field or district scale, even through the cloud cover that makes optical-only approaches unreliable in humid growing regions.
Sensors
- Sentinel-1 A/B (C-band SAR): 10 m ground range resolution in Interferometric Wide Swath mode, 250 km swath, 6-day repeat at the equator (12-day per satellite). VV and VH polarisations respond to canopy volume scattering and biomass accumulation. Penetrates cloud and rain at C-band frequencies (5.405 GHz), making it the primary data source during monsoon or rainy-season growing windows.
- Sentinel-2 A/B (MSI): 10 m resolution in visible and near-infrared bands, 20 m in red-edge and shortwave infrared. Five-day combined revisit. Delivers NDVI, EVI, LAI and red-edge chlorophyll index. Cloud contamination is the binding constraint in humid regions; a full growing season may yield only three to five usable cloud-free composites in some tropical zones.
- MODIS Terra/Aqua: 250 m to 500 m resolution, near-daily global coverage. Useful for regional-scale LAI and NDVI time-series infilling where Sentinel-2 acquisitions are cloud-contaminated, and for historical yield-model calibration over multi-year archives dating to 2000. Spatial resolution limits its utility below administrative-unit scale.
- Planet SuperDove: 3 m resolution, up to daily revisit in tasked mode, eight spectral bands including red-edge. Adds within-field spatial detail for calibration and validation of coarser model outputs. Commercial licence required; not open-access.
What the radar backscatter signal is actually measuring
C-band SAR backscatter intensity is not a photograph of a crop. It is a measure of how much microwave energy the canopy scatters back toward the sensor, which depends on canopy geometry, moisture content, and the roughness of the underlying soil. In cereals, VH-polarisation backscatter rises through the vegetative phase as stem density and leaf area increase, peaks around heading, then drops as the canopy dries down toward maturity. That trajectory, integrated across a growing season, correlates with above-ground biomass accumulation and, by extension, with harvestable yield.
The physics matters for honest model design. C-band does not penetrate a dense closed canopy the way L-band (1.2 GHz, used by ALOS-2 PALSAR) does. Once a maize or rice canopy fully closes, C-band backscatter saturates and loses sensitivity to further biomass increases. Published saturation thresholds for C-band VH in rice are typically cited in the range of 2 to 4 kg per square metre of above-ground dry biomass, depending on crop architecture. Models that ignore saturation overestimate yield in high-productivity fields.
Why fusion outperforms either sensor alone
Optical vegetation indices such as NDVI and LAI track canopy greenness and photosynthetically active radiation interception, which are the inputs to most process-based crop growth models. The problem is that a single cloud-contaminated acquisition can corrupt an entire phenological stage in the time series. SAR fills those gaps, not by measuring the same thing as NDVI, but by providing a structurally complementary signal that is available regardless of cloud.
The fusion approach treats the two data streams as independent predictors in a statistical or machine-learning model. A random forest or gradient-boosted regressor trained on multi-year yield survey data can learn which combination of SAR backscatter trajectory features (peak VH, rate of rise, date of peak) and optical index features (maximum NDVI, integrated LAI) best predicts the observed yield. Studies published in Remote Sensing and similar journals consistently show that fused models reduce root-mean-square error by 15 to 30 percent compared with optical-only baselines in regions with more than 60 percent cloud cover during the growing season. The improvement is smaller in arid or semi-arid regions where optical data are rarely missing.
Process-based models such as ORYZA or DSSAT can also assimilate SAR-derived LAI estimates through data assimilation schemes, replacing the need for dense optical observations at the cost of greater computational complexity and the need for local soil and weather inputs.
Model architectures and their honest limits
Three broad model classes appear in operational deployments. Empirical regression models (linear, random forest, gradient boosting) are fast to train and interpret but require multi-year ground-truth yield data at the spatial unit of interest. They generalise poorly to new crop varieties or to years with anomalous weather outside the training distribution. Process-based models with satellite data assimilation are more physically grounded but require meteorological forcing data and soil parameter maps that are rarely available at field scale in data-sparse regions. Deep learning approaches, including LSTM networks applied to dense SAR and optical time series, show promise but demand large labelled datasets and are opaque to agronomists who need to explain a forecast to a ministry.
Uncertainty quantification is not optional in food security applications. A point estimate of 2.4 tonnes per hectare means little without a credible interval. Ensemble methods, conformal prediction wrappers, or Bayesian regression can all produce calibrated prediction intervals. The honest answer is that field-scale yield prediction uncertainty is typically plus or minus 15 to 25 percent at one standard deviation even in well-calibrated systems. At administrative-unit scale, averaging over many fields reduces random error substantially, and district-level forecasts can achieve better than 10 percent mean absolute percentage error in published studies for rice in Southeast Asia and wheat in Europe.
The cloud problem is not uniform: where SAR matters most
In the Sahel, the growing season coincides with a short rainy period but cloud cover is intermittent enough that Sentinel-2 typically yields several usable acquisitions per month. SAR adds value at the margins. In humid tropical zones, the picture is different. Across much of Southeast Asia, West Africa's forest-savanna transition zone, and Pacific island nations, cloud cover during the main growing season can exceed 80 percent of acquisition days. A Sentinel-2 time series for such a region may contain only two or three cloud-free scenes between planting and harvest.
This is the operational context for Satellize's crop-estimation programme in the Kingdom of Tonga. Tonga's agricultural parcels are small, fragmented, and subject to persistent cloud cover for significant parts of the year. SAR-optical fusion is not a methodological preference there; it is a practical necessity. The programme uses Sentinel-1 backscatter time series as the primary phenological signal, with Sentinel-2 composites providing calibration of vegetation indices where cloud-free observations are available.
Calibration, validation, and the ground-truth bottleneck
Every yield model needs ground truth. Crop-cut surveys, where trained enumerators harvest a known area and weigh the output, are the international standard for calibration data. They are expensive, logistically demanding, and often available only at the district level from national agricultural statistics offices. Farmer-reported yields introduce recall bias. Administrative statistics may be politically smoothed.
The practical consequence is that model validation is often conducted on a subset of the same survey data used for training, which flatters performance metrics. Independent validation against official post-harvest statistics, run in a held-out year, is the only credible test. Analysts should report validation results separately for different agro-ecological zones and yield ranges, because models frequently underestimate high yields and overestimate low ones due to training data imbalance. Spatial cross-validation, which holds out geographically contiguous units rather than random samples, gives a more honest picture of generalisation.
From pixel to policy: what a ministry actually receives
A yield map at 10 m resolution is not a decision tool. A ministry of agriculture or a food security agency needs production estimates aggregated to the administrative units used for procurement and relief planning, delivered before the harvest so that import or distribution decisions can be made in time to matter. The typical operational target is a forecast issued four to six weeks before harvest, with an update at two weeks.
Delivery formats range from tabular district-level estimates with uncertainty bounds, suitable for integration into existing food balance sheet models, to GIS layers showing within-district spatial variability for extension services. Automated pipelines that ingest new Sentinel-1 acquisitions as they arrive and update a running forecast reduce latency to one or two days after each satellite pass. Ministries that want to interrogate the model rather than simply receive its output benefit from interactive dashboards that expose the underlying time-series data alongside the forecast.
Typical figures
| Primary SAR spatial resolution | 10 m (Sentinel-1 IW mode, ground range) |
| Primary optical spatial resolution | 10 m visible/NIR, 20 m red-edge/SWIR (Sentinel-2 MSI) |
| SAR revisit (Sentinel-1) | 6 days at equator (both satellites); 12 days per satellite |
| Optical revisit (Sentinel-2) | 5 days combined; usable cloud-free acquisitions vary from daily (arid) to 2-3 per season (humid tropics) |
| SAR frequency | C-band, 5.405 GHz; VV and VH polarisations in IW mode |
| Typical forecast accuracy (district scale) | Mean absolute percentage error 8-15% for rice and wheat in published calibrated systems; field-scale uncertainty typically ±15-25% (1 SD) |
| Minimum mapping unit | Practically ~0.5 ha for SAR-optical fusion at 10 m; sub-hectare fields require Planet or similar commercial imagery for reliable delineation |
| Sentinel-1 archive depth | From April 2014 (Sentinel-1A launch); MODIS archive from 2000 for multi-year calibration |
| Forecast lead time | Operational target: 4-6 weeks before harvest; update at 2 weeks |
| Delivery formats | GeoTIFF yield maps, district-level CSV with confidence intervals, GeoJSON for GIS ingestion, PDF bulletin for ministry use |
Analytics Satellize can run
| End-of-season yield forecast (district and national) | Random forest or gradient-boosted regression on fused SAR backscatter trajectory features and Sentinel-2 vegetation index time series, calibrated against crop-cut or official statistics | Tabular district estimates with 80% prediction intervals, updated at 6 weeks and 2 weeks pre-harvest |
| In-season yield trajectory (phenological stage tracking) | Time-series analysis of Sentinel-1 VH backscatter to identify heading date, canopy closure and senescence onset; cross-checked with Sentinel-2 NDVI where cloud-free | Running GIS layer updated within 2 days of each Sentinel-1 pass |
| LAI time series (gap-filled) | Sentinel-2 LAI retrieval using ESA SNAP biophysical processor; SAR-based gap-filling using published regression between VH backscatter and LAI for the target crop | 10 m raster LAI stack per growing season, delivered as GeoTIFF archive |
| Production estimate (area × yield) | Yield forecast combined with crop-area mask from classified multispectral imagery; area uncertainty propagated into production confidence interval | National and sub-national production bulletin in PDF and CSV |
| Anomaly detection (below-normal yield risk) | Comparison of current-season SAR and NDVI trajectories against multi-year baseline climatology; z-score thresholding for early warning | Alert report flagging administrative units tracking more than 1.5 SD below historical mean at mid-season |
| Model calibration and validation report | Spatial cross-validation on held-out administrative units; independent validation against post-harvest official statistics in withheld years | Methodology and accuracy report suitable for submission to national statistics office or donor audit |
Who does the work
We can get this done for you. Satellize runs its own analyst desk and a strong science team. You do not buy a data feed and work out what it means; our people source the imagery, run the analysis described on this page, and hand you the answer with its confidence limits stated. Discuss this requirement.