Comparable site identification by spectral and morphological matching
Satellite-derived spectral signatures, building-footprint density and road-network proximity can be combined into reproducible feature vectors that identify objectively comparable development sites across large geographies, replacing subjective valuer judgement with nearest-neighbour search.
Sensors
- Sentinel-2 MSI: 13 spectral bands from 443 nm to 2190 nm at 10–60 m ground sampling distance; 5-day revisit at the equator with twin satellites. Provides the land-cover spectral signature backbone: NDVI, NDBI (Normalised Difference Built-up Index) and bare-soil indices that characterise site surface type with high consistency across dates and geographies.
- Copernicus DEM GLO-30: Global digital elevation model at 1 arc-second (~30 m) posting, derived from TanDEM-X radar. Supplies terrain slope and relative elevation for each candidate site, which matter for drainage, construction cost and flood exposure. Freely available; no tasking required.
- Maxar WorldView-3: 0.31 m panchromatic, 1.24 m multispectral, 3.7 m SWIR across 16 bands. Used selectively to extract building-footprint polygons and rooftop morphology at sites where open-data building layers are absent or outdated. Tasking adds cost and latency but resolves ambiguity that 10 m imagery cannot.
- Planet SuperDove: 3 m ground sampling distance, 8 spectral bands, near-daily revisit globally. Useful for tracking short-cycle land-cover change at candidate sites between Sentinel-2 acquisitions, and for verifying that a spectral match is stable rather than a seasonal artefact.
Why spectral matching produces a different answer from a valuer's spreadsheet
A traditional comparable analysis selects sites by postcode proximity, broad land-use category and transaction date. The result is defensible in a report but rarely reproducible: two valuers given the same brief will return overlapping but non-identical sets. That inconsistency is not incompetence; it reflects the genuine difficulty of encoding dozens of physical site attributes into a mental model.
Spectral and morphological matching starts from the opposite direction. It encodes each site as a numerical vector, then searches for the nearest neighbours in that vector space across whatever geography the client defines. The search is deterministic. Run it twice and you get the same answer. Change a feature weight and you can see exactly which comparables move in or out. That auditability is the point.
Building the feature vector: what goes in and why
A useful feature vector for development-site comparison typically draws on four groups of attributes. First, spectral land-cover fractions derived from Sentinel-2: the proportions of built surface, vegetation, bare soil and water within a defined buffer around the site centroid. The Normalised Difference Built-up Index (NDBI = (SWIR1 - NIR) / (SWIR1 + NIR)) is a well-established proxy for impervious surface density. NDVI captures greenery. Together they describe the immediate character of a site's surroundings without requiring a classified map.
Second, morphological density: building-footprint coverage ratio and, where height data are available, floor-area ratio estimated from shadow length or lidar-derived canopy height models. Third, accessibility: road-network distance to the nearest motorway junction, town centre or transit node, computed from OpenStreetMap or equivalent open graph data. Fourth, terrain: mean slope and elevation from GLO-30, which proxy construction difficulty and flood exposure. Each attribute is normalised to zero mean and unit variance before the nearest-neighbour search, otherwise distance in metres overwhelms distance in spectral units.
The choice of feature weights is where honest analysis gets uncomfortable. Doubling the weight on road proximity will pull comparables from a wider geographic radius but similar accessibility band. Doubling the spectral weight will pull comparables that look physically similar but may sit in different economic submarkets. There is no objectively correct weighting. The weights encode a hypothesis about which attributes drive value, and that hypothesis must be stated explicitly and tested.
Nearest-neighbour search at scale: practical geometry
Once each site is represented as a vector, comparison is a standard k-nearest-neighbour problem. For a national-scale database of, say, 200,000 parcels, approximate nearest-neighbour algorithms (FAISS, Annoy and similar open-source libraries) return results in milliseconds per query. The computational cost is in building and maintaining the feature database, not in the search itself.
The search can be constrained geographically (find the five closest comparables within 50 km) or left unconstrained (find the five globally closest comparables regardless of location). Unconstrained search is genuinely useful for cross-border portfolio benchmarking, where a client wants to know whether a logistics site in Poznań is more similar, on physical grounds, to sites in Łódź or in Brno. Sentinel-2 spectral data is radiometrically consistent across ESA's global archive, which makes cross-border comparison valid in a way that stitched national datasets often are not.
Validation: how to know whether the matches are any good
The standard validation approach uses known comparable pairs: transactions where two sites were explicitly treated as comparables in a published valuation or planning appeal decision. If the spectral-morphological model ranks those pairs highly, the feature weights are plausibly calibrated. If it does not, the weights need adjustment or additional features.
A second check is internal consistency: if site A is returned as a comparable for site B, then site B should appear in the top comparables for site A. Asymmetric results usually indicate a normalisation error or a feature with a skewed distribution.
Cloud cover is a genuine limit on Sentinel-2 spectral quality. In persistently cloudy climates, a single-date spectral snapshot may be unrepresentative. The standard mitigation is to compute spectral features from a cloud-free composite over a defined season rather than a single image. ESA's Sentinel Hub offers pre-built compositing services. Even so, a site undergoing active construction will show a spectral signature that does not reflect its completed character, and that ambiguity should be flagged rather than silently averaged away.
Resolution is a second honest limit. At 10 m per pixel, Sentinel-2 cannot distinguish a detached house from a terrace; it sees aggregate surface type. For fine-grained morphological matching within a single urban district, WorldView-3 or equivalent sub-metre imagery is necessary. The cost and latency of commercial tasking mean that sub-metre data is best reserved for a shortlist of candidate sites rather than the initial national-scale screening.
Where the method fits into a valuation workflow
Spectral-morphological matching is a screening tool, not a valuation. It narrows a population of thousands of candidate comparables to a tractable shortlist of ten or twenty that share measurable physical characteristics with the subject site. A qualified valuer then applies market knowledge, transaction evidence and judgement to that shortlist. The satellite step does not replace professional expertise; it changes what that expertise is applied to.
The output is most useful when the subject site is in a geography where transaction evidence is thin: secondary cities, emerging markets, or asset classes (data centres, logistics sheds, renewable energy sites) where comparables are sparse and geographically dispersed. Satellize applies this class of spectral feature analysis operationally, most visibly in the Tonga crop-estimation programme where land-cover vectors are used to stratify sampling zones. The same vector-construction logic transfers directly to property submarket delineation.
Clients typically receive a ranked comparable list as a GIS layer with attached feature-vector scores, a weight-sensitivity report showing how the ranking changes under alternative weighting assumptions, and a cloud-cover quality flag for each site's spectral data. The ranked list is reproducible: archive it, and a different analyst running the same query six months later will get the same result unless the underlying data has been updated.
Typical figures
| Primary spectral resolution | 10 m (Sentinel-2 visible and NIR bands); 20 m (SWIR bands used for NDBI) |
| Revisit cadence | 5 days at equator (Sentinel-2A+B combined); near-daily for Planet SuperDove verification layer |
| Elevation model posting | ~30 m (Copernicus DEM GLO-30, TanDEM-X derived) |
| Sub-metre morphological imagery | 0.31 m pan / 1.24 m multispectral (WorldView-3, on-demand tasking) |
| Spectral bands used in feature vector | Sentinel-2 B4 (Red), B8 (NIR), B11 (SWIR1) for NDVI and NDBI; B2–B8A for full spectral signature |
| Minimum site area for reliable spectral characterisation | Approximately 0.5 ha at 10 m resolution; smaller sites require sub-metre imagery |
| Archive depth | Sentinel-2 from 2015; Landsat-compatible spectral indices from 1984 via USGS archive |
| Geographic coverage | Global (Sentinel-2 and GLO-30 cover all land surfaces) |
| Typical feature-vector dimensionality | 10–30 attributes depending on data availability; reduced to 6–10 principal components before search |
| Deliverable formats | GeoPackage, GeoJSON, CSV with feature scores; optional WMS/WMTS tile layer |
Analytics Satellize can run
| National-scale site feature database | Sentinel-2 cloud-free seasonal composite; NDVI, NDBI and bare-soil index extraction per parcel centroid buffer; GLO-30 slope and elevation summary statistics | GeoPackage of parcels with normalised feature vectors, updated quarterly |
| Comparable site ranked shortlist | k-nearest-neighbour search (Euclidean distance in normalised feature space) with configurable geographic constraint and k value | GeoJSON of top-k comparable sites with per-attribute distance scores and aggregate similarity rank |
| Weight-sensitivity report | Repeated nearest-neighbour query across a grid of feature weight combinations; rank-stability analysis for each candidate comparable | PDF report showing which comparables are stable across weight assumptions and which are marginal |
| Spectral quality flags | Cloud-cover fraction from Sentinel-2 scene classification layer (SCL) over the composite window; flagging of sites with less than 3 clear-sky observations per season | Per-site data-quality column in the feature database; alert for sites requiring commercial tasking to fill gaps |
| Submarket boundary delineation | Unsupervised clustering (k-means or DBSCAN) on the same feature vectors used for comparable search; cluster boundaries exported as polygons | GIS polygon layer of objectively defined property submarkets with cluster-centroid feature profiles |
| Change-detection alert for comparable sites | Bi-temporal NDBI and NDVI differencing between baseline composite and current quarter; threshold-based flag for sites whose spectral character has materially shifted | Automated alert feed (GeoJSON or email digest) identifying comparables whose physical character may no longer match the subject site |
Who does the work
We can get this done for you. Satellize runs its own analyst desk and a strong science team. You do not buy a data feed and work out what it means; our people source the imagery, run the analysis described on this page, and hand you the answer with its confidence limits stated. Discuss this requirement.