JULY 10TH, 2026

Representative of What? When a New Sample Falls Outside Your Model's Domain

Overview

“Take a representative sample” carries two different jobs: whether a sample reflects the area you want to describe, and whether a new location sits inside the conditions a calibrated model has actually seen. The second question, the applicability domain, is the one most monitoring programmes skip. How to test whether a new or resampled point is in-domain, with auditable checks a verifier will accept, under the GHG Protocol LSR, Verra VM0042, and the EU CRCF. Two interactive tools to move the dials yourself.

Topics

Representativeness // Applicability domain // Sampling // MRV

Authors

Dr. Thomas Fungenzi

Share

LinkedInEmail

A cocoa sourcing programme builds a soil organic carbon baseline for its first origin. The team collects samples across the supply shed, fits a model that predicts carbon stock from soil and landscape covariates, and validates it with cross-validation. The R² is 0.78. Good enough to report. A year later the programme expands to a second origin in a different country, and reuses the same model to estimate the new baseline, because rebuilding it would cost a field campaign the budget does not have. The verifier asks one question: what is your evidence that the second origin falls within the range of conditions the model was calibrated on? The team points to the 0.78. But that number describes how well the model predicts inside the first origin's conditions. It says nothing about the second. The room goes quiet.

This is the representativeness problem that monitoring teams meet most often and name least clearly. It is not really about the sample. It is about the relationship between three things: where you measured, where you now want a number, and the slice of the world the model has actually seen.

Two questions hiding in one word

Representativeness gets used for two distinct claims, and conflating them is where trouble starts. The first claim is about a sample and a population: that the sample reflects the area you want to describe, so an average computed from it is a fair estimate of the true average. This is the domain of sampling design, and the sense we work through in How Large Should Your Soil Carbon Sampling Campaign Be? and Signal versus Noise.

The second claim is about a new point and a model. Most monitoring today does not stop at a sample average. It fits a model, an allometric equation, a soil-carbon map, a spectral calibration, a process model such as RothC, or a default emission factor, and applies it to locations where nothing was measured. A prediction at a new location is only as trustworthy as the model's experience of conditions like that location. Outside the range of conditions in the training data, the model is extrapolating, and its error is neither bounded nor knowable from the fit statistics.

These map onto the two classical modes of inference. Design-based inference gets its validity from how the sample was selected, and answers the sample-to-population question. Model-based inference gets its validity from the model being correct where it is applied, and answers the new-point-to-model question 78. A sample can be a textbook probability sample of its own population and still land a prediction well outside a borrowed model's experience. The two questions are independent, and the second is the one that silences the room.

The applicability domain, defined

The concept that governs the second question already has a name, borrowed from a field that learned the lesson early. In chemistry, quantitative structure-activity models predict a molecule's properties from its structure, and regulators noticed these models made confident, wrong predictions for molecules unlike anything in the training set. The response was to define an applicability domain: the region of input space where a model produces reliable predictions, with an explicit rule for deciding whether a new case falls inside it 12.

Environmental monitoring uses the same idea, usually without the name. Every allometric equation is fit over a range of tree diameters, and applying it beyond that range is extrapolation, which is why the practical rule is to plot your inventory's diameter distribution against the equation's fitted range, as we describe in How to Calculate the Amount of Carbon in a Tree? A diameter range is an applicability domain in one dimension. A modern soil model depends on many covariates at once, the SCORPAN factors that digital soil mapping formalises 9, and the training data occupy a cloud in that high-dimensional space. The applicability domain is the region of that space the cloud actually covers, and the whole question becomes geometric: does the new point sit inside the cloud, or off in a corner the model never visited?

How to tell whether a new point is in-domain

Domain membership is measurable, and the methods range from a five-minute check to a defensible statistic you can put in front of an auditor. Use them in order of increasing rigour.

Range and multivariate range. For each covariate, does the new site fall inside the minimum and maximum of the training data? This catches the crudest extrapolation. Its weakness is that a point can be inside the range of every variable separately and still be an unrealistic combination the model never saw, for example the rainfall of a wet site with the clay content of a dry one.

Distance to the training data. The workhorse methods measure how far the new point sits from the training cloud, accounting for the correlations among variables. Leverage (the hat-matrix diagonal) does this for linear models; Mahalanobis distance scales each direction by the training covariance; the distance to the k nearest training points, and Gower distance for mixed continuous and categorical variables, are robust non-parametric alternatives 1210.

The area of applicability, for maps. When the model produces a wall-to-wall map, the question is spatial: which pixels can the model be applied to? The area of applicability method answers exactly this 3. It weights each covariate by its importance, measures the distance in that weighted space from each pixel to the nearest training point, and compares it to the typical distance among the training points themselves. Pixels beyond a cross-validation-derived threshold are outside the area of applicability, and the model's reported error does not apply there. A regional map can look complete and authoritative while large parts of it are, by this measure, out of domain 4.

Spectral calibrations have their own version. If soil carbon is predicted from mid-infrared spectra rather than covariates, principal component analysis of the calibration set defines the space it knows, and a new spectrum's Hotelling's T² (how far inside that space) and Q-residual (how far outside it) flag samples the calibration was never trained to read.

The validation trap: why a good R² can lie

There is a specific way this goes wrong that flatters the practitioner. When a model is validated by ordinary random cross-validation, held-out points are scattered among the training points. If the data are spatially autocorrelated, and soil and vegetation data almost always are, then every held-out point has a near neighbour in the training set. The model is being tested almost entirely inside its domain, on easy points. Spatially structured validation, which holds out whole blocks of space, repeatedly reveals that models with excellent random cross-validation scores predict poorly when they have to reach 56. A regional map can be unbiased on average across its extent and still be wrong at every individual project inside it.

The honest response is not simply to swap random for spatial cross-validation and move on, because the right validation design is itself debated: spatial blocking can be pessimistic if it holds out regions the model would never be asked to predict, and the most defensible option, where the budget allows, is a separate validation sample drawn by probability sampling from the target area, which estimates map accuracy without leaning on the model at all 78. The point that survives the debate is simpler: an accuracy number is only meaningful together with a statement of where, in covariate and geographic space, it was measured, and whether the place you now want a prediction resembles that.

A decision rule you can defend

Put together, domain testing is a short, repeatable workflow that fits on one page of a monitoring plan. Define the calibration domain (the multivariate distribution of covariates in the training data). Locate each new point relative to it: the per-variable range check first, then a distance metric. Classify it as in-domain, borderline, or out-of-domain against thresholds fixed in advance. Then act on the class: in-domain points get the prediction as is; borderline points get it with an explicitly inflated uncertainty and a flag; out-of-domain points do not get a borrowed number at all, they trigger local reference sampling and re-estimation.

The value of writing this down is that it converts an argument (“we think the new region is similar enough”) into evidence (“here is where each new point falls, and here is the rule we applied”). That is the difference between a claim a verifier accepts and one a verifier probes.

A worked example: one model, two regions

A soil-carbon model is calibrated on full-sun and light-shade farms in one sourcing region. Sourcing expands to a second region of established shaded agroforestry. Five of the model's covariates, checked against the new region's representative site:

CovariateCalibration range (A)New site (B)In range?
Clay (%)12–3441No (above)
Mean annual rain (mm)900–15001850No (above)
Mean annual temp (°C)23–2725Yes
Elevation (m)40–320180Yes
Shade canopy cover (%)0–1555No (above)

Three of five covariates fall outside the calibrated range, and the two that matter most for carbon, rainfall and shade, are among them. A Mahalanobis distance or an area-of-applicability index would put this site well beyond the training cloud. The verdict is unambiguous: the model is out of domain here, and applying it would report a number with no support. This is the same finding, in miniature, that independent evaluations reach at continental scale, where soil-carbon relationships calibrated on temperate row-crop systems do not transfer to shaded perennial systems without local re-estimation. We take up model choice and local recalibration in How Do You Model Soil Carbon Change Over Time?

Resampling, and the domain that moves

There is a second, quieter way representativeness fails, and it shows up precisely when a programme does everything else right. Monitoring is built on resampling: revisit the same network over time, compare, and report the change. The unstated assumption is that the network stays representative of the system it monitors. Systems move. Management changes, a shade programme matures, a drought reshapes the soil moisture regime, land use shifts at the edges of the supply shed. A model calibrated once against the original system can slide out of domain relative to the system that now exists, without anyone re-running a check. Statisticians call this covariate shift or concept drift; in the field it looks like a baseline that quietly stops describing the present.

Why verifiers are starting to ask

This is no longer only a methodological nicety. The accounting frameworks that govern carbon claims are converging on the demand that uncertainty be quantified and disclosed, and a prediction applied outside a model's domain is uncertainty that has been hidden rather than quantified. The GHG Protocol Land Sector and Removals Standard requires defensible quantification with disclosed uncertainty. The IPCC 2019 Refinement pushes toward higher-tier, locally-appropriate factors over borrowed defaults precisely because defaults are out-of-domain for many of the places they get applied 12. Verifiers are beginning to ask for spatially-structured or independent validation rather than a headline R², and for evidence that a reused model fits the site it is being reused on. The teams that will pass these questions are the ones who tested domain membership before they were asked, and can show the work.

Key takeaways

01

Representativeness carries two separate questions: is the sample representative of the population (sampling design), and is a new point inside the model's calibrated domain (applicability). The second is the one most programmes skip.

02

A sample can be perfectly representative of its own population and still fall outside a borrowed model's applicability domain. The two are independent.

03

The applicability domain is the region of covariate space the training data actually cover. An allometric equation's diameter range is this idea in one dimension; a soil-carbon map's domain is the same idea in many.

04

Domain membership is measurable: per-variable range, then distance to the training data (leverage, Mahalanobis, k-nearest-neighbour, Gower), the area of applicability for maps, and Hotelling's T² with Q-residuals for spectral calibrations.

05

A high random cross-validation score can be an artefact of spatial autocorrelation. Judge a model with spatially-structured or, better, independent probability-based validation, and always report where the accuracy was measured.

06

Use a fixed decision rule: define the domain, locate the new point, classify it, and either accept, inflate uncertainty, or trigger local recalibration. The domain drifts, so re-check membership at every resampling cycle.

References

  • 1.Netzeva, T.I. et al. (2005). Current status of methods for defining the applicability domain of (quantitative) structure-activity relationships. ATLA Alternatives to Laboratory Animals, 33(2), 155–173.
  • 2.Jaworska, J., Nikolova-Jeliazkova, N., Aldenberg, T. (2005). QSAR applicability domain estimation by projection of the training set in descriptor space, a review. ATLA, 33(5), 445–459.
  • 3.Meyer, H., Pebesma, E. (2021). Predicting into unknown space? Estimating the area of applicability of spatial prediction models. Methods in Ecology and Evolution, 12(9), 1620–1633.
  • 4.Meyer, H., Pebesma, E. (2022). Machine learning-based global maps of ecological variables and the challenge of assessing them. Nature Communications, 13, 2208.
  • 5.Ploton, P. et al. (2020). Spatial validation reveals poor predictive performance of large-scale ecological mapping models. Nature Communications, 11, 4540.
  • 6.Roberts, D.R. et al. (2017). Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography, 40(8), 913–929.
  • 7.Wadoux, A.M.J-C., Heuvelink, G.B.M., de Bruin, S., Brus, D.J. (2021). Spatial cross-validation is not the right way to evaluate map accuracy. Ecological Modelling, 457, 109692.
  • 8.Brus, D.J., Kempen, B., Heuvelink, G.B.M. (2011). Sampling for validation of digital soil maps. European Journal of Soil Science, 62(3), 394–407.
  • 9.McBratney, A.B., Mendonça Santos, M.L., Minasny, B. (2003). On digital soil mapping. Geoderma, 117(1–2), 3–52.
  • 10.Sheridan, R.P. et al. (2004). Similarity to molecules in the training set is a good discriminator for prediction accuracy in QSAR. Journal of Chemical Information and Computer Sciences, 44(6), 1912–1928.
  • 11.Chave, J. et al. (2014). Improved allometric models to estimate the aboveground biomass of tropical trees. Global Change Biology, 20(10), 3177–3190.
  • 12.IPCC (2019). 2019 Refinement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories. Intergovernmental Panel on Climate Change.

Share this article

LinkedInEmail