SEPTEMBER 4TH, 2026

How Do You Calibrate and Validate a Model? A Field Guide from Soil Carbon to Tree Biomass

Overview

Calibration means four different things, and most arguments about validated models are two people using different ones. This piece separates the senses, then walks the workflow they share: why fitting well proves almost nothing, what each level of a data split actually estimates, how the cross-validation scheme you pick encodes a claim about the prediction task, which metrics are worthless alone, and what it takes for a stated uncertainty to be honest. Three interactive tools, on overfitting, on validation schemes, and on the many parameter sets that fit equally well.

Topics

Modelling // Validation // MRV // Uncertainty

Authors

Dr. Thomas Fungenzi

Share

LinkedInEmail

A model produces a number. Someone then has to decide whether to believe it, and that person is usually not the person who built it. They are a verifier reading a project document, or a sustainability lead deciding whether to put the figure in a disclosure, or a reviewer who has been handed a manuscript. The number arrives with a name attached, RothC or a Chave equation or a random forest, and the name does a surprising amount of work. It should do almost none.

Nothing about a model's architecture makes its output trustworthy. Trust comes from evidence about how the thing behaves on data it was not built from, and that evidence is either in the documentation or it is not. In practice it is often not. A project document will describe the model in detail, name the parameters, cite the original publication, and say remarkably little about what was held back, how it was held back, and what the errors looked like when the model met it.

This piece is about that missing half. It sits between two others on this site. How to choose a soil carbon model comes first, and how to know where a fitted model may be applied comes after. This is the middle: having picked something, how do you fit it, how do you test it, and what would it take for the test to mean anything.

The answer, before the reasoning

Calibration is four words

Instrument traceability, empirical fitting, process-model parameter estimation, and the calibration of a probability statement. They have different success criteria and different failure modes.

Fitting well proves nothing

Training error falls as a model grows more flexible, without limit and without meaning. Only error on data withheld from fitting carries information about the next prediction.

The split scheme is a claim

Random k-fold, leave-site-out and temporal holdout answer different questions. A cross-validated error reported without its scheme is not interpretable, and the schemes are not comparable to each other.

No metric works alone

Bias, spread and skill against a baseline are three separate questions, and correlation answers none of them. A model can correlate perfectly with the truth and be wrong by half.

Coverage is checkable

A ninety percent interval either contains the truth about ninety percent of the time on held-out data or it does not. This is a measurement, not a claim, and it is rarely reported.

None of this is advanced statistics. It is the ordinary discipline of the field, applied consistently, and most of the distance between a defensible modelled number and an indefensible one is covered by doing it at all.

One word, four jobs

Sit in on a validation call and count the senses. The developer says the model was calibrated on twelve long-term trials. The verifier asks whether it has been validated. The remote-sensing lead mentions that her calibration used four hundred spectra with a hundred held back. The statistician, quiet until now, asks whether the intervals are calibrated. Four people, one word, and at least three different operations on the table.

None of them is using the word wrongly. Calibration carries four established meanings across the disciplines that meet in a carbon project, and they are not shades of one idea. They differ in what gets adjusted, in what counts as success, and in how they fail. Sliding between them is how a room ends up agreeing on a word while disagreeing about everything the word was supposed to settle.

The oldest sense is metrological. An instrument produces an indication, and calibration relates that indication to values supplied by measurement standards, each carrying its own uncertainty, then turns that relation into a way of getting a measurement result from a reading1. A dry combustion analyser is calibrated against certified reference materials of known carbon content. One detail here is the source of a good deal of loose talk later: calibration does not change the instrument. Setting a system so that it gives prescribed indications is a separate operation with its own name, adjustment, and the vocabulary is explicit both that the two should not be confused and that calibration is a prerequisite for adjustment1. Calibration in this sense produces knowledge about an instrument, not a better instrument.

The second sense is the statistical fitting of an empirical model. An allometric equation relating diameter to biomass, or a spectral model relating reflectance to carbon concentration, has coefficients estimated from data. The data used to estimate them is the calibration set; whatever is held back is the validation set. Nothing here is being made traceable to a standard. A functional form is being fitted, and the question that matters is how it behaves on data it has not seen. Soil spectroscopy is where the first two senses collide most often. Predicting a soil property from a diffuse reflectance spectrum means fitting a regression on a calibration set and testing it on samples held back2, so the word there names the regression rather than any adjustment of the spectrometer, even though there is a spectrometer in the room.

The third sense is parameter estimation for a process model. RothC's decomposition rate constants, or the analogous parameters in Century and DayCent, are not quantities anyone can go and measure on a given farm. They are inferred by running the model and adjusting them until it reproduces an observed time series. Every candidate parameter set needs another full run to learn how far it misses, which is what makes this an inverse problem and not a fit3. The difference from the second sense is not technique but obligation: these parameters are supposed to mean something physical. A decomposition rate tuned to whatever value minimises the residuals has stopped being a decomposition rate and become a free parameter wearing a mechanistic label. It is also where several different parameter sets turn out to reproduce the same observations about equally well, a problem with its own name and its own literature4.

The fourth sense concerns probability statements rather than predictions. A forecast is calibrated when the events it assigns ninety percent probability happen about ninety percent of the time, and a ninety percent prediction interval is calibrated when it contains the truth about ninety percent of the time. This is the one sense in which a model can be flawlessly calibrated and completely useless. Predict the long-run mean every time, attach intervals wide enough to swallow the variation, and coverage will be excellent while the forecast tells nobody anything. Calibration here is necessary but has to be judged alongside how narrow the statement is, which the forecasting literature calls sharpness5.

1 // Instrument calibrationMetrology

What gets adjusted

Nothing on the instrument. The relation between its indication and a reference value of known uncertainty.

Success looks like

The result is traceable to a standard, with a stated uncertainty.

Fails when

Drift between calibrations, or a traceability chain assumed rather than documented.

2 // Fitting an empirical modelStatistics // chemometrics

What gets adjusted

Coefficients of a chosen functional form, estimated from a calibration set.

Success looks like

Predictive skill on data held back from fitting.

Fails when

Overfitting, and validation data that is not independent of the calibration data.

3 // Process-model parameter estimationBiogeochemical modelling

What gets adjusted

Parameters meant to carry physical meaning, inferred by inversion against observed time series.

Success looks like

The model reproduces sites it was not tuned on, with parameters that stay physically plausible.

Fails when

Equifinality, and parameters bent past physical sense to absorb structural error.

4 // Calibration of a probabilityForecasting

What gets adjusted

The width and shape of the stated uncertainty, not the central prediction.

Success looks like

Achieved coverage matches nominal coverage, at useful sharpness.

Fails when

Intervals so wide that coverage is met and the statement is empty.

The consequence for a carbon project is specific enough to watch for. A verifier asks whether the model has been validated. The developer answers that it was calibrated against local long-term trials, which is a true and relevant statement about the third sense. The verifier may have been asking about the second, meaning a held-out test, or about the fourth, meaning whether the uncertainty attached to the final number has any coverage behind it. Three reasonable questions, one answer, and the gap between them does not show up in the minutes.

An older argument sits underneath this one. Verification belongs to closed systems: a purely formal structure can be proved by symbolic manipulation, and an algorithm inside a program can be checked the same way. A model built from such components is not itself closed, because the parameters it needs are never completely known, and the claim that a model of an open environmental system has been validated in any absolute sense has been contested since the mid-nineties6. Textbooks written for modellers draw the same line for practical reasons. One treats a model as a scientific hypothesis when the question is whether its processes behave as the real ones do, a hypothesis that can be invalidated but never definitively validated. It treats the same model as an engineering tool when the question is not how it works but how well it works, and calls that evaluation7. The other declines the word validation for the third phase of a modelling project and calls that phase evaluation. Comparing predictions with observations is only one of several things that make a model useful for its purpose, and validation suggests that a correct model exists to be found8. Nothing in this piece requires taking a side in that argument. It does require reading validation as shorthand for a specific test against specific data, rather than as a certificate.

Fit is cheap

Take twelve trees, measure diameter and destructively sample biomass, and fit a polynomial. At degree two the curve passes near the points. At degree eleven it passes through every one of them exactly. The residuals are zero, the coefficient of determination is one, and the fit is perfect in the only sense the word can bear in-sample. Between the points the curve swings to physically impossible values, and a thirteenth tree will land nowhere near it.

This is not a pathology of polynomials. It is the general behaviour of flexible models, and the reason in-sample performance carries so little information. As a model gains freedom, its training error falls monotonically, with no floor other than zero and no point at which the decline signals anything. Error on data withheld from fitting behaves differently: it falls while the added flexibility is buying structure that generalises, then turns and climbs while the same flexibility starts absorbing noise specific to the sample in hand9.

The mechanism has a standard decomposition. Expected error at a new point splits into a bias term, which is how far the model is wrong on average, a variance term, which is how much the fitted model moves when the training sample is redrawn, and an irreducible term, the variance of the target around its own true mean, which no amount of modelling can remove. Flexibility trades the first against the second. A rigid model is more biased and steadier, a flexible one less biased and jumpier, and the useful place is not at either end9.

Expected prediction error at a new point
E[(y − ŷ)²] = Bias(ŷ)² + Var(ŷ) + σ²irreducible
y, ŷ
The true value at a new point, and the model's prediction of it.
Bias(ŷ)²
How far the model is wrong on average.
Var(ŷ)
How much the fitted model moves when the training sample is redrawn.
σ²
The variance of the target around its own true mean. No amount of modelling removes it.

One corollary is worth stating on its own, because it survives in project documentation long after it should have been retired. The in-sample coefficient of determination never falls when a predictor is added. Add a column of random numbers to a regression and R² will rise. Adjusted forms penalise the parameter count and correct part of the problem, but no in-sample statistic answers the question that matters, which is what happens on the next site rather than on this one.

In our own field the trap has a specific shape. A locally fitted allometric equation, estimated on a few dozen trees from one plantation, will beat a published pantropical equation on those same trees essentially every time. That comparison is not evidence that the local equation is better, and it is routinely presented as if it were. The honest comparison holds back trees the local equation never saw, which is exactly the comparison a small destructive dataset makes painful, because every tree withheld is a tree not used for fitting. The tension is real. Pretending it away by reporting in-sample skill is not a resolution of it.

The same failure scales up to maps. Large-scale ecological models validated by randomly holding out pixels have reported strong predictive performance that collapsed once the held-out data was separated spatially from the training data, because neighbouring pixels are not independent observations and a random split quietly hands the model the answers10. That specific failure is the subject of a companion piece and is not repeated here; what belongs here is the general lesson that the split does the work, not the metric computed after it.

Anatomy of a validation

The word validation covers a hierarchy that is worth taking apart, because the levels estimate different things and are routinely collapsed into one. In the cleanest version there are three roles for data, and a given observation plays exactly one of them.

Training set

Estimates the parameters. Error measured here tells you the model can reproduce data it has already seen, which is not in doubt.

Validation set

Chooses between candidates: which functional form, how many terms, which hyperparameters. Error here guides selection, and is biased low the moment it is used for it.

Test set

Estimates how the chosen model will perform on new data. It is unbiased only while it remains untouched, and it can be spent exactly once.

The distinction between the second and third roles is the one that goes missing. Suppose you try eight model configurations, evaluate each on the same held-out block, pick the one that scores best, and then report that score as the model's expected performance. The number is too good, and it is too good for a reason that has nothing to do with dishonesty: you selected the maximum of eight noisy estimates, so you have selected partly for genuine skill and partly for favourable noise. The more configurations you tried, the more of the reported score is noise you went looking for11.

When data is too scarce to hold back a test set that is never touched, the established answer is to nest the procedure: an inner loop that does selection, an outer loop that only ever sees models already chosen, so the outer estimate is never used to choose anything11. It costs computation and it costs nothing in data, which makes its absence from most applied work hard to defend on practical grounds.

Underneath all three levels sits a load-bearing word that the arithmetic cannot check for you. Held-out data has to be independent, and independence here is a claim about the world rather than about the spreadsheet. Two cores from the same field ten metres apart are two rows in a table and roughly one observation about soil. Two years from the same site are not two independent tests of a temporal model. Splitting a table at random guarantees that the rows are separated. It guarantees nothing at all about whether the information is.

The cross-validation family, and what each member assumes

Holding back a single test set wastes data, and with a few hundred observations that waste hurts. Cross-validation recovers most of it by rotating the role of the holdout: partition the data, train on all parts but one, test on the part left out, repeat until every part has taken a turn, then average. The idea is old, simple, and almost universally reported in a form that omits the only detail a reader needs.

The detail is how the partition was made. Cross-validation is not one procedure but a family, and its members differ in how they assign observations to folds. That assignment is not a technical preference. It is a statement about what the model will be asked to do, and each choice estimates the error of a different task12.

Random k-fold

Question it answers

How well does this predict another observation drawn from the same pool?

Appropriate when

Genuinely exchangeable observations with no grouping, no spatial clustering and no time ordering.

Breaks down when

Clustered or autocorrelated data, where a fold's near neighbours sit in the training set and leak the answer.

Leave-one-out

Question it answers

The same question, with folds of one.

Appropriate when

Small datasets, and linear models where it can be computed in closed form rather than refitted n times.

Breaks down when

Its estimate has high variance, and it inherits every leakage problem of random k-fold.

Leave-site-out

Question it answers

How well does this predict a site that contributed nothing to fitting?

Appropriate when

Any model meant to be applied somewhere it has not been calibrated, which covers most project work.

Breaks down when

Pessimistic if the intended use really is interpolation within known sites, and it needs enough sites to average over.

Temporal holdout

Question it answers

How well does this predict forward from the data in hand?

Appropriate when

Anything projecting a trajectory, where training on the future to predict the past would be meaningless.

Breaks down when

Spends the most recent and usually most relevant data on testing, and one holdout window is one noisy estimate.

Run the same dataset through several of these and the reported error moves, often a great deal. That movement is the point rather than a nuisance. A cross-validated error is a conditional statement, and the condition is the partition. Two numbers produced under different schemes are not two estimates of one quantity, so comparing them, or comparing your leave-site-out result against a published random k-fold result, is a category error dressed up as a benchmark. The tool below runs one site-clustered dataset through three schemes so the size of the movement is visible rather than asserted.

The spatial case has more structure than this section gives it: how far apart a block has to be before it counts as independent, how blocking can overshoot into pessimism, and what to do when the intended prediction targets are themselves clustered. That belongs to the piece on representativeness, which starts where this section stops.

Metrics that only work in company

Once the split is settled, something has to be computed across it, and here the habit is to report one number. No single number describes a model's errors, because errors have at least three independent properties: whether they lean, how big they typically are, and whether the model beats a trivial alternative. A metric answers one of those and is silent on the others.

Mean error, the average signed residual, measures lean. It is the one that matters most for carbon accounting, because a systematic offset does not shrink when the project gets bigger: bias aggregates, while random error partly cancels. It is also the metric most easily satisfied by accident, since large positive and negative errors average to a comfortable zero.

Root mean square error measures typical size, in the units of the quantity, and weights large misses heavily because it squares before averaging. Mean absolute error asks the same question without that weighting. Reporting both is informative: when RMSE greatly exceeds MAE, the error distribution has a tail, and a tail in a validation set usually means a subpopulation the model handles badly rather than bad luck.

Neither says whether the model is worth having. That needs a baseline, which is what model efficiency supplies: it compares the model's squared error against the error of simply predicting the mean of the observations, and reports the proportion of that initial variance the model accounts for13. One is perfection, zero says the mean would have done as well, and negative says the mean would have done better. Negative values turn up in published soil and biomass work more often than the tone of the surrounding text usually admits. That is an impression from reading, not a tallied figure.

Model efficiency, against the mean of the observations
EF = 1 − Σ(oᵢ − pᵢ)² / Σ(oᵢ − ō)²
EF
Model efficiency. One is perfect, zero says the mean would have done as well, negative says it would have done better.
oᵢ
The observed value for sample i.
pᵢ
The value the model predicted for that same sample.
ō
The mean of the observed values, which is the no-model baseline.

The metric to treat with the most suspicion is the correlation coefficient and its square. Correlation measures whether two series move together, not whether they agree. A model that returns exactly half the true biomass of every tree correlates with the truth perfectly, R² of one, and is wrong by fifty percent everywhere. This is not a contrived case: a systematically scaled prediction is precisely what a miscalibrated model produces, and it is precisely what R² is blind to. Lin's concordance correlation coefficient exists for this reason. It rewards agreement with the one-to-one line rather than with any line, so it falls when a model is proportionally biased even while correlation stays perfect14.

Which leaves the plot everyone draws and a surprising number draw wrongly. When observed and predicted values are plotted against each other and a line is fitted to test for a slope of one and an intercept of zero, the choice of axis is not cosmetic. Regressing predicted on observed and regressing observed on predicted give different slopes, and only one of them tests the hypothesis about model performance that the plot is meant to test. The convention with the statistical argument behind it puts observed values on the vertical axis and predicted on the horizontal, because the predictions are the quantity being evaluated and the regression assumptions fall where they should15.

What a fuller evaluation looks like is on record. When the STICS soil-crop model was tested with one standard parameter set across fifteen crops and a wide range of French soils and climates, the assessment was split three ways. Accuracy came from RMSE, relative RMSE and model efficiency, together with whether the error was mostly bias or dispersion. Robustness asked whether the size of the error depended on the crop or the environment. Behaviour asked whether the model reproduced the differences observed between contrasting environmental conditions and management practices16. The three answer different questions, and a validation section that reports only the first has answered one of them.

Calibrating a process model is a different job

Everything above assumed a model whose parameters are free to take whatever values fit best, because they carry no meaning beyond the fit. Process models break that assumption. RothC's pool decomposition rates, Century's turnover constants and DayCent's nitrogen parameters are written as physical quantities, and the model's claim to project beyond its calibration data rests entirely on their being physical. That claim is what makes tuning them delicate.

The instruments for the job are the long-term experiments, some running for over a century, where soil carbon has been followed under fixed treatments. They are the only data that constrains a decadal model on the timescale it claims to work, which also means there are not many of them, they are unevenly distributed across the world's climates and soils, and the same handful appears in calibration after calibration. The workshop that began testing soil organic matter models against long-term datasets in a systematic way gave the reason: changes in soil organic matter are slow, so only long-term data can test a model of them. Its datasets came, by its own account, from within the temperate region17. A model tuned on temperate arable trials and applied to tropical agroforestry has been calibrated against an instrument that never measured anything resembling the target18.

Practice runs along a spectrum. At one end, a modeller adjusts parameters by hand until simulated and observed curves look close, which is fast, undocumented and unrepeatable. In the middle, an optimiser minimises a loss function, which is repeatable and returns a single best set with no statement of how much better it was than the alternatives. At the far end, Bayesian calibration treats the parameters as random variables, combines a prior with the likelihood of the observations, and returns a posterior distribution: not one answer but a description of every answer the data tolerates, with its own uncertainty attached19.

That framing also makes room for something the other two ignore. Real models are structurally wrong, and structural error does not behave like measurement noise. Calibration returns a best-fitting value, not a physical one, and a decomposition rate can end up at whatever value compensates for a process the model does not represent. It still fits. It is no longer a decomposition rate, and a projection that leans on its physical meaning is leaning on something that has quietly stopped being physical19.

The other reason to prefer a distribution over a best fit is that the best fit is usually not much better than a large number of quite different alternatives. Many parameter sets reproduce an observed carbon trajectory about equally well, and they do not agree with one another about the future, because fitting the same past is a much weaker constraint than it looks. The response is not to pick the winner and report its projection, but to carry the acceptable set forward and let it produce a range4. The tool below makes that visible on a one-pool toy model with two unknowns: the region of parameter space consistent with the data is wide, it narrows as more data arrives, and it never collapses to a point.

One further interaction is easy to miss. Where a process model starts is itself a calibration decision. The split of the initial stock across fast, slow and inert pools is rarely measurable, usually estimated, and materially changes a thirty-year projection. Calibrating parameters against a trajectory whose starting point was itself assumed means fitting two unknowns to one curve. That problem, and the spin-up procedure that is the defensible answer to it, is worked through in the piece on choosing a soil carbon model.

An uncertainty statement is a prediction, and can be tested

Return to the fourth sense of calibration, which is where most reported model uncertainty quietly fails. A project states its estimate as a central value with a ninety percent interval. That interval is not a caveat or a gesture at humility. It is a falsifiable claim: across many such statements, the truth should fall inside about ninety percent of them. The fraction that actually do is the achieved coverage, and comparing it against the nominal ninety is an ordinary measurement that almost nobody performs.

When it is performed, intervals are usually too narrow. Three causes account for most of it. Parameter uncertainty gets dropped, so the interval reflects residual scatter around one fitted model rather than uncertainty about which model was right. Structural error gets ignored, because no term in the arithmetic represents the possibility that the model form is wrong. And correlated observations get counted as independent ones, which inflates the effective sample size and shrinks every interval computed from it, a mechanism worked through in the representativeness piece.

The correction is not to widen intervals until they cover, which is the failure mode from the other direction and produces statements too vague to act on. Coverage and sharpness are separate axes, and the target is the narrowest interval that still covers at its nominal rate5. What makes this tractable in practice is that both are measurable on held-out data, using the same splits already built for evaluating the central prediction. A validation that reports achieved coverage alongside RMSE has answered a question a verifier is entitled to ask and usually does not know how to phrase.

What the standards actually ask for

Two questions come up whenever this material meets a real project. Does anyone require this, and does anyone check. The short answers are that model-based accounting does carry validation requirements, that they are less specific than the statistical discussion above, and that the gap between what a standard requires and what makes a number defensible is where most of the professional judgement lives.

The IPCC framework sets the vocabulary. Tier 3 is the level at which an inventory uses models or direct measurement rather than default emission factors, and the requirement attached to the model-based version is stated plainly: Tier 3 model-based inventories require measurements to calibrate and validate the models used to estimate carbon stock changes20. Both halves of this article sit in that sentence, joined by an and rather than treated as one thing. Measurements to calibrate, and measurements to validate. Whether a given inventory keeps them apart is not something the sentence can enforce.

In the voluntary market the requirements are more explicit. Verra's methodology for improved agricultural land management permits a biogeochemical model to simulate soil carbon change, and attaches machinery to it: a model validation report, assessed by an independent modeling expert and resubmitted as later measurement campaigns extend the calibration and validation dataset, and an uncertainty deduction applied to the issued reductions21. A separate module carries the calibration and validation requirements themselves22. The deduction is the part worth naming, because the design principle is a good one. It scales with the relative standard error of the estimate, so uncertainty is not merely disclosed but priced, and a project that cannot demonstrate a well-constrained model issues fewer credits for the same measured carbon.

On the partition itself, the module is more specific than most practitioners expect, and more demanding. It requires that calibration and validation data be demonstrably independent, and it says what that means: the datasets must not overlap in experimental research location, and must not be drawn from the same experimental study22. Splitting a single pool is allowed, and k-folding is named as a way to do it, but the independence condition still binds, so the folds have to respect study and location rather than being drawn at random across rows. In substance that is much closer to leave-site-out than to the random holdout most people picture, and a team that satisfied itself with a random k-fold over clustered plots has not met the requirement, however good the resulting number looks. How a verifier reads the resulting claim is the subject of the companion piece.

What to ask of any modelled number

None of the following requires the person asking to be a modeller. They are the questions that separate a number with evidence behind it from a number with a citation behind it, and they can be asked in a meeting.

Ten questions

  1. 1When you say the model is calibrated, which of the four senses do you mean?
  2. 2What data was held back from fitting, and how was it chosen?
  3. 3In what way is that held-back data independent: different site, different year, different region?
  4. 4How many model configurations were evaluated before this one was reported?
  5. 5Was the data used to choose the model the same data used to report its performance?
  6. 6What is the bias, separately from the RMSE, in the units of the quantity?
  7. 7Against what baseline is the skill measured, and does the model beat predicting the mean?
  8. 8On the observed-against-predicted plot, what are the slope and the intercept?
  9. 9Has the stated uncertainty been checked for coverage, and what fraction of held-out observations fell inside it?
  10. 10Are the conditions where this will be applied inside the range it was calibrated on?

A team that can answer all ten has done the work. A team that treats several of them as unreasonable has told you something useful too.

The reason to ask them is not suspicion. Most modelled numbers in this field are produced by people doing their best with small datasets and real deadlines, and the gaps are usually inherited conventions rather than corner-cutting. But a number that goes into a disclosure, a credit issuance or a policy document has left the modeller's hands, and at that point the only thing standing between it and a decision is whether anyone asked what it was tested against.

Key takeaways

  1. Calibration names four different operations: making an instrument traceable, fitting an empirical model, estimating the parameters of a process model, and calibrating a probability statement. They fail in different ways, and most disputes about validated models are a mismatch between two of them.

  2. Training error falls without limit as a model gains flexibility and carries no information about the next prediction. Only error on data withheld from fitting does.

  3. The in-sample coefficient of determination never falls when a predictor is added, including a predictor of random numbers. A locally fitted allometry beating a global one on its own calibration trees is not evidence of anything.

  4. Data used to choose a model cannot also measure its performance. Selecting the best of several configurations on one holdout means the reported score is partly the noise you selected for; nesting the procedure fixes it at no cost in data.

  5. The cross-validation scheme is a claim about the prediction task. Random k-fold, leave-site-out and temporal holdout answer different questions and are not comparable, so an error reported without its partition cannot be interpreted.

  6. No metric works alone. Bias, spread and skill against a baseline are separate properties, and correlation measures none of them: a model returning exactly half of every true value correlates perfectly with the truth.

  7. For process models, many parameter sets fit the same trajectory about equally well and disagree about the future. Carrying the acceptable set forward is more defensible than reporting the projection of the single best fit.

  8. A stated uncertainty is a testable prediction. Achieved coverage on held-out data either matches the nominal rate or does not, it is measurable with splits already built for the central estimate, and it is almost never reported.

References

  • 1.JCGM (Joint Committee for Guides in Metrology) (2012). International vocabulary of metrology: Basic and general concepts and associated terms (VIM), 3rd edition. JCGM 200:2012. Bureau International des Poids et Mesures. bipm.org
  • 2.Viscarra Rossel, R.A., Walvoort, D.J.J., McBratney, A.B. et al. (2006). Visible, near infrared, mid infrared or combined diffuse reflectance spectroscopy for simultaneous assessment of various soil properties. Geoderma, 131(1–2), 59–75. doi:10.1016/j.geoderma.2005.03.007
  • 3.Haefner, J.W. (2005). Modeling Biological Systems: Principles and Applications, 2nd edition. Springer, New York. doi:10.1007/b106568
  • 4.Beven, K., Freer, J. (2001). Equifinality, data assimilation, and uncertainty estimation in mechanistic modelling of complex environmental systems using the GLUE methodology. Journal of Hydrology, 249(1–4), 11–29. doi:10.1016/s0022-1694(01)00421-8
  • 5.Gneiting, T., Balabdaoui, F., Raftery, A.E. (2007). Probabilistic Forecasts, Calibration and Sharpness. Journal of the Royal Statistical Society Series B: Statistical Methodology, 69(2), 243–268. doi:10.1111/j.1467-9868.2007.00587.x
  • 6.Oreskes, N., Shrader-Frechette, K., Belitz, K. (1994). Verification, Validation, and Confirmation of Numerical Models in the Earth Sciences. Science, 263(5147), 641–646. doi:10.1126/science.263.5147.641
  • 7.Wallach, D., Makowski, D., Jones, J.W. et al. (2019). Model evaluation. In: Working with Dynamic Crop Models: Methods, Tools and Examples for Agriculture and Environment, 3rd edition, pp. 311–373. Academic Press. doi:10.1016/B978-0-12-811756-9.00009-5
  • 8.Grant, W.E., Swannack, T.M. (2008). Ecological Modeling: A Common-Sense Approach to Theory and Practice. Blackwell Publishing. ISBN 978-1-4051-6168-8. wiley.com
  • 9.Hastie, T., Tibshirani, R., Friedman, J. (2009). The Elements of Statistical Learning. Springer Series in Statistics. doi:10.1007/978-0-387-84858-7
  • 10.Ploton, P., Mortier, F., Réjou-Méchain, M. et al. (2020). Spatial validation reveals poor predictive performance of large-scale ecological mapping models. Nature Communications, 11(1), 4540. doi:10.1038/s41467-020-18321-y
  • 11.Varma, S., Simon, R. (2006). Bias in error estimation when using cross-validation for model selection. BMC Bioinformatics, 7(1), 91. doi:10.1186/1471-2105-7-91
  • 12.Roberts, D.R., Bahn, V., Ciuti, S. et al. (2017). Cross‐validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure. Ecography, 40(8), 913–929. doi:10.1111/ecog.02881
  • 13.Nash, J.E., Sutcliffe, J.V. (1970). River flow forecasting through conceptual models part I — A discussion of principles. Journal of Hydrology, 10(3), 282–290. doi:10.1016/0022-1694(70)90255-6
  • 14.Lin, L.I.K. (1989). A Concordance Correlation Coefficient to Evaluate Reproducibility. Biometrics, 45(1), 255. doi:10.2307/2532051
  • 15.Piñeiro, G., Perelman, S., Guerschman, J.P. et al. (2008). How to evaluate models: Observed vs. predicted or predicted vs. observed?. Ecological Modelling, 216(3–4), 316–322. doi:10.1016/j.ecolmodel.2008.05.006
  • 16.Coucheney, E., Buis, S., Launay, M. et al. (2015). Accuracy, robustness and behavior of the STICS soil–crop model for plant, water and nitrogen outputs: Evaluation over a wide range of agro-environmental conditions in France. Environmental Modelling & Software, 64, 177–190. doi:10.1016/j.envsoft.2014.11.024
  • 17.Powlson, D.S., Smith, P., Smith, J.U. (eds.) (1996). Evaluation of Soil Organic Matter Models Using Existing Long-Term Datasets. NATO ASI Series I: Global Environmental Change, vol. 38. Springer, Berlin. doi:10.1007/978-3-642-61094-3
  • 18.Smith, P., Smith, J.U., Powlson, D.S. et al. (1997). A comparison of the performance of nine soil organic matter models using datasets from seven long-term experiments. Geoderma, 81(1–2), 153–225. doi:10.1016/s0016-7061(97)00087-6
  • 19.Kennedy, M.C., O'Hagan, A. (2001). Bayesian Calibration of Computer Models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 63(3), 425–464. doi:10.1111/1467-9868.00294
  • 20.IPCC (2019). 2019 Refinement to the 2006 IPCC Guidelines for National Greenhouse Gas Inventories. Volume 4 (AFOLU), Chapter 2. IPCC, Switzerland. ipcc.ch
  • 21.Verra (2024). VM0042 Methodology for Improved Agricultural Land Management, v2.2. Verified Carbon Standard. Verra, Washington DC. verra.org
  • 22.Verra (2025). VMD0053 Model Calibration, Validation and Uncertainty Guidance for Biogeochemical Modeling for Agricultural Land Management Projects, v2.1. VCS Module, 25 March 2025. Verra, Washington DC. verra.org

Share this article

LinkedInEmail