A laboratory does not report a soil carbon stock. It reports a carbon concentration, in grams per kilogram, on a sample somebody else collected. Turning that into tonnes per hectare takes three more numbers: how dense the soil is, how deep the sample went, and what fraction of the volume was stones rather than soil. Multiply the four, and the errors of all four come with them.
This is where most arguments about soil carbon numbers actually live. Our benchmark piece, how much soil carbon a hectare can gain per year, showed that a claimed rate is only interpretable against the stock it changes. This one asks the prior question. Given how the stock is assembled, how well can it be known at all, and what is the smallest change a programme could honestly claim to have seen?
The answer, before the reasoning
Spatial heterogeneity
Usually the largest term, often over half the variance. It is a property of the field, not of your instrument, and no amount of laboratory precision touches it.
Bulk density
Rarely the biggest random term, but the usual home of systematic bias. Get it wrong the same way twice and the error does not average out.
Coarse fragments
Ignored more often than any other term, and capable of inflating a stock by a third on stony ground.
Depth and soil mass
Fixed-depth and equivalent-soil-mass accounting of the same field can differ by more than a hundred percent once bulk density shifts.
Together these set the floor, the minimum detectable difference. On an ordinary arable field sampled at thirty points, it sits high enough that a decade of good practice can pass underneath it without ever becoming statistically visible.
A stock is a product of four measurements, so four errors compound
The calculation itself is simple. Carbon concentration times bulk density times depth, times the fraction of the volume that is fine earth rather than stone. What is less obvious is how the uncertainties travel through it. For a product of independent terms, the relative variances add, so the coefficient of variation of the stock is the square root of the sum of the squared coefficients of variation of its parts.
CV(stock)² ≈ CV(C)² + CV(BD)² + CV(depth)² + CV(coarse)²
That squaring is the practical point. A term contributes in proportion to the square of its own variability, so the budget is dominated by its largest entry and almost indifferent to its smallest. Halving a term that carries 60 percent of the variance is worth far more than perfecting one that carries 5 percent, and a great deal of monitoring effort goes into the second kind of improvement. The formal treatment using error propagation and Monte Carlo cross-checks converges on the same picture2.
Spatial heterogeneity usually carries most of the budget
Soil carbon varies over metres. Two cores taken ten paces apart in the same field can differ by more than the change a decade of management will produce, and that variability is a property of the ground rather than of the instrument. An intensive campaign across one rangeland site and seven cropland sites in California, more than a thousand samples in total, found spatial heterogeneity to be the primary driver of uncertainty, with dry combustion assays contributing relatively little3. A field-scale decomposition on cultivated soil put natural variation at 49 to 68 percent of the total variance, sample preparation at 11 to 26 percent, and the analytical step itself at only 5 to 9 percent11.
The same study set that quantified heterogeneity also delivered the uncomfortable corollary: in heterogeneous agricultural landscapes, the sample sizes people actually use, ten to thirty cores, cannot reliably detect the modest changes that a few years of ordinary management produce3. That is not a criticism of anyone's diligence. It is arithmetic, and it is the reason this article exists.
Land use shifts the picture. Repeated inventories at twelve European sites with a hundred cores each found croplands, being relatively uniform, had the lowest minimum detectable differences, while grassland and forest sites were roughly twice as hard to read1. Scale shifts it too: the coefficient of variation of a stock estimate rises from around 5 percent at plot scale to 35 percent at landscape scale, and runs higher for grassland than for cropland2.
Bulk density is where the systematic bias hides
Bulk density is often a modest contributor to random error. In one Australian survey, carbon percentage carried 84 to 99 percent of the uncertainty in stocks while bulk density carried under 5 percent7. That statistic is frequently quoted as a reason to relax about bulk density, and it is the wrong lesson, because the danger from bulk density is not random. It is systematic.
Management changes bulk density. Reduced tillage loosens the surface, so the same 30 cm core holds less soil after the practice than before. Sampling to a fixed depth therefore compares unequal masses, and the resulting bias runs the same direction every time. A realistic shift from 1.5 to 1.1 grams per cubic centimetre underestimates the stock change by 17 percent5. Using a pedotransfer function to estimate bulk density rather than measuring it makes matters worse in a subtler way: it substantially underestimates the total variance of the stock unless the function's own error is carried through, which inventories rarely do1.
Method choice alone can move a national number. Across nearly 2,900 plots in the United States forest inventory, three common ways of calculating bulk density produced mean stocks differing by up to 13 tonnes per hectare, driven mostly by inconsistent treatment of rocks and roots, which overstated carbon by 32 percent of the mean12. At the other end of the scale, average bulk density values by soil type serve national inventories adequately but not field-scale carbon schemes, where site-specific measurement is needed13.
Coarse fragments are the term most often left out
Stones hold no carbon, but they occupy volume. Failing to subtract them inflates a stock in direct proportion to how stony the ground is, and the mistake is widespread: a review of the way bulk density and rock fragment content are handled found stocks systematically overestimated through misuse of these two parameters14. That paper drew a published comment disputing parts of the analysis15, and the exchange is worth reading precisely because the disagreement is about how to correct, not whether to.
Where stones are abundant the term stops being a correction and becomes a controlling variable. Work on the single-layer equivalent soil mass method across 393 Swiss fields found the coarse fraction to be a major source of error whenever it exceeds about 10 percent of the layer volume, though handling it properly did not degrade detectability4. Landscape-scale error budgets likewise put rock fragment content alongside carbon concentration as one of the two predominant sources of uncertainty2.
Fixed depth and equivalent soil mass can differ by more than a hundred percent
Of all the choices above, this one has the largest single effect on a reported change. Comparing the two accounting methods over twenty years at two long-term American experiments, estimates of stock change differed by over 100 percent, and the difference tracked bulk density change in the surface soil almost exactly6. The authors conclude that equivalent soil mass accounting and sampling to 60 cm should be treated as best practice for annual row-crop systems, and that the cost of doing so belongs in the price of a credit.
The direction of the bias is not fixed. In boreo-temperate tillage trials, ignoring soil mass overstated the apparent benefit of reduced tillage by 15 to 47 percent16, while in tropical land-use change it understated the effect by 28 percent17. Fixed-depth accounting is not a conservative simplification. It is an error whose sign you cannot predict without doing the correction you were trying to avoid.
The practical objection to equivalent soil mass has always been cost, and that objection is weakening. The Swiss work found that a simplified single-layer method, with layer mass determined by gouge auger, produced a minimum detectable change corresponding to under ten years, against several decades for fixed-depth sampling with assumed bulk density4. Doing it properly turned out to be the cheaper route to an answer, because the alternative never produces one.
The floor has a name: minimum detectable difference
Everything above converges on one number, and the number has a name. The minimum detectable difference, or MDD, is the smallest real change a monitoring design can reliably separate from its own sampling noise. The convention is 95 percent confidence and 80 percent power: a change of exactly the MDD gets caught four times out of five, while an unchanged soil raises a false alarm one time in twenty. The floor is a property of the design alone, computable before the practice begins, and a true change smaller than it does not show up as a smaller signal. It usually shows up as nothing, and the nothing proves nothing: a result below the MDD is not evidence that the soil stood still, only that this design could not have seen it move.
Cores needed to see a change Δ: n = 2kσ² / Δ²
Each symbol is a lever, and they are not equal. The standard deviation σ is the output of the whole error budget above, and it enters linearly: halve it and the floor halves. The sample count sits under a square root, so buying that same halving with cores costs four times as many. The constant k is the price of rigour, set the moment you choose your confidence and power; relaxing it lowers the floor on paper and nowhere else. And the factor of two is there because a change is two campaigns, each contributing its own noise, which is why resampling the same physical points and differencing the pairs removes so much variance from the comparison before the formula ever runs. Run the arithmetic the other way, fixing the change you must see and solving for n, and it sizes a campaign instead; we work through that direction in the sample-size piece.
The reason to care about the formula is not the formula. It is that it can be evaluated at a desk, before any money is spent, from numbers you can borrow: a variance estimate from a regional survey or a pilot, the sampling intensity in the budget, the interval in the contract. Take the ordinary arable field this article has used throughout, a 55-tonne stock with a standard deviation of 11 tonnes per hectare. Thirty cores per campaign put the floor near 8 tonnes of carbon per hectare, roughly 15 percent of the stock. A practice change accruing 0.3 tonnes per hectare per year, a perfectly respectable rate, accumulates 1.5 tonnes over a five-year contract: a fifth of the floor. That campaign cannot succeed, and the arithmetic said so before the field crew was booked. Seeing the five-year signal would take more than eight hundred cores; seeing it with thirty cores would take nearly three decades.
This is the calculation that keeps a programme from paying full price for a number it can never use, because a sampling campaign costs the same whether or not it is capable of detecting anything. When the numbers refuse, the honest responses are to lengthen the interval, add cores, shrink the variance by stratifying and pairing, or decline to promise a measured change on that timescale at all. Each of these is cheaper than the fifth response, which is running the campaign anyway and finding out at verification.
The tool below runs the arithmetic live. Take the standard deviation the budget tool produced, choose a sampling intensity and an interval, and read off the floor, then see whether the rate you had in mind ever clears it.
Published estimates of how long this takes are consistent and sobering. Thirteen sites across the north central United States gave minimum detectable differences spanning 6 to 31 percent of the stock, with the time to detect a change under no-till running from 11 to 71 years depending on site and replication8. Regional monitoring in southern Belgium could detect a 20 percent stock change within 11 years only where rates ran as high as 1 tonne of carbon per hectare per year2. European inventories with a hundred cores per site put detection at 2 to 15 years, but only for changes in the top 10 cm of stone-poor soils1. Ten-year resampling intervals are the usual compromise for good reason18.
Cheaper methods buy coverage by spending precision
If heterogeneity dominates the budget, the obvious response is more samples, and the obvious way to afford them is a cheaper measurement. Infrared spectroscopy is the usual candidate, and the trade is real in both directions. A handheld near-infrared analyser used across three Massachusetts farms gave less precise predictions than elemental analysis but still resolved the same statistical differences in stock between farms, at a large cost saving9. Used well, spending precision to buy coverage is a sound bet when the thing you are fighting is spatial variance.
Used badly it is not. Across 151 plots in western Kenya, a portable scanner correlated only weakly with laboratory values, two laboratories agreed only moderately with a systematic offset between them, and repeat measurements within one laboratory were poorly repeatable. Simulated forward into a payment scheme, an unbiased error caused a 22 percent deviation in estimated stock change for a group of fifty farmers, while a systematic bias of just 1 gram per kilogram produced an 18 percent error that persisted no matter how large the group grew10.
That contrast is the whole lesson of this article in miniature. Random error is diluted by scale; bias is not. A cheap method with a known, stable, calibrated offset can be an excellent instrument. A cheap method with an uncharacterised offset is a liability that grows with the size of the programme resting on it.
What actually moves the floor
The design choices that matter are not evenly weighted. A handful of them do nearly all the work.
Attack the largest variance term first. On most agricultural fields that is spatial heterogeneity, which means stratifying the field or resampling fixed locations, not buying a better analyser.
Resample the same points. Pairing measurements to locations removes the between-point variance from the comparison entirely, which is the single cheapest improvement available.
Measure bulk density rather than estimating it, and if you must use a pedotransfer function, carry its error through the propagation.
Correct for coarse fragments explicitly, by volume. Above roughly a tenth of the layer volume this stops being a refinement.
Use equivalent soil mass, and sample deeper than 30 cm. Both cost money; both prevent an error whose direction you cannot otherwise know.
Keep the method identical between campaigns. A consistent imperfect method beats an improved one, because a change in method is indistinguishable from a change in the soil.
Report the floor next to the result. A stock change quoted without the minimum detectable difference cannot be assessed by the person reading it.
None of this argues that soil carbon should not be measured. It argues that a measurement programme should be designed backwards from the change it needs to detect, rather than forwards from the budget it happens to have. The question worth asking at the start is not how many samples can we afford, but what is the smallest change we would need to see, and what would it take to see it. Sometimes the honest answer is that this field, at this intensity, will not produce a defensible number for a decade. Knowing that early is worth a great deal more than discovering it at verification19.
Poin-poin utama
A soil carbon stock is a product of four measurements, and their relative variances add as squares. The budget is dominated by its largest term and nearly indifferent to its smallest.
Spatial heterogeneity is usually that largest term, contributing around half to two-thirds of the variance on cultivated land, against 5 to 9 percent for the analytical step.
Sample sizes of ten to thirty cores, which is what most programmes use, cannot reliably detect the changes a few years of ordinary management produce.
Bulk density is rarely the biggest random term but is the usual home of systematic bias, because management changes it in a consistent direction. A shift from 1.5 to 1.1 g/cm³ understates stock change by 17 percent.
Fixed-depth and equivalent-soil-mass accounting of the same experiment can differ by over 100 percent, and the sign of the error is not predictable without doing the correction.
Coarse fragments stop being a minor correction above about a tenth of the layer volume, and inconsistent treatment of rocks and roots overstated US forest soil carbon by 32 percent.
Cheap methods trade precision for coverage, which is a good trade against random error and a bad one against bias. A 1 g/kg systematic offset produced an 18 percent error that no amount of scale removed.
Design the programme backwards from the change it must detect. The minimum detectable difference can be computed at a desk before any core is paid for; report it beside the result, so a reader can tell signal from silence.
References
- 1.Schrumpf, M., Schulze, E.D., Kaiser, K., Schumacher, J. (2011). How accurately can soil organic carbon stocks and stock changes be quantified by soil inventories? Biogeosciences, 8(5), 1193–1212. doi:10.5194/bg-8-1193-2011
- 2.Goidts, E., Van Wesemael, B., Crucifix, M. (2009). Magnitude and sources of uncertainties in soil organic carbon (SOC) stock assessments at various scales. European Journal of Soil Science, 60(5), 723–739. doi:10.1111/j.1365-2389.2009.01157.x
- 3.Stanley, P., Spertus, J., Chiartas, J., Stark, P.B., Bowles, T. (2023). Valid inferences about soil carbon in heterogeneous landscapes. Geoderma, 430, 116323. doi:10.1016/j.geoderma.2022.116323
- 4.Boivin, P., Lemaître, C., Clark, J. et al. (2025). The single-layer equivalent soil mass method for the evaluation of soil organic carbon stocks: Sources of errors, simplification, and associated detectable change. Geoderma, 456, 117279. doi:10.1016/j.geoderma.2025.117279
- 5.Fowler, A.F., Basso, B., Millar, N., Brinton, W.F. (2023). A simple soil mass correction for a more accurate determination of soil carbon stock changes. Scientific Reports, 13, 2242. doi:10.1038/s41598-023-29289-2
- 6.Raffeld, A.M., Bradford, M.A., Jackson, R.D. et al. (2024). The importance of accounting method and sampling depth to estimate changes in soil carbon stocks. Carbon Balance and Management, 19, 2. doi:10.1186/s13021-024-00249-1
- 7.Holmes, K.W., Wherrett, A., Keating, A., Murphy, D.V. (2012). Meeting bulk density sampling requirements efficiently to estimate soil carbon stocks. Soil Research, 49(8), 680–695. doi:10.1071/SR11161
- 8.Necpálová, M., Anex, R.P., Kravchenko, A.N. et al. (2014). What does it take to detect a change in soil carbon stock? A regional comparison of minimum detectable difference and experiment duration in the north central United States. Journal of Soil and Water Conservation, 69(6), 517–531. doi:10.2489/jswc.69.6.517
- 9.Sanderman, J., Partida, D., Safanelli, J.L. et al. (2025). Application of a handheld near infrared spectrophotometer to farm-scale soil carbon monitoring. European Journal of Soil Science, 76(2), e70053. doi:10.1111/ejss.70053
- 10.Schilling, F., Beesigamukama, D., Tanga, C.M. et al. (2026). Reliability of approaches for measuring soil organic carbon and implications for results-based payments for smallholder carbon farming. Journal of Environmental Management, 398, 128562. doi:10.1016/j.jenvman.2026.128562
- 11.Samsonova, V.P., Meshalkina, J.L., Dobrovolskaya, N.G. et al. (2023). Investigation of uncertainty in organic carbon stock estimates on a field scale. Eurasian Soil Science, 56, 1765–1775. doi:10.1134/S106422932360183X
- 12.Lang, A.K., Pastore, M.A., Walters, B.F. et al. (2025). Bulk density calculation methods systematically alter estimates of soil organic carbon stocks in United States forests. Biogeochemistry, 168, 47. doi:10.1007/s10533-025-01235-6
- 13.Harbo, L.S., Olesen, J.E., Liang, Z. et al. (2022). Estimating organic carbon stocks of mineral soils in Denmark: Impact of bulk density and content of rock fragments. Geoderma Regional, 30, e00560. doi:10.1016/j.geodrs.2022.e00560
- 14.Poeplau, C., Vos, C., Don, A. (2017). Soil organic carbon stocks are systematically overestimated by misuse of the parameters bulk density and rock fragment content. SOIL, 3(1), 61–66. doi:10.5194/soil-3-61-2017
- 15.Hobley, E., Murphy, B., Simmons, A. (2018). Comment on “Soil organic stocks are systematically overestimated by misuse of the parameters bulk density and rock fragment content”. SOIL, 4(2), 169–171. doi:10.5194/soil-4-169-2018
- 16.Meurer, K.H.E., Haddaway, N.R., Bolinder, M.A., Kätterer, T. (2018). Tillage intensity affects total SOC stocks in boreo-temperate regions only in the topsoil: A systematic review using an ESM approach. Earth-Science Reviews, 177, 613–622. doi:10.1016/j.earscirev.2017.12.015
- 17.Don, A., Schumacher, J., Freibauer, A. (2011). Impact of tropical land-use change on soil organic carbon stocks: A meta-analysis. Global Change Biology, 17(4), 1658–1670. doi:10.1111/j.1365-2486.2010.02336.x
- 18.Smith, P., Soussana, J.-F., Angers, D., Schipper, L., Chenu, C. et al. (2020). How to measure, report and verify soil carbon change to realize the potential of soil carbon sequestration for atmospheric greenhouse gas removal. Global Change Biology, 26(1), 219–241. doi:10.1111/gcb.14815
- 19.VandenBygaart, A.J. (2006). Monitoring soil organic carbon stock changes in agricultural landscapes: Issues and a proposed approach. Canadian Journal of Soil Science, 86(3), 451–463. doi:10.4141/S05-105