Abstract:
Objective The aim was to address the challenge of simultaneously achieving high prediction accuracy and mechanistic interpretability in spatial prediction for soil organic matter (SOM). This study investigated the applicability of the Geographical Optimal Similarity (GOS) model and evaluated its performance across different land-use types.
Method Taking Zhaodong, Heilongjiang Province as the study area, 1,780 topsoil samples were collected. Fifteen environmental covariates, including topography, climate, vegetation, and land-use types, were selected and incorporated into the modeling process after dimensionality reduction using principal component analysis (PCA). The GOS model was introduced for SOM spatial prediction, and the sensitivity analysis was conducted by applying different sampling density gradients (5% – 90%). The performance of GOS was systematically compared with that of Random Forest (RF), Support Vector Machine (SVM), Geographically Weighted Regression (GWR), and kriging methods (Ordinary Kriging, OK; Co-kriging, CK) in terms of prediction accuracy, computational efficiency and spatial uncertainty.
Result ① Under the full-sample condition, the predictive performance of the GOS model was comparable to that of the RF model (R2 = 0.88), and both significantly outperformed GWR (R2 = 0.82) and OK (R2 = 0.79). ② The GOS model required only about one-fifth of the computational time of RF. In contrast to the black-box characteristics of machine learning models, GOS provided geographical interpretability through similarity-weighted predictions and produced spatial uncertainty maps based on sample attribute dispersion, highlighting northern uplands and environmental transition zones as low-confidence areas. ③ In specific habitat contexts, the GOS model showed a slight negative bias (−0.1 g kg−1) in dryland (maize) areas, whereas in paddy field (rice) regions, the incorporation of land-use constraints resulted in a positive bias of + 0.4 g kg−1, effectively alleviating the underestimation of high SOM values commonly observed in conventional approaches. ④ The sensitivity analysis indicated that RF exhibited greater robustness under sparse sampling conditions (N < 300), whereas the prediction accuracy of GOS rapidly converged to a level comparable to RF when the sample size was sufficient (N > 600).
Conclusion Under dense sampling conditions, the GOS model achieves predictive accuracy comparable to advanced machine learning methods, while offering advantages of computational efficiency, mechanistic transparency, and the ability to quantify pointwise uncertainty. The method is well suited for high-resolution digital soil mapping in black soil areas with large sample sizes and can inform regional sampling optimization and targeted management of black soils.