基于地理最优相似性的黑土区土壤有机质空间预测与评价

Spatial Prediction and Evaluation of Soil Organic Matter in Black Soil Region Using the Geographically Optimal Similarity Approach

  • 摘要:
    目的 针对土壤有机质(SOM)空间预测中模型精度与机理可解释性难以兼顾的问题,探究地理最优相似性(GOS)模型的适用性及其在不同土地利用类型下的表现。
    方法 以黑龙江省肇东市为研究区,采集1780个表层土壤样点,选取地形、气候、植被及土地利用类型等15个环境协变量,经主成分分析(PCA)降维后参与建模。引入GOS模型进行SOM空间预测,并设置不同采样密度梯度(5% ~ 90%)进行敏感性分析,系统对比其与随机森林(RF)、支持向量机(SVM)、地理加权回归(GWR)及克里格插值(OK、CK)在预测精度、计算效率及空间不确定性方面的差异。
    结果 ①在全样本条件下,GOS与RF模型的预测精度相当(R2 = 0.88),均显著优于GWR(R2 = 0.82)与OK(R2 = 0.79)。②GOS模型计算耗时仅为RF的约20%;与机器学习的“黑箱”特征不同,GOS通过相似度权重机制清晰追溯预测结果的参考样本来源,并基于样点属性离散度生成空间不确定性分布图,识别出北部岗地及环境过渡带为低置信度区域。③在特定生境下,GOS模型在旱地(玉米)的预测偏差为 −0.1 g kg−1;在水田(水稻)区域,得益于土地利用变量的约束,预测偏差为 +0.4 g kg−1,有效缓解了传统方法在高值区的低估倾向。④敏感性分析表明,RF在稀疏样本(N < 300)下鲁棒性较强,而GOS在样本充足(N > 600)时精度迅速收敛至与RF持平。
    结论 GOS模型在密集采样条件下具备与先进机器学习方法同等的预测精度,且兼具计算高效、机理透明及可量化逐点不确定性的优势。该方法适用于大样本量的黑土区高分辨率数字土壤制图,可为区域采样布局优化及黑土差异化管理提供科学依据。

     

    Abstract:
    Objective The aim was to address the challenge of simultaneously achieving high prediction accuracy and mechanistic interpretability in spatial prediction for soil organic matter (SOM). This study investigated the applicability of the Geographical Optimal Similarity (GOS) model and evaluated its performance across different land-use types.
    Method Taking Zhaodong, Heilongjiang Province as the study area, 1,780 topsoil samples were collected. Fifteen environmental covariates, including topography, climate, vegetation, and land-use types, were selected and incorporated into the modeling process after dimensionality reduction using principal component analysis (PCA). The GOS model was introduced for SOM spatial prediction, and the sensitivity analysis was conducted by applying different sampling density gradients (5% – 90%). The performance of GOS was systematically compared with that of Random Forest (RF), Support Vector Machine (SVM), Geographically Weighted Regression (GWR), and kriging methods (Ordinary Kriging, OK; Co-kriging, CK) in terms of prediction accuracy, computational efficiency and spatial uncertainty.
    Result ① Under the full-sample condition, the predictive performance of the GOS model was comparable to that of the RF model (R2 = 0.88), and both significantly outperformed GWR (R2 = 0.82) and OK (R2 = 0.79). ② The GOS model required only about one-fifth of the computational time of RF. In contrast to the black-box characteristics of machine learning models, GOS provided geographical interpretability through similarity-weighted predictions and produced spatial uncertainty maps based on sample attribute dispersion, highlighting northern uplands and environmental transition zones as low-confidence areas. ③ In specific habitat contexts, the GOS model showed a slight negative bias (−0.1 g kg−1) in dryland (maize) areas, whereas in paddy field (rice) regions, the incorporation of land-use constraints resulted in a positive bias of + 0.4 g kg−1, effectively alleviating the underestimation of high SOM values commonly observed in conventional approaches. ④ The sensitivity analysis indicated that RF exhibited greater robustness under sparse sampling conditions (N < 300), whereas the prediction accuracy of GOS rapidly converged to a level comparable to RF when the sample size was sufficient (N > 600).
    Conclusion Under dense sampling conditions, the GOS model achieves predictive accuracy comparable to advanced machine learning methods, while offering advantages of computational efficiency, mechanistic transparency, and the ability to quantify pointwise uncertainty. The method is well suited for high-resolution digital soil mapping in black soil areas with large sample sizes and can inform regional sampling optimization and targeted management of black soils.

     

/

返回文章
返回