Journal Paper Digests 2026 #21
- An energy-based model of soil fragmentation by tillage
- Stacking machine learning models to estimate soil moisture using gravimetric and in-field sensed data
- Do data from the RISMA soil monitoring network support agricultural modeling requirements?
- Combinatorial group testing for efficient scaling across biological applications
- Putting the soil health principles to the test in Iowa
- A Copula-Based Regression Method to Predict Sediment Concentration and Load in Rivers Using Streamflow Data
A Copula-Based Regression Method to Predict Sediment Concentration and Load in Rivers Using Streamflow Data
Pollution and deposition of sediments in rivers and streams are critical environmental, ecological, navigational, and recreational concerns. While several well-known watershed models such as SWAT (Soil and Water Assessment Tool) and HSPF (Hydrological Simulation Program-FORTRAN) are applied to predict sediment concentrations and loads in rivers and streams, their application requires the acquisition of a large amount of climatic, GIS, land use, and watershed input data, which is very time-consuming and cost-prohibitive. This study developed a copula-based regression method to predict sediment concentrations and loads in rivers using streamflow data, which is difficult to achieve using traditional methods. Predicted sediment concentrations from the method were verified and validated by field measurements from three US Geological Survey (USGS) gauge stations across the US, as well as by statistical metrics. Prerequisites, advantages, uncertainties, and limitations of the method were presented and discussed. Results revealed that sediment loads were not proportional to watershed drainage area, indicating that land use and anthropogenic activities also play an important role in sediment load. Overall, no significant increasing or decreasing trends of annual sediment loads were found at study sites in New York and Florida over the 11-year period from 2011 to 2021. This study suggests that the copula-based regression approach, which is time-saving and cost-effective compared to the traditional watershed models, is a promising alternative for predicting sediment concentrations and loads using streamflow data when a good dependence (or correlation) exists between the sediment content and streamflow after copula transformation.
Putting the soil health principles to the test in Iowa
One of the most popular soil conservation campaigns is based on the USDA Natural Resource Conservation Service’s Soil Health Principles (NRCS-SHPs). The NRCS-SHP program identifies four principles—maximize presence of living roots, minimize disturbance, maximize soil cover, and maximize biodiversity—with the underlying assumption that the more principles one follows, the greater improvements in soil health. Despite the popularity of the NRCS-SHPs, this underlying assumption has not been rigorously tested. To do so, we used nine long-term experiments all located in central Iowa, but with varying degree of NRCS-SHP adoption, to determine if greater adoption increases three slow-changing (maximum water holding capacity, bulk density [BD], and soil organic carbon) and three dynamic (microbial biomass carbon [MBC], potentially mineralizable carbon [PMC], and permanganate oxidizable carbon [POXC]) soil health indicators. We regressed these indicators with a soil health principle score that can scale soil management based on adoption of the NRCS-SHPs. Of the slow-changing soil properties, increased adoption of NRCS-SHPs only decreased soil BD (R2 = 0.22, p = 0.024). On the other hand, increased adoption of NRCS-SHPs strongly predicted increases in both MBC and PMC and across two sampling dates (R2 > 0.23, p < 0.015); POXC, however, did not increase with greater adoption. The consistent increases in MBC and PMC with greater adoption of NRCS-SHPs supports their usefulness as sensitive indicators of positive soil health change. Our study provides scientific evidence to support the NRCS-SHPs concept, improving its usefulness as an extension campaign, and stands as a step toward evidence-based soil conservation
Combinatorial group testing for efficient scaling across biological applications
Combinatorial group testing can reduce experimental costs and turnaround time by strategically pooling samples to minimize the number of measurements needed for a given experiment. Despite broad potential utility, it remains underutilized due to its intrinsic complexity and the lack of implementation tools. Here we present PoolPy, a unified end-to-end framework and web platform to benchmark, automate, and decode combinatorial group testing strategies. PoolPy tailors pooling designs to application-specific constraints, such as time, cost, or signal dilution, across experiment types. By implementing ten different pooling algorithms, which we comprehensively benchmark in silico across >100,000 conditions, we identify key design trade-offs that define pooling applicability to specific use cases. We experimentally validate PoolPy across diverse applications, including protein-ligand interaction screening, RT-qPCR viral testing and genome-wide protein-DNA interaction profiling, achieving a 60 to 93% reduction in number of measurements needed. Overall, PoolPy provides a scalable, user-friendly ecosystem to increase throughput and reduce costs across biological applications. PoolPy is available at https://poolpy.trouillonlab.org for open use.
Do data from the RISMA soil monitoring network support agricultural modeling requirements?
The Real-Time In-Situ Soil Monitoring for Agriculture (RISMA) network provides soil water and agri-environmental data for conceptualization, calibration, and validation of remote sensing and modeling products for Canadian agriculture at 32 stations in seven watersheds in Saskatchewan, Manitoba, and Ontario. The data support work in precision agriculture, flood and drought management, and greenhouse gas management. However, project-based development has led to inconsistent data collection. This work assesses the adequacy of data collected by the RISMA network to support agricultural modeling. The work provides expert modeler review of parameter requirements for Versatile Soil Moisture Budget (VSMB), Soil, Vegetation, and Snow (SVS), Environmental Policy Integrated Climate (EPIC), Agricultural Policy Environmental eXtender (APEX)/ArcAPEX, Hydrogeosphere, HYDRUS, Denitrification-Decomposition (DNDCv.CAN), and Hydrologic Engineering Center’s Hydrologic Modeling System (HEC-HMS) and qualitatively assesses the alignment between network observations and model input requirements. All RISMA stations collect soil volumetric water content, soil temperature, soil electrical conductivity, and liquid precipitation. Select sites also collect precipitation by weight, temperature, relative humidity, wind speed and direction, solar radiation, and crop-related data such as yields, tillage, and crop species. To better support modeling, the RISMA network should standardize collection of weather, soil, snowfall, and crop production data, including planting and harvest dates, crop type, tillage, and fertilizer application. Standardizing these observations will enhance the network’s value for supporting site-specific modeling, regional upscaling, and assessing variability where local data are unavailable.
Stacking machine learning models to estimate soil moisture using gravimetric and in-field sensed data
Even though nutrient and irrigation management could be improved by accurate soil moisture monitoring, a cost-effective approach is not available for obtaining this information. The objective of this study was to use soil- and weather-derived predictors to estimate volumetric water content at multiple depths. Two multi-site data sets were used in this study. One dataset was derived from preplant soil samples collected from 19 sites located in soils with an ustic soil moisture regime between 2019 and 2022, and the second dataset was collected at sites that had udic soil moisture regime. At the second dataset, the soil sensors collected volumetric soil information near-continuously between 2021 and 2023. In the study, the performance of multiple linear regression, support vector regression, random forest, extreme gradient boosting, and a stacked feed-forward neural network meta-learner model were compared. The machine learning models for the preplant gravimetric soil moisture collected from sites that had an ustic soil moisture regime had low predictability (R2 < 0.3) and high root mean square errors. These results suggest that the models did not adequately capture the nonlinear and spatio-temporal drivers governing soil moisture variability. Machine learning models that estimated in-season soil moisture at sites located in the udic soil moisture regime had higher predictability, and the best models had R2 values as high as 0.99. The high performance of these models was attributed to the use of temporal predictors, and in general stacking provided consistent gains (up to Δ 𝑅2 =+0.114 at 10 cm; mean ̅̅̅̅̅̅̅̅ 𝑅2 =0.618 across depths) but offered limited benefit when temporal features are included. Overall, the results show that incorporating temporal context and, when needed, stacking complementary learners can produce more reliable soil moisture estimates for field-scale management.
An energy-based model of soil fragmentation by tillage
Tillage greatly modifies soil structure, yet existing approaches to modelling tillage-induced soil structural change remain largely qualitative or over-simplified. Here, we present a quantitative, energy-based reformulation of the classical tillage equation that predicts soil fragmentation from the initial soil state and the applied energy. The new model partitions tillage energy into surface creation through fragmentation, displacement of existing soil fragments, and plastic deformation of the soil. Soil moisture and mechanical properties control the energy partitioning among these processes. As fragmentation progresses, an increasing proportion of tillage energy is dissipated through the displacement of existing fragments, whereas at elevated soil water contents energy is consumed by plastic deformation. Soil fragmentation resulting in the creation of new fragment surface area scales with the remaining energy. Literature-derived soil fragmentation data from drop-shatter tests and tillage experiments were used for model parameterisation. The model was subsequently evaluated against separate, independent literature datasets not used for parameterisation, covering tillage-induced soil fragmentation across a range of soil conditions. Illustrative applications demonstrate the model’s ability to capture (i) the texture-dependent soil workability range and (ii) the diminishing effectiveness of repeated tillage operations. We outline how the model-derived fragment surface area can be linked to fragment size distributions under simplifying assumptions. Assuming spherical fragment geometry and a Weibull distribution of fragment sizes allows an estimation of the tillage-induced pore size distribution. Altogether, our physically grounded model provides a basis for predicting tillage-induced soil structural change and can be incorporated into agroecosystem models. Further model refinement would benefit from datasets that jointly quantify tillage energy input, soil water status, as well as pre- and post-tillage soil structure and hydraulic properties, enabling a more mechanistic description of the transition from brittle fragmentation to plastic deformation.