Correction of Diurnal Errors in Dielectric-Based In-Situ Soil Moisture Measurements via a Stacked LSTM Framework
Abstract. Soil moisture (SM) is a fundamental variable in land-atmosphere interactions, yet in-situ measurements from dielectric-based sensors often suffer from systematic temperature sensitivity known as the Maxwell-Wagner polarization effect. This sensitivity induces spurious daytime peaks that contradict the physical reality of evapotranspiration-driven daytime dry-down. In this study, we propose a data-driven approach using stacked Long Short-Term Memory (LSTM) networks to correct these temperature-induced diurnal errors in the International Soil Moisture Network (ISMN) by integrating in-situ observations with physically constrained diurnal patterns from ERA5-Land and MERRA-2 reanalysis datasets.
The LSTM-based correction (ISMNLSTM) effectively reverses the spurious positive correlations between diurnal cycles of SM and soil temperature to physically consistent negative correlations, demonstrating superior robustness across various sensor technologies compared to a recent reanalysis-informed Fourier filtering method that can struggle with regional biases. Performance evaluations reveal a significant enhancement in diurnal temporal correlation, with ISMNLSTM achieving R=0.89 against reference observations compared to the raw ISMN (R=0.14). Gradient-based sensitivity analysis confirms that the model's predictive logic is rooted in physical processes, with short-lag (1–3 hours) thermal forcing from high-frequency components and broadly distributed sensitivity to low-frequency components representing background land surface state, reflecting a multi-scale information integration. Furthermore, the diurnally adjusted SM in flux tower observations reveals physically realistic land-atmosphere coupling with negative SM-latent heat flux correlations throughout the diurnal cycle, which are fundamentally misdiagnosed as weakly positive in the original observations. Overall, the LSTM-based correction framework provides a reliable foundation for the initialization of numerical weather prediction models, the validation of sub-daily satellite SM retrievals, and advancing the understanding of sub-daily land-atmosphere interactions.
This paper addresses an important problem, "temperature-induced diurnal problem in dielectric soil-moisture measurements". The global scope, use of ISMN and flux-tower observations, and comparison with the recent Fourier-based correction method are valuable. The figures are generally clear, and the manuscript is readable.
However, the current evidence does not establish that the LSTM recovers the true soil-moisture signal. The model is trained using records selected for having a negative soil-moisture–temperature correlation, and success is then judged largely by whether the corrected data exhibit that same negative correlation. This creates a substantial circularity. Several other issues like sample leakage, weak ground-truth validation, non-causal preprocessing, insufficient baselines, and overinterpretation of the land–atmosphere analysis affect the central conclusions. My recommendation is to reject.
Major comments:
1. The training procedure described in Lines 213–218 raises a concern about circular validation, because only targets showing a sufficiently negative soil moisture-temperature correlation, R(SM, TS)<−0.34, are retained, while the negative correlations obtained after correction in Figures 3 and 4 are subsequently presented as evidence that the model has recovered physically realistic behaviour. Consequently, it is unclear whether the method is genuinely removing temperature-related artefacts, simply learning and imposing a negative diurnal pattern, or replacing local variability with a reanalysis-informed climatological cycle. Although a negative SM–TS correlation can be a useful diagnostic, it should not serve simultaneously as the training-selection criterion, the underlying physical assumption, and the main validation measure. The method therefore needs independent validation using, for example, colocated temperature-insensitive and dielectric sensors, controlled laboratory experiments, independently calibrated field observations, or synthetic tests in which thermal errors of known magnitude are added to clean soil-moisture records and the model’s ability to recover both the original signal and the injected error is evaluated.
2. The validation based on 16 electrical-reference sensor pairs within a 50 km radius requires stronger justification because stations separated by such distances may experience considerable differences in rainfall, elevation, vegetation, soil properties, land management, and local hydrological conditions. In addition, cosmic-ray sensors and point-scale dielectric probes represent very different spatial footprints. This concern is particularly relevant to Figure 5, where Cache Junction and TWDEF are approximately 40 km apart; therefore, the observed correlation may reflect common regional meteorological forcing rather than accurate correction of local soil moisture. The authors should provide detailed information for each pair, including separation distance, elevation difference, measurement depth, sensor type and footprint, soil and land-cover characteristics, temporal overlap, and rainfall agreement, together with a clear explanation of why the reference station is representative of the target station. The robustness of the results should also be examined using different pairing radii, such as 1, 5, 10, 25, and 50 km. Ideally, the primary validation should be based on genuinely colocated or near-colocated instruments, while more distant pairs should be treated only as supplementary regional comparisons.
3. The training–validation split described in Lines 220–224 needs clarification, as it is unclear whether the 70:30 division was performed by station, day, continuous time period, or individual sliding-window sample. If overlapping 18-hour windows from the same station and time period were distributed between the training and validation sets, closely related observations would occur in both subsets, leading to data leakage and overly optimistic validation results. The data should therefore be partitioned before generating the sliding windows, with model performance assessed on independent stations, networks, continuous temporal blocks or years, sensor technologies, and, where possible, climatic regions. The manuscript should also clearly distinguish the validation dataset used for hyperparameter selection from a separate final test dataset. Results from stations or periods that contributed to model training should not be presented as independent evidence of model generalization.
4. The manuscript appears to treat afternoon drying and a negative soil moisture–temperature correlation as nearly universal indicators of physically realistic behaviour. Although this inverse relationship is commonly expected during rain-free evaporative conditions, sub-daily soil moisture can also be influenced by irrigation, dew formation, hydraulic redistribution, upward capillary movement, lateral flow, post-rainfall drainage, vegetation shading, frozen-soil or snow conditions, and sensor installation effects. The timing and direction of the relationship may also vary with measurement depth, as soil moisture at depths approaching 10 cm can respond more slowly than surface temperature and evaporation. Consequently, a model trained to enforce a standard inverse diurnal pattern may suppress genuine hydrological variability rather than simply remove thermal artefacts. The authors should therefore evaluate performance separately across measurement depths, climates, vegetation types, seasons, soil textures, wetness regimes, and irrigation conditions, and demonstrate that the correction preserves real wetting, drainage, and redistribution processes.
5. The claim of real-time applicability is not supported by the current methodology. The preprocessing relies on a 24-hour centred moving average, which requires observations from subsequent hours and therefore cannot be applied causally at the time a measurement is received. In addition, ERA5-Land and MERRA-2 are used as model inputs, although these reanalysis products may not be available with the latency required for operational correction. This also appears inconsistent with the statement in Lines 108–110 that inference can be performed without further auxiliary reanalysis data. Moreover, the Physics Constraint Layer requires the low-frequency component derived from ISMN observations. The claims of “instantaneous correction,” real-time monitoring, and operational deployment in Lines 450–458 should therefore be removed or substantially revised unless the authors implement and evaluate a causal trailing filter, use operationally available forcing data, clearly state all data-latency assumptions, and demonstrate the method through prospective streaming experiments. In its present form, the framework should be described as an offline retrospective correction method.
6. The Physics Constraint Layer appears to enforce only the non-negativity of soil moisture and does not incorporate water conservation, rainfall response, evapotranspiration limits, drainage behaviour, soil porosity, temporal smoothness, or consistency between soil-moisture changes and surface fluxes. Describing it as a physics-constraint layer therefore overstates its physical content; “non-negativity output layer” would be more accurate. The explanation that negative reconstructed values provide a corrective gradient to the network also requires reconsideration because ReLU has a zero derivative for negative inputs, meaning that the clipping operation itself does not transmit a useful gradient through this branch, even if the resulting loss remains nonzero. The authors should consider a differentiable physical penalty or a bounded parameterization, such as softplus or a scaled sigmoid function. The latter could constrain predictions between zero and a physically meaningful upper limit based on soil porosity or the plausible range of volumetric water content.
7. The evaluation relies too heavily on correlation, which measures similarity in temporal phase but is insensitive to systematic bias and differences in signal magnitude. A corrected series may therefore show a high correlation with the reference even when its diurnal amplitude is substantially incorrect, particularly here where the model could learn a smooth and recurring diurnal pattern. This concern is visible in Figure 6, where the LSTM appears to differ from the reference around the afternoon minimum despite the reported correlation of R=0.89. The assessment should therefore be expanded to include RMSE, MAE, unbiased RMSE, mean bias, variance ratio, diurnal amplitude error, timing errors in the daily maximum and minimum, dry-down slope, water-balance consistency, spectral coherence and phase, and Kling–Gupta efficiency or another suitable composite metric. The authors should also report the fraction of corrected values outside physically plausible limits and demonstrate that genuine rainfall and irrigation responses are preserved. Finally, uncertainty should be quantified using confidence intervals and station-level block-bootstrap tests, because a pooled or median correlation alone is insufficient to establish statistically robust improvement.
8. The proposed stacked LSTM contains three recurrent layers with 256 units each, followed by dense layers, but the manuscript does not demonstrate that such a complex architecture is necessary or provides a meaningful advantage over simpler approaches. The comparison is largely limited to the authors’ previous Fourier-based method, whereas additional benchmarks should include hourly climatology, linear or ridge regression, a generalized additive model, random forest or gradient boosting, a shallow multilayer perceptron, a smaller or single-layer LSTM, direct use of the reanalysis high-frequency component, and a conventional temperature-regression correction. A systematic ablation analysis is also needed to determine the contribution of ERA5-Land, MERRA-2, ISMN temperature, low- and high-frequency features, latent heat flux, the negative-correlation training filter, and the non-negativity layer. Each component should be removed separately and the resulting change in independent test performance reported. In particular, the claim that ERA5-Land and MERRA-2 provide complementary information cannot be established from gradient magnitudes alone, because these values indicate model sensitivity rather than the unique predictive contribution of each dataset.
9. The gradient-based saliency analysis should be interpreted more cautiously because it measures the model’s local mathematical sensitivity to each input rather than establishing physical causality or a unique physical contribution. The resulting gradients can be influenced by input scaling, saturation of the LSTM gates, correlations among predictors, model parameterization, and the distribution of samples used for testing. This is particularly important because soil moisture, soil temperature, and latent heat flux from ERA5-Land are strongly interdependent, making it difficult to isolate the contribution of any single variable. Therefore, the statement that the saliency analysis “confirms” the model’s physical logic should be softened. Stronger evidence could be obtained through permutation importance, systematic feature-ablation experiments, integrated gradients, repeated training with multiple random seeds, and uncertainty estimates across stations. The authors should also resolve the inconsistency between the lag definition τ∈[0,17] in the text and the 18-to-1-hour labels shown in Figure 7.
10. The rainfall-screening procedure requires stronger justification because rainy days at ISMN stations are inferred from a daily soil-moisture increase exceeding 0.5 standard deviations, whereas rainfall in the reanalysis products is identified using a 0.1 mm threshold. These criteria are not physically equivalent and may select different events. A soil-moisture-based criterion may miss rainfall when infiltration is offset by evaporation or drainage, incorrectly identify irrigation or sensor artefacts as rainfall, and depend strongly on the local variability of each station. It may also preferentially retain the regular diurnal patterns the model is intended to learn. The collocated-precipitation sensitivity analysis mentioned from the authors’ previous study should therefore be reproduced or, at minimum, summarized in this manuscript. The robustness of the results should also be evaluated across several soil-moisture-change and precipitation thresholds, with the number of retained stations and events reported for each case.
11. Although the application is potentially useful, the methodological novelty appears limited because the proposed framework mainly combines a standard stacked LSTM with frequency-separated inputs and non-negativity clipping. It also closely follows the problem formulation and reanalysis-informed correction strategy introduced by Han et al. (2026), which was developed by members of the same research team. The principal methodological change is therefore the replacement of the earlier Fourier-based adjustment with a learned sequence model. This could still represent a publishable contribution, but only if the authors demonstrate clear improvement against independent reference observations, reliable generalization to unseen networks and sensor technologies, meaningful advantages over simpler machine-learning approaches, preservation of genuine hydrological events, and practical operational feasibility. At present, these aspects have not been sufficiently demonstrated; therefore, the contribution remains largely incremental, and the available validation does not support the broader claims made in the manuscript.
12. The manuscript does not provide sufficient information to assess the reproducibility of the proposed framework. The authors should report the complete station list and selection workflow, including the number of stations and samples used for training, validation, and independent testing, together with sensor types, observation periods, and the temporal and spatial matching procedures. They should also explain how time zones and local solar time were handled, how missing observations and edge effects from the centred moving average were treated, whether normalization was global, station-specific, or feature-specific, and whether all scalers were fitted exclusively using the training data. Further details are required on the model parameter count, hyperparameter-selection procedure, random seeds, variability across independent runs, software and dependency versions, early-stopping epochs, and computational cost. To enable genuine reproduction, the authors should make available the exact data partitions, station metadata, fitted scalers, preprocessing parameters, trained model weights, and resulting corrected dataset.
13. This paper uses point-scale ISMN observations with coarse-resolution ERA5-Land (~9 km) and MERRA-2 (~0.5°) products datasets. Since these datasets represent different spatial resolutions, reanalysis inputs may not accurately reflect soil and vegetation conditions. Reanalysis ERA5 and SM, TS, and LH used as physically constrained inputs inherit their own land-surface model biases and may not represent local soil and vegetation heterogeneities at all sites. This is explicitly acknowledged for FDR sensors poor FFT performance, but the same limitation could equally affect the LSTM output. While this limitation is discussed for the FFT approach, its potential impact on the LSTM model is not examined. It would strengthen the manuscript to evaluate whether model performance degrades at sites where reanalysis products poorly represent local soil moisture.
Minnor comments
Abstract: “significant enhancement” is used without a statistical significance test. Use “substantial” unless significance has been formally assessed.
Lines 47–51: The discussion of dawn maxima and hydraulic lift is too universal. Hydraulic lift is ecosystem- and condition-dependent.
Lines 100–110: The claim that auxiliary reanalysis is not needed at inference contradicts the listed ERA5-Land and MERRA-2 inputs.
Lines 204–207: “post-precipitation dry drifting” appears to be an error; probably “dry-down.”
Lines 209–212: Clarify how the 14 features are counted. They appear to be LF and HF versions of seven underlying variables.
Lines 215–217: The p<0.1 interpretation for a 24-hour correlation likely assumes independent samples, which hourly data do not satisfy because of autocorrelation.
Lines 220–224: State explicitly whether the split is by station, day, or sliding window.
Lines 240–252: Revise the claims concerning ReLU gradient propagation and physical penalization.
Figure 4: The station counts are shown only for panel (a). Provide counts for every method and clarify whether these stations contributed to training.
Figure 5: A short, favourable five-day example may be cherry-picked. State the selection criterion and provide multiple representative successes and failures.
Figure 6: Explain exactly how the correlations were calculated-composite 24-hour climatologies, all hourly values, or anomalies across overlapping dates.
Figure 7: Align the plotted lead times with the mathematical lag definition and provide uncertainty across stations or model seeds.
Section 4.3: “Realization in land-atmosphere interactions” would be clearer as “Implications for diagnosed land–atmosphere coupling.”
Conclusions: Claims regarding numerical weather prediction initialization, drought early warning, precision agriculture, satellite validation, and instantaneous correction are speculative because none of these applications is tested.