the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Soil water forecasting for Australian dryland agriculture
Abstract. Soil water availability is a critical constraint to agricultural productivity. While soil water forecasting has previously been conducted in the literature for irrigated fields and surface soil (0–10 cm), there is a limited understanding of subsurface soil water in dryland paddocks, particularly in Australia. Due to rooting depth typically extending below 10 cm, subsurface soil water forecasts in cropping systems, for example, enable decision-making relating to sowing, fertiliser use, and seasonal yield potential. This study aimed to identify the best performing methods, and the underlying variables that affect accuracy when forecasting subsurface soil water (30–100 cm). The methods appraised in this study are Random Forest, XGBoost, Multi-Layer Perceptron (MLP), Long Short-Term Memory (LSTM), and an Encoder-decoder LSTM (ENCDEC). Various in situ and remotely sensed meteorological data were used to forecast soil water up to multiple months ahead, at 54 probe sites across Australia.
We find all models can perform with relatively similar accuracy for sub-monthly forecasts. For longer lead times, MLP and XGBoost performed better for sites with uniform rainfall throughout the year, and the memory-based models, LSTM and ENCDEC, better at sites with seasonally-dominant rainfall. Furthermore, we determine an upper limit to model utility, as accuracy can become poor many months ahead, and can be replaced by a historic average as an estimate. We find our models to be 'useful' up to 4 months ahead on average, after which accuracy is too low, or a historic average outperforms a model. We find rainfall intensity and seasonality, and model choice, to be key drivers of forecast accuracy at a new site. This study highlights the capability of elementary forms of each ensemble and deep learning model to provide sufficiently accurate soil water forecasts, particularly in the absence of a rainfall forecast feature.
- Preprint
(19427 KB) - Metadata XML
- BibTeX
- EndNote
Status: final response (author comments only)
- RC1: 'Comment on egusphere-2026-2438', Anonymous Referee #1, 04 Aug 2026
-
RC2: 'Comment on egusphere-2026-2438', Anonymous Referee #2, 14 Sep 2026
The manuscript compares five machine learning methods (Random Forest, XGBoost, MLP, LSTM, and an encoder-decoder LSTM) for forecasting subsurface soil water content (30–100 cm) at 54 dryland sites across Australia, at lead times ranging from one week to 52 weeks. The authors classify sites into four "zones" based on rainfall seasonality and geography (Uniform, Winter, Summer, WA), evaluate model performance using RMSE and a scaled Kling-Gupta Efficiency (KGE's), define a practical "model utility" horizon (benchmarked against a historic climatological average), and use multiple linear regression to identify drivers of forecast error. The topic is relevant, and the dataset is a genuine strength; the paper fills a documented gap regarding subsurface, dryland, non-irrigated soil moisture forecasting. However, several methodological points need clarification before the results can be fully assessed, and the presentation would benefit from tighter quantitative support in several places. I recommend major revisions.
Major comments
Train/validation/test separation is insufficiently described (Sections 4.1, 4.3).
Hyperparameter tuning is performed via grid search to maximize KGE's at a 30-day lead time, then the selected combination is checked for generalizability across lead times, and finally models are evaluated on "the final 2 years of data." It is not clear whether this final 2-year window is disjoint from the data used during hyperparameter selection, or whether the same held-out period informs both. Please clarify explicitly whether tuning was performed with nested/rolling cross-validation on the training portion only, and state the exact boundaries of train, validation, and test splits.Unequal sample sizes across zones are not flagged where they materially affect conclusions.
Table 2 reports that 100% of Summer sites at 60–100 cm hit the RMSE benchmark, but the text later reveals this zone/depth combination consists of a single site ("the only available Summer site at 60-100"). Presenting this as a percentage in a table alongside proportions computed from many more sites (e.g., Uniform, n presumably >10) is misleading without a sample-size caveat. Please report n per zone/depth combination directly in Table 2, or add footnotes, and temper claims about the Summer zone accordingly throughout Section 5.3.1 and the Discussion.Explanatory power of the regression model (Section 5.4) is modest and underexplored.
The linear regression explains only 36–45% of RMSE variance. This is reported honestly, but given that soil moisture forecasting is known to involve non-linear interactions, it would strengthen the paper to either (a) test for interaction terms explicitly, or (b) briefly benchmark against a non-linear explanatory model to see whether substantially more variance can be explained, and to check whether the linear coefficients' signs and significance are stable.Zone definition logic is heterogeneous and should be justified more rigorously.
Three zones (Uniform, Winter, Summer) are defined by rainfall seasonality, while the fourth (WA) is defined by geography/soil texture independent of its (winter-dominated) rainfall pattern. The authors note that WA sites happen to all be winter-dominant, but this raises the question of why WA is not simply nested within "Winter." A short quantitative comparison (e.g., a table cross-tabulating clay: sand ratio and rainfall seasonality by zone) would make the rationale for treating WA as a separate axis more transparent and defensible, particularly since Section 5.4 partly attributes WA's anomalous 1-week regression coefficient to method-averaging artifacts rather than the zone definition itself.Attribution of RNN underperformance is speculative (Section 6, lines ~543–551).
The claim that LSTM/ENCDEC do not outperform non-sequential methods because of root-water interactions, irregular rainfall input, and depth is plausible but not directly tested in the current analysis. Since this is a central and somewhat counter-literature finding, it would be considerably strengthened by a targeted diagnostic rather than resting on narrative interpretation alone.Minor comments
- Eq. (1): The RMSE formula as rendered is missing the squared term in the numerator; "RSME" is also misspelled at line 280.
- Table 2 / Figure captions: Percentages in Table 2 should sum in a way that is checkable against Figures 4–5; consider adding raw counts alongside percentages.
- Figure 1 is referenced but never discussed in the text.
- Line 100: "this study focuses on understand forecasting capabilities" --> grammatical fix needed.
- Line 407: "forests" --> "forecasts."
- Reference list: The two Bureau of Meteorology "Meteorology, a" / "Meteorology, b" entries appear to be a bibliography-export artifact and should be corrected to proper author/organization citations.
Citation: https://doi.org/10.5194/egusphere-2026-2438-RC2 -
RC3: 'Comment on egusphere-2026-2438', Anonymous Referee #3, 29 Sep 2026
Soil water forecasting for Australian dryland agriculture
General comments:
- The introduction reflects that GRU architecture has shown the most promising results for soil moisture prediction, but it is not included in the tested models. Is there any particular reason for not including it?
- Mean annual rainfall is used to look at the forecast performance which is not really indicative of model accuracy for regions with seasonal rainfall.
- Site specific models will be fitted to site specific climatology and would have limited to no utility outside that specific site. It would be better to train the models with data from all the sites in each zone and compare with site specific results. Figure 6 suggests that no one model outperforms the other models when looking at the proportion of sites. Would it not be better to pool the ‘zonal’ performance via training using one model for all sites?
- Further details are needed about the model training data. What time periods are used for each variable? How is the spatial discrepancy between the site and 5km feature value dealt with?
- Nice figures!
Specific comments:
- Line 53: typo: MLPs consistently outperform conventional
- Line 59: Are MLPs or LSTMs the best performing? The paragraph above states MLPs but then in this paragraph it says LSTM.
- Line 70: Please rephrase the sentence: ‘They develop an architecture that pass their predictors and a LSTM encoded-decoded forecast of soil moisture at the preceding time steps to the target lead time into an LSTM.’
- Line 115: grammar mistake
- Line 140: Why is WA a different ‘zone’? How is different from ‘winter’ zone?
- Line149 : ‘The rainfall is calculated as the maximum between the in situ measured rainfall and the SILO modelled estimated rainfall at that site.’ How do you account for the spatial representativeness of the two different rainfall sources?
- Fig 4. Why is the RMSE shown as %VWC instead of the actual soil moisture values?
- Line 561: Rainfall is already used as an input. Is only historical rainfall used as input, not forecast rainfall?
Citation: https://doi.org/10.5194/egusphere-2026-2438-RC3
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 257 | 211 | 28 | 496 | 14 | 14 |
- HTML: 257
- PDF: 211
- XML: 28
- Total: 496
- BibTeX: 14
- EndNote: 14
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Amir et al. evaluate several ML models for soil moisture forecasting. Their results indicate that (a) DL models generally outperform ML methods such as MLP and XGBoost, and (b) how limited the models are compared to a "standard" threshold. While the application is relevant and the Australia-wide aggregation of probe records is potentially valuable, I do not see the manuscript as fit for publication. Several major concerns, in my opinion, should be addressed:
1) The methodological contribution appears limited. Benchmarking established models such as MLP and XGBoost, without introducing a novel architecture, theoretically grounded formulation, or substantive methodological advance, is insufficiently innovative for a 2026 study. In addition, the site-specific and geographically imbalanced dataset limits the generalizability of the findings and their broader contribution to data-driven soil-water forecasting. The authors should also justify the exclusion of meteorological forecasts and reanalysis products, such as ERA5 soil moisture. Incorporating these sources or systematically evaluating their added value could support more transferable models and provide a stronger benchmark for assessing the proposed approach.
2) The methodology is not clear. The authors should define one training example explicitly: forecast issue date `t`, target date `t+h`, every feature timestamp available at `t`, sequence/lookback window, target, and lead time `h`. This is essential because the paper calls the task forecasting but uses retrospective gridded and remotely sensed products. It is not clear what the forecast step is, what the issue date is, and what inputs are used. Does the LSTM use past observed or past series of meteorological forcing? If so, what is the lookback, and is this the case with other methods? If not, what is the usage of LSTM?
3) Validation, test use, and hyperparameter selection are ambiguous and may contaminate the test set. The text says tuning and training occur on a training set and the models are validated on a test set. Then it says high-performing combinations are "tested across all lead times" and selected for low KGE variance. What are the train, validation, and test sets? The final one- or two-year test period must remain untouched until all preprocessing, architecture choices, and hyperparameters are fixed.
4) How was the architecture of the encoder-decoder selected?
5) If my guess is correct, for each lead time, a model has been developed. This point should be very clear in the text, but unfortunately, it is not. If this is the case, the autoregressive property is being discarded.
6) Model implementations are not reproducible. Code and data should be provided. Based on EGU policies, this should have been flagged before archiving.
7) Search spaces differ greatly in capacity and completeness. Random Forest does not tune `max_features` or depth; XGBoost does not tune the regularization parameters discussed at lines 181–183; MLP width is restricted to 5–12 units and its depth is undisclosed; the recurrent networks use small, different grids. Consequently, conclusions about algorithm classes may instead reflect unequal tuning budgets or constrained implementations.
8) A fixed seed improves reproducibility but does not quantify variability. MLP/LSTM/ENCDEC results can vary with initialization and minibatch order. For fair comparison, an ensemble of seeds needs to be evaluated.
9) Models are trained separately at each site. Testing later dates at those same sites measures temporal evaluation, not evaluation at a new site. Zone summaries likewise do not constitute out-of-site validation. Therefore, please rephrase all "new site" claims or add leave-one-site-out/grouped-site validation in which the held-out site contributes no training, preprocessing, climatology, or tuning information.
10) Hyperparameters are optimized using scaled KGE′ but models are ranked and interpreted primarily by RMSE. Explain why this objective/evaluation mismatch is appropriate and report both metrics on untouched test data. RMSE alone does not reveal bias, event timing, extremes, or uncertainty; add MAE, bias, correlation/KGE components, and performance during wetting/drying events.
11) Without timeseries plots and more metrics, an evaluation cannot be completed.
12) The referencing needs an overhaul. For example, see the LSTM section. No reference for the properties is provided.
13) Many abbreviations are not defined or are defined late. e.g., ML, PCA, ConvLSTM, StemGNN, ....
14) Many abbreviations have been defined several times. e.g., MLP, LSTM, RNN, ...