the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
NeuralFAO56 v1.0: A Scalable Physics-Informed Deep Learning Framework for Real-Time Evapotranspiration Estimation Across CONUS
Abstract. Accurate estimation and forecasting of reference evapotranspiration (ETo) are essential for irrigation demand estimation. The FAO-56 Penman–Monteith formulation remains the physical standard for ETo computation, while recent advances in deep learning (DL) have demonstrated strong predictive skill for ETo forecasting. However, real-time ETo forecasting remains constrained by manual meteorological station identification, heterogeneous data acquisition, and labor-intensive preprocessing workflows. Existing software tools primarily support physics-based ETo estimation without real-time data integration or forecasting capability, whereas DL-based approaches often require manual data preparation, limiting automation and real-time applicability. This study introduces NeuralFAO56 Python package, a hybrid physics–data DL computational framework that embeds neural network–driven forecasting architectures for on-demand ETo estimation and forecasting across the continental United States (CONUS). NeuralFAO56 couples physics-based FAO-56 with DL sequence modeling within a unified pipeline that enables automated data acquisition, standardized preprocessing, and scalable deployment. The framework operates in dual modes: (i) physics-based FAO-56 ETo estimation using observed and forecasted meteorological inputs, and (ii) data-driven ETo forecasting using Long Short-Term Memory (LSTM) and Transformer architectures for multi-horizon (up to 7-day lead time) real-time forecast. The framework is evaluated across 867 stations at continental US spanning different climate regions. Results demonstrate strong short-term predictive skill, with performance degradation at longer lead times driven by reduced temporal predictability. Higher forecasting skill is observed in climatologically stable regions, while comparatively lower performance occurs in humid, convectively active regions. Overall, NeuralFAO56 provides a scalable, real-time framework that integrates physically based ETo modeling grounded in energy and mass conservation with DL forecasting and automated meteorological data pipelines to support short- to medium-range irrigation planning and management.
- Preprint
(3191 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 19 Sep 2026)
- RC1: 'Comment on egusphere-2026-2300', Anonymous Referee #1, 24 Aug 2026 reply
-
RC2: 'Comment on egusphere-2026-2300', Anonymous Referee #2, 02 Sep 2026
reply
This manuscript presents NeuralFAO56 v1.0, an automated Python framework that combines FAO-56 Penman–Monteith reference evapotranspiration calculations with LSTM- and Transformer-based forecasting and integrates NWS/NCEI meteorological data acquisition within a single workflow. I see practical value in the automated data pipeline, the CONUS-wide implementation across 845 stations, and the attempt to compare physics-based and data-driven forecasting approaches within the same operational framework.
However, after reading the manuscript, I remain uncertain about the central scientific and operational contribution of the proposed framework. In particular, the manuscript needs to more clearly establish what is meant by “physics-informed,” why the DL pathway is needed when meteorological forecasts can already be propagated directly through FAO-56, and exactly what information is used by the DL models. These issues are central to the interpretation and reproducibility of the study.
Major comments
1. The use of “physics-informed” in the title and throughout the manuscript needs stronger justification.
The title presents NeuralFAO56 as a “Physics-Informed Deep Learning Framework,” and the manuscript frequently refers to coupling physics-based FAO-56 with DL. However, based on the current methodology, I understand the framework primarily as two forecasting pathways integrated within a common software environment: FAO-56 uses observed or forecast meteorological forcing, whereas the LSTM and Transformer models perform data-driven forecasting from historical records.
I could not identify a mechanism through which the FAO-56 physics directly informs the learning process—for example, through a physics-based loss term, physical constraints, conservation requirements, physics-derived internal states, or another explicit interaction with the neural-network architecture.
The authors should therefore explain precisely what makes the DL component “physics-informed.” If the integration occurs mainly at the workflow/software level rather than within the learning process, terminology such as “hybrid physics-based and data-driven framework” may more accurately describe the contribution. This is not simply a terminology issue because the current wording implies a stronger methodological integration between physics and machine learning than is evident from the manuscript.
2. The central added value of the DL pathway over direct FAO-56 forecasting is not yet convincingly demonstrated. Figure 10 should play a much more central role in establishing this value.
The practical motivation for the DL component is currently unclear to me. If short-term meteorological forecasts are already available from NWS, future ETo can be obtained directly by propagating those forecasts through FAO-56. The key question is therefore not simply whether an LSTM or Transformer can reproduce or predict a historical FAO-56-derived ETo time series, but whether the DL pathway provides meaningful additional predictive skill. If short-term weather forecasts are already available, users can directly calculate future ETo with FAO-56. Why is an additional DL forecasting pathway needed?
Figure 10 appears to be the most important experiment for answering this question because it directly compares forecast-driven FAO-56, LSTM, and Transformer predictions against a common reference derived from observed meteorology. However, this prospective comparison is limited to a single seven-day period, 8–14 November 2025.
Given that this is the only direct prospective comparison between the two forecasting strategies, the manuscript should explain why this particular period was used and whether it was selected a priori based on data availability or another criterion. More importantly, a single November window does not establish whether the relative performance shown in Figure 10 is representative of other meteorological conditions. For much of CONUS, November is also outside the primary growing and irrigation season, whereas the manuscript makes broader claims regarding irrigation planning and operational readiness.
I therefore hope the authors could strengthen this part of the study substantially. Multiple independent forecast windows covering contrasting conditions—particularly growing-season/high-ETo periods—would allow the manuscript to answer the much more useful question of when, where, and at what lead time the DL pathway provides added value over forecast-weather-driven FAO-56. Without such evidence, it remains difficult to understand why a user should select the DL pathway rather than directly use available meteorological forecasts with FAO-56.
3. The DL predictor set and the information available to the models at forecast time are not sufficiently documented. This is important both for reproducibility and for interpreting the regional results.
The manuscript lists meteorological variables retrieved from NWS/NCEI and explains that input variables exceeding the 5% missing-data threshold are removed before training. However, I could not find a clear specification of the exact predictor vector ultimately supplied to the LSTM and Transformer models. Please explicitly report the exact meteorological variables used as DL predictors, clarify whether any forecast meteorological information is supplied to the DL models during the forecast horizon, and indicate whether the same predictor set is used across all stations.
Specific comments
1. Solar-radiation approximation and fixed kRS.
The FAO-56 reference ETo depends on incoming solar radiation, but Rs is estimated using the Hargreaves–Samani formulation with kRS fixed at 0.16 for all stations. Because the study spans nine NOAA climate regions and interprets spatial differences in ETo forecasting performance, the uncertainty associated with applying a spatially uniform kRS deserves further consideration. I do not assume that this choice necessarily explains the reported regional performance patterns; however, the manuscript currently provides little information on how sensitive the derived ETo values are to this assumption. A limited sensitivity analysis using representative stations or climate regions would therefore be useful to determine whether reasonable variation in kRS materially affects the magnitude of ETo or the main regional conclusions. If the impact is small, demonstrating this would also strengthen confidence in the robustness of the current framework.2. Line 52.
The citation contains an apparent editing artifact: “(4/30/2026 12:05:00 PMAllen et al., 1998; ASCE, 2005).” Please correct this citation.3. Line 137.
“Data acquisition componenet” should be corrected to “Data acquisition component.”4. Lines 711–716 / Conclusion.
The statement “Across 845 stations selected to represent four distinct NOAA climate regions” appears incorrect. The full analysis includes 845 stations across nine NOAA climate regions, whereas four representative stations from four regions are used for the detailed station-level analysis. Please revise this sentence accordingly.Citation: https://doi.org/10.5194/egusphere-2026-2300-RC2 -
RC3: 'Comment on egusphere-2026-2300', Anonymous Referee #3, 07 Sep 2026
reply
This manuscript presents NeuralFAO56 v1.0, a Python-based hydroinformatics framework that integrates automated acquisition of meteorological data from NWS and NCEI, FAO-56 Penman-Monteith reference evapotranspiration calculations, and station-specific LSTM- and Transformer-based forecasting. The implementation across 845 NWS-NCEI matched stations is ambitious and demonstrates that the framework can be deployed over a broad geographic domain. The automated workflow, software architecture, and effort to integrate physics-based and data-driven approaches within a common operational framework are potentially useful contributions within the scope of Geoscientific Model Development.
However, in its current form, I do not think the manuscript yet satisfies the level of methodological traceability and scientific support expected for final publication in GMD. My principal concerns relate to the physical and computational traceability of the meteorological preprocessing chain, the spatial representativeness of the station-selection procedure, the distinction between continental-scale deployment and true spatial generalization, and the reproducibility of the real-time forecasting experiment.
These issues are important because the FAO-56 ETo series generated by the framework serves both as a physical benchmark and, effectively, as the target against which the DL forecasts are assessed. Therefore, the transformation from heterogeneous raw meteorological data to daily FAO-56 inputs needs to be documented with enough precision that an independent scientist could construct a scientifically equivalent implementation.
I recommend reconsideration after major revisions. I believe the framework may become suitable for publication if the authors substantially improve methodological documentation, address the unsupported claims identified below, and conduct targeted diagnostic analyses where necessary.
Specific comments
- The manuscript states: “NWS provides meteorological data at sub-daily temporal resolution, whereas NCEI archives provide data on a daily scale. Following data acquisition, these variables are preprocessed and transformed into a standardized set of daily inputs for the FAO-56 ETo computation.” The manuscript also states that sub-daily NWS data are “aggregated to daily values prior to FAO-56 ETo computation.” This transformation is central to the model, but the aggregation procedure is not described in sufficient detail. Please specify, for every meteorological variable, the native temporal resolution, native units, unit conversion, daily aggregation operator, definition of a meteorological day, treatment of time zones, handling of partial days, and treatment of forecast periods that cross calendar-day boundaries. The manuscript should also clarify whether historical NCEI records, real-time NWS observations, and NWS forecasts are placed on exactly the same daily temporal basis before FAO-56 calculation and DL training. This is particularly important for Tmin, Tmax, humidity-related variables, and wind speed. A preprocessing table showing source variable, source dataset, native resolution, aggregation method, target unit, derived FAO-56 variable, and QC rule would substantially improve reproducibility.
- The FAO-56 formulation defines “u2 [as] the mean wind speed at 2 m height,” whereas the data-acquisition section only states that “wind speed” is retrieved. Please specify the wind variable retrieved from each data source, the measurement/reference height, whether historical and real-time datasets follow the same convention, and whether wind speed is normalized to 2 m prior to FAO-56 calculation. If a height correction is applied, please provide the equation and describe how missing measurement-height metadata are handled. Because wind speed directly enters the aerodynamic component of the Penman-Monteith equation, this is an important physical requirement rather than only a software implementation detail.
- The framework retrieves both relative humidity and dew-point temperature, while the FAO-56 calculation requires saturation and actual vapor pressure. Please explain the hierarchy used to derive actual vapor pressure. For example, is dew-point temperature used preferentially when available? How are relative-humidity data aggregated to daily values? What happens if one humidity-related variable is unavailable? Is the same procedure used for NCEI historical data, NWS observations, and NWS forecasts? A short decision tree or pseudocode would improve traceability considerably.
- The manuscript states that the LSTM and Transformer models are “trained on station-specific historical datasets.” However, the conclusion refers to “spatial generalizability.” Training and testing a separate model using historical data from each station demonstrates that the workflow can be applied across many stations and climate regions. It does not demonstrate that a trained model generalizes to an unseen location. I recommend either revising the terminology throughout the manuscript to refer to continental-scale deployment, multi-site applicability, or robustness across a geographically diverse station network, or adding a true spatial holdout experiment in which stations or geographic regions are withheld from model training. The current 845-station experiment is valuable, but it should not be interpreted as evidence of spatial transferability unless unseen-location testing is performed.
- The manuscript states: “Stations are ranked by geodesic distance, and the nearest suitable station is selected.” It then states: “real-time meteorological observations are retrieved from the selected NWS stations, while short-term forecasted meteorological variables are obtained from gridded NWS forecast products corresponding to the user-specified location.” The manuscript subsequently describes this as enabling forecasting with “fine spatial resolution across CONUS.” This claim requires additional support. The observation station and user-requested coordinate may represent different locations, while the forecast forcing is obtained for the requested location. Please clarify the maximum allowed station distance, the distribution of user-location-to-station distances, whether elevation difference is considered, which latitude and elevation enter the FAO-56 calculation, and whether any correction is applied when the station and requested location differ substantially in elevation or physiographic setting. I suggest reporting station-selection distance and elevation-difference distributions and implementing a representativeness or quality flag. Without such evaluation, the term “fine spatial resolution” should be moderated.
- The manuscript states: “To evaluate NeuralFAO56 under true real-time forecasting conditions, model forecasts were assessed during the post-testing deployment window spanning 8–14 November 2025.” However, the forecast initialization/issue time is not clearly defined. It is therefore difficult to determine whether the seven lead times originate from one forecast initialization, a rolling sequence of daily forecasts, or another retrieval procedure. Please report the forecast retrieval timestamp, issue/initialization time, valid forecast time, lead-time definition, NWS product/API endpoint, and whether the experiment is based on one forecast cycle or rolling updates. The exact forecast meteorological fields used in the analysis should ideally be archived with the study data. This is necessary for reproducibility because statements such as “day 1” through “day 7” are meaningful only relative to a clearly defined forecast initialization time.
- Some regional hydroclimatic explanations are stronger than the evidence presented. For example, the manuscript attributes higher skill to “strong seasonal forcing and relatively stable atmospheric conditions” and lower skill in other regions to “higher humidity variability, convective weather systems, and stronger synoptic-scale disturbances.” These explanations are physically plausible, but they are not directly tested in the presented analysis. I recommend either moderating the language to “may reflect,” “is consistent with,” or similar wording, or supporting the interpretation with a quantitative diagnostic analysis linking station-level forecast skill to measurable properties such as humidity variability, temperature variability, wind variability, ETo persistence, or other atmospheric indicators.
- The term “significant” should not be used without an inferential test. The manuscript states that “no significant differences are observed between the two DL architectures.” The reported results show very similar regional means and strongly overlapping distributions, but I did not identify a formal statistical test supporting the word “significant.” Please either conduct an appropriate paired station-level comparison with effect sizes and confidence intervals or revise the wording to indicate that the models showed nearly identical regional mean performance and overlapping station-level distributions.
- The conclusion that model architecture is not important is stronger than the experimental evidence supports. The manuscript states that the similar performance of LSTM and Transformer “indicat[es] that ETo forecasting is primarily constrained by data characteristics and inherent predictability rather than model architecture.” The results demonstrate that the two specific lightweight configurations evaluated here perform similarly. However, the manuscript itself notes the absence of station-specific hyperparameter tuning, and both architectures were intentionally constrained for computational efficiency. The current experiment therefore does not distinguish conclusively between intrinsic predictability limits and limited architectural differentiation. I recommend revising this conclusion to state that, under the lightweight configurations evaluated, increasing architectural complexity did not provide an evident performance advantage.
- Please distinguish calculated FAO-56 ETo from directly observed evapotranspiration. The manuscript repeatedly uses terminology such as “observed FAO-56 ETo.” FAO-56 ETo is calculated from observed meteorological inputs; it is not itself directly observed in this study. I recommend consistently using “FAO-56 ETo calculated from observed meteorological inputs” or “observation-driven FAO-56 ETo.” This distinction is particularly important because the same calculated ETo also serves as the reference target used in the model evaluation.
- Computational scalability and operational scalability should be distinguished. The runtime results demonstrate that station-specific model training is computationally feasible on the reported hardware. This is useful evidence of computational scalability. Operational scalability, however, also involves model retraining frequency, data latency, API outages, model caching, model versioning, changes in station metadata, station discontinuity, and fallback behavior. Please clarify whether models are trained at every request or cached, when retraining occurs, how failures in data retrieval are handled, whether fallback from Mode 2 to Mode 1 is communicated to the user, and whether model/data versions are logged. This would make the operational claims more concrete.
- A worked physical example would substantially strengthen the model description. I strongly encourage the authors to provide one worked example tracing a single station/day through the complete workflow: raw API variables – QC -- temporal aggregation -- unit conversion--Tmin/Tmax -- humidity/vapor-pressure calculation -- wind-height processing -- radiation calculation -- FAO-56 intermediate variables -- final ETo.
- Technical corrections
- Replace “observed ETo” or “observed FAO-56 ETo” with “FAO-56 ETo calculated from observed meteorological inputs,” where appropriate.
- Section 2.4.4 uses the heading “ETo Forecasts Across Gauging Stations.” These appear to be meteorological/observation stations rather than hydrologic gauging stations. Please revise the terminology.
- Define and use consistently the terms forecast horizon, lead time, forecast day, and prediction horizon.
- Define clearly the distinction among “real-time,” “near-real-time,” and “on-demand.”
- Quantify or moderate the claim of “fine spatial resolution across CONUS.”
- Ensure that mechanistic interpretations of regional forecast skill are clearly identified as interpretations unless quantitatively demonstrated.
- Please review the manuscript carefully for consistency in station counts, station terminology, units, module names, and references.
Citation: https://doi.org/10.5194/egusphere-2026-2300-RC3
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 125 | 51 | 13 | 189 | 17 | 17 |
- HTML: 125
- PDF: 51
- XML: 13
- Total: 189
- BibTeX: 17
- EndNote: 17
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The manuscript presents an operationally useful, well-designed Python framework of NeuralFAO56 v1.0 that integrates FAO-56 Penman–Monteith reference-ET calculation with LSTM- and Transformer-based ET forecasting, using NWS real-time and forecast meteorology and NCEI historical archives, across a continental network of NWS–NCEI matched stations. The scope is suitable for a GMD Model description. The supporting evidence for the current publication, however, is not sufficient. The manuscript should be rethought after the authors make a response to the following substantive concerns. I would suggest the following comments, and hope that it will be helpful for the authors. I would suggest another resubmission: