Detecting soil moisture drought impacts on ecosystem physiology from earth observation data
Abstract. Satellite remote sensing is widely used to monitor vegetation drought impacts, yet water stress triggers both structural changes and physiological adjustments, whose relative importance varies across space and time. Structural responses, captured by multispectral vegetation indices are well documented, while the extent to which readily available satellite data can detect physiological drought effects remains insufficiently quantified.
Here, we assess how much information on drought-induced physiological responses is contained in multispectral reflectance, land surface temperature (LST), and climate reanalysis data. We combine MODIS reflectance and thermal data with an ERA5-Land potential cumulative water deficits metric (PCWD), and train a machine-learning model using eddy covariance data. The target variable represents drought-induced reductions in light use efficiency (fLUE), explicitly separating physiological regulation from structural canopy changes.
Our resulting data-driven model captures variations in drought-related physiological response more accurately than previously documented indices commonly applied for drought monitoring (R² = 0.64, RMSE = 0.112, spatial cross-validation). Application across central Europe during two recent summers demonstrates that the model detects drought impacts on photosynthesis earlier and more sensitively than NDVI, particularly in evergreen ecosystems where structural signals are muted.
Our results show that a systematically trained, data-driven integration of multispectral reflectance, thermal signals, and climate data can extract a substantial portion of the physiological drought response from Earth observation data. This approach enables a more confident and spatially consistent assessment of drought impacts on photosynthesis using the full combined information content of satellite and climate datasets, beyond what structural vegetation indices alone can provide.
I want to thank the authors for their originality in the modelling and for their scientific rigour. Paper defenitly is not boring. Reading it, I do feel the paper tries combining a few different points, and it lacks a clear direction, making the paper hard to read. Most of my comments are in that spirit. I know few models that consider cummulated precipitation deficit. My concerns of this paper are mostly about what this paper is for. It sometimes reads too technical, at other instances too general/conceptual. Most of my comments are therefore on the readability and internal structure of the paper. The main question to answer to me is: who is the reader you have in mind for this paper? Are you writing to the modelling community or to the ecophysiology community. Of course it's a bit of both, and I must admit this is not an easy question to answer. My feeling is that the paper is relevant for ecophysiologists who are not particularly knowlegable in the field of modelling. Try making sure the paper can be understood by such a person (while keeping the scientific rigour).
Good luck with the revisions.
line 44: I'm certain you know this, but I would mention 'canopy greenness and canopy structure here' as two different points.
Line 98: your marks already suggest you thinkg that that ground truthing is an ugly word. I agree with that. Rather say ground referencing
Line 110: 2018 is not that recent anymore. But I like the work.
Introduciton general: I am sure that this approach is complementary to the SMAP GPP product, but elaborate in the added value of your approach over the SMAP GPP product. I presume this is linked to cumulative water deficit. Explain the need for considering cumulative water deficit in the introduction. The cumulative water deficit is not included in most models, so it is a great selling point and I would even call it the raison d'être of this model (as many others do include non-greenness indicators such as LST).
Methodology: I presume the drought-related information
Part 2.4.1. NIRvP is also a popular proxy for photosynthesis. Would be great if you could check that one, too.
Line 286-294: highly technical language. I think what you did is correct, but add in a describe sentence or two what you did on a more conceptual level. Or try doing something else to make this part more readable for someone without a background in machine learning (in my head: we want machine learning models to interpolate rather than extrapolate).
Line 313: do not call it "observed" fLUE but repeat the precise data source (FLUXNET if I'm not mistaken?)
there is no section 3.3.
3.4.: this is a section that is relevant for vegetation ecologists that think more in terms of processes and field observations than in terms of statistic metrics. Make sure the text is understandable by such a profile.
3.5.: link the paragraphs 3.5 and 3.4 more clearly
figure 4: what is the point of subplots e and f? I cannot find that clearly in the text, or at least not in pararaph 3.4.
Section 4 (before 4.1.). Bit of an odd section, as it summarizes the introduction. Either move it to the introduction or make a clearer link to your own results. You can use a theoretical recap to introduce the discussion, but make the link clearer.
The fact that nd_NR_B1_NR_B4 is the biggest predictor, suggests that greenness still plays an important role in your machine learning algorithm. Please elaborate on that.
Make a clearer link to figure 2 in the discussion (and to the results of it) when discussing the drivers.
A thing to clarify better within the paper: is it your main goal to study drought in europe (difference between regions and ecosystems) or is the dataset/model itself the thing you are trying to sell. If the latter: try answering more clearly who you believe would benefit from this model.
paragraph 4.2. is very good, clearly shows the advantages. However, using NDVI as reference for a reflectance-only/greenness-only approach is a straw man argument. Look up some studies that have used more refined vegetaion indices (like MTCI/MTVI2 etc) to tackle this issue. I do think your initial argument still stands though.
The main mechanism that LST is sensitive to stress conditions goes over the latent heat flux, which reduces under stomatal closure. I am sure you already know this mechanism, but you will make the paper readable to a larger audience if you explain that mechanism explicitly in the introduciton.
Data availability: think about who your target audience is. I strongly believe this work is relevant plant biologists who are to non-ML experts. If you want them to use your model (which I think would be great), I would recommend you to publish the model results alonside this paper (in a zenodo link or something). Either way, downloading someone else's model, getting all the input data yourself, it's a tedious thing to do. Researchers are more inclined to using your model if you prepare a dataset that's readily available.
If you are more writing to the modelling community (which you have the right to do), then add an in-depth comparison where you compare this model to other drought models using (i) thermal remote sensing (ii) meteorological indicators such as SPEI (iii) remote sensed soil moisture/SIF. What is the selling point of your model in the landscape of all the others?