the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Satellite Data Records as a Tool to Monitor Changes in Air Temperature
Abstract. We analyze satellite-derived lower-tropospheric temperature (TLT) data for the period 1981–2025 and examine their relationship to the pronounced warming observed in in situ measurements during 2023–2024. To reduce uncertainty and improve the robustness of detecting long-term changes in Earth's warming rate, we construct an adjusted TLT record by removing variability associated with the El Niño–Southern Oscillation (ENSO) and atmospheric aerosols. This adjustment reduces the magnitude of annual TLT variability by nearly 50 %, substantially enhancing the robustness and reliability of the observed satellite temperature trends. Using the adjusted TLT record, we identify statistically significant warming trends of up to 0.482 ± 0.113 °C decade⁻¹ after 2015 across all satellite and reanalysis datasets examined in this study. These trends are approximately four to five times larger than those during the pre-2015 period. However, these trend estimates are likely conservative. At the upper end, statistically significant increases in the warming rate of up to 0.48 ± 0.12 °C decade⁻² are inferred near 2024, indicating that the pronounced warming observed during 2023–2024 are part of an ongoing increase in the underlying warming rate that was further amplified by the El Niño event. Projections based on these increased warming rates suggest the potential for an additional 0.5–1.0 °C of lower-tropospheric warming over the next decade. The physical mechanisms responsible for this unusually rapid warming, however, remain unclear. Resolving these mechanisms is therefore essential for improving future climate projections and informing effective mitigation and adaptation strategies.
Status: open (until 07 Oct 2026)
- RC1: 'Comment on egusphere-2026-3691', Anonymous Referee #1, 30 Aug 2026 reply
-
RC2: 'Comment on egusphere-2026-3691', Anonymous Referee #2, 14 Sep 2026
reply
This manuscript uses satellite-derived lower-tropospheric temperature (TLT) data from 1981–2025 to evaluate whether there is support for an acceleration in the underlying warming TLT trend. The authors report an astonishingly high warming rate of +0.482±0.113 °C decade-1 since 2015, after statistically mitigating the contribution of aerosols and the El Nino-Southern Oscillation.
Unfortunately, I believe there are statistical issues with the manuscript that mean I cannot support publication. Substantial work would be needed to reproduce the analysis and account for these issues, and I am not sure that the conclusions will remain robust once they are addressed. I list my major comments below.
I find the topic interesting and the authors bring a strong background of knowledge about the details of the datasets involved, factors which are often brushed over in simple statistical analyses. I also liked the appendix, e.g. Figure A4 is a good example of researchers evaluating the consequences of joint choices.
I believe it is possible to do a similar assessment that would yield robust results. After the Major concerns, I will briefly comment on Monte-Carlo approaches that might allow development of such a method. I end with some Minor concerns if the authors continue their analysis.
Major concern 1: model complexity
The statistical fits have a lot of parameters, once you include ENSO, AOD and either polynomials or the change points. It’s not clear these are justified: some justification or demonstration beyond R2/errors would be needed. There are information criterion methods, or brute-force Monte-Carlo as I mention below.
The model complexity also raises strong questions about Figure 4. I can perhaps see the argument for displaying “continuing new trend after breakpoint” as an illustration, it is at least somewhat parsimonious and could be used as a reference as new data arrive. But expanding out a cubic fit does not seem to have any physical or statistical basis.
Major concern 2: autocorrelation
Even after removing the parts that correlate with ENSO and AOD, I would expect there to be autocorrelation in the residuals. I do not see if this has been accounted for, and the consequences for uncertainties can be large. I am very sceptical of the narrowness of the reported confidence intervals. Foster & Rahmstorf’s 2011 paper (https://iopscience.iop.org/article/10.1088/1748-9326/6/4/044022/meta) is useful here; the Appendix shows how temperature data require corrections that are even more sophisticated than AR(1).
This concern, since it potentially means that the effective sample size is different from that assumed in all of the tests, could affect all conclusion statements. And it can also affect some information criterion values.
Major concern 3: interpretation of and uncertainty in change/acceleration points, including multi-trial effects
The selection of the existence & location of a breakpoint involves multiple trials. I do not see how this was accounted for in the uncertainty calculations here. I think the reported significance is derived for “if I randomly select one breakpoint, is this significant” rather than what you actually did: “I tried a lot of breakpoints, based in part on prior knowledge of the dataset, and then used the most significant”.
The derived confidence intervals could be grossly understated. Similarly, Figure 3 shows an “acceleration year” for cubic fits – my problems with this are that the cubic model is suspect (see e.g. Major Concern 1) and that the uncertainties in the year are not shown.
Major concern 4: sequential regression
The decision to regress against NINO3.4 and aerosol first, and then against other factors also concerns me. In a sequential regression, if the first terms (ENSO/AOD) share any low-frequency covariance with the underlying trend or polynomial terms, the first step can inappropriately attribute some of the TLT variation to the ENSO/AOD terms. I believe a simultaneous regression with all parameters in the design matrix should be used, including the ENSO index, aerosols, and whatever other fit terms are selected.
SUGGESTED SOLUTION:
I do like the use of different datasets and the discussion of features that introduce time-varying differences in them. That is a partial exploration of systematic error, and if combined with a more robust statistical analysis we might be able to conclude there is evidence for acceleration in these satellite records. However, the issues identified above mean that I simply do not believe the reported confidence intervals and therefore the claimed conclusions have not, in my view, been rigorously tested.
My solution would be to use Monte-Carlo methods to help develop and validate statistical methods that work. Generate large numbers of pseudorandom series and/or use large ensemble output, as in e.g. Glaser et al. (2024, https://iopscience.iop.org/article/10.1088/2752-5295/ad69b6/meta). Apply your tests, including the calculated confidence intervals, and see whether your statistical approach appropriately estimates p values and the likes.
MINOR COMMENTS:
The discussion of aerosol was unclear to me. My understanding is that the general approach of statistically “removing” aerosol targets the effect of episodic volcanism. It seemed like here there were also possibly assertions being made about tropospheric aerosol, which is a complex topic on its own.
The NINO3.4 index uses a rolling 30-year baseline. When long-term warming occurs within the final 30 years of the record, there is an artificial drift in this index. I would expect consequences for an analysis like this, where the end points of the data are important. See van Oldenborgh et al. (2021, https://doi.org/10.1088/1748-9326/abe9ed) for the Relative ONI/RONI alternaive. I attached a plot of the recent RONI minus ONI difference: there’s a trend that could affect your results, with a recent difference of ~0.5 units. If I’m thinking this through correctly, your choice of NINO3.4 would actually result in a negative bias in derived acceleration.
LOESS is referenced: I like LOESS. But I don’t see where the window size or weighting are described. Is it e.g. a common default of 1/3 of the data length with a tricube weighting scheme?
Data sets
ENSO-Aerosol Adjusted Annual Global Mean TLT Time Series (NOAA V5.0, UAH V6.1, RSS V4.0, and ERA5-equiv) Xianjun Hao http://wamis.gmu.edu/cdr/products.html
Model code and software
Source Code for Publications Xianjun Hao http://wamis.gmu.edu/cdr/pub/code_paper/index.html
Viewed
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 122 | 0 | 1 | 123 | 0 | 0 |
- HTML: 122
- PDF: 0
- XML: 1
- Total: 123
- BibTeX: 0
- EndNote: 0
Viewed (geographical distribution)
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
There are huge issues with this submission, all making accelerated warming appear more likely than it statistically is (right now). While I also suspect that planetary warming has accelerated recently, the authors' analysis here seems rather misguided.
1) First, subtracting/adjusting for the effects of ENSO, solar aerosols, etc. on the temperatures, analyzing the adjusted series, and then claiming that any trend findings for it also pertain to the original temperature series is fallacious. The authors justify this via the references [8] and [40-41], which make similar mistakes to varying degrees. We could also fit and remove additional factors like PDO, AO, CO_2, etc. and obtain even more significant findings. Where does it stop? The point is what you have found does not apply to the original temperatures (the ones we experience on Earth) any more.
2) Second, annual temperature series are usually positively autocorrelated. I do not see where autocorrelation is taken into account in the standard error margins or p-values for the various model fits. If you don't account for correlation, your standard errors and p-values will be too small, making insignificant fluctuations appear significant.
3) Third, the changepoint techniques are naive. At least, one can't fit the model for every potential breakpoint time b, find the p-value assuming this value of b is known, and then report results for the best b. This never accounts for the variability of the changepoint time, leading again to overstated conclusions.
There are other issues, but the above three dominate.
In the end, I am afraid there's nothing salvageable here: 1) above isn't fixable, and while 2) is, doing 3) right is hard (at least, current research papers on joinpin changepoint models are just now appearing in the statistics literature).