Automated Homogenisation of Early Instrumental Temperature Data at the Global Scale Using the HCLIM Dataset
Abstract. Early instrumental observations provide a unique opportunity to investigate climate variability before the establishment of modern meteorological networks. However, their effective use is often hindered by incomplete records, uneven spatial coverage, and non-climatic inhomogeneities. The HCLIM dataset provides a recently compiled global collection of early instrumental temperature records, yet systematic evaluations of automated homogenisation methods for such data remain limited. Here, we evaluate three widely used homogenisation methods, CLIMATOL, PHA, and BART, applied to the HCLIM dataset. The assessment considers data retention following pre-processing, breakpoint detection in the homogenised series, and agreement with the Twentieth Century Reanalysis version 3 (20CRv3), using French and Southeast Asian station networks as representative case studies. The pre-processed datasets differ only modestly in overall structure, although BART retains fewer records (80 %) than CLIMATOL (96 %) and PHA (99 %) because of its stricter requirement for a minimum record length of 15 consecutive years. Consequently, BART preserves longer average record lengths and detects substantially more breakpoints, indicating greater sensitivity to potential inhomogeneities. Breakpoint occurrence varies considerably across regions and time periods, reflecting differences in station density and observational coverage. Comparisons with 20CRv3 show generally high agreement for all methods, with stronger consistency in the dense French network than in the sparser Southeast Asian network. Despite differences in data retention, breakpoint detection, and adjustment strategy, the homogenised datasets show strong agreement with one another and preserve the principal climatic signal. Notably, differences among the homogenisation methods are consistently smaller than the differences between the original and homogenised datasets, indicating that the primary challenge is not selecting a single optimal algorithm but reconstructing temporally coherent climate information from incomplete historical observations. The homogenised datasets produced in this study constitute HOM-HCLIM, a new global database of homogenised early instrumental temperature records that provides an important resource for climate reconstruction, reanalysis, and long-term climate variability studies.
General comments:
The authors have applied three homogenization methods to improve the quality of the HCLIM historical climate observations dataset. They apply several statistics to compare the results and conclude that all of them seem to successfully remove the non climatic biases in the series. Differences between the methods can be ascribed to the density of observatories in time and space, and to the default parametrization of the methods. Their conclusions reassures the validity of homogenization for a better assessment of climate variability and trend studies based on observational series.
Apart from this article, the climatological community will greatly benefit if the authors undertake the intended future work including daily observations, additional reanalyses and novel approaches for breakpoint detection and adjustment.
Specific comments, preceded by the section or line number:
73-74: “SNHT identifies potential inhomogeneities through comparison of accumulated anomalies.” This is more a description of the Craddock method. SNHT works by standardizing the data and calculating a test statistic that identifies the peak point in the series where the before and after means differ the most.
2.2.1 Exclusions of records: Having read other papers comparing the performance of several homogenization methods, some reasons for the rejection of records seem quite inconsistent to me. This is the case, for example, with the exclusion of records from nHom_CLIMATOL because of outliers of temperature, since CLIMATOL automatically deletes outliers greater than a prescribed threshold and goes on with the homogenization process, so there is no reason to exclude series because this software detects some outliers in them. (But if by “records” you mean individual data items rather than series, then please disregard this comment.) Also strange is the exclusion because of low correlations (Table 2), since correlation has no role in the selection of neighbors in CLIMATOL, or “Too long distance (defined by the program)” and “Big gaps”, when no limits are set by the program to these parameters. It is clear that too long distance to neighbors may compromise the reliability of the homogenization, but this will be true also for BART, PHA and any other relative homogenization methods. As to BART and PHA, readers would probably like to know which are those “Other reasons” in Table 2.
205-207: CLIMATOL provides a function for the automatic conversion of SEF files into its own input formats. Apart from that, the sentence “CLIMATOL requires CSV input files, with all series combined into a single spreadsheet and accompanied by an .esd metadata file.” is wrong, since both the data and metadata files are plain text.
224: “frequency-based Standard Normal Homogeneity Test”. Frequency-based? Please correct or explain what you mean with this sentence.
428-429: “No statistically significant differences were detected among hom_CLIMATOL, hom_BART, and hom_PHA.” This sentence is repeated in line 427.
Fig. 11, caption: “While the hom_BART dataset generally has smaller RMSD values compared”. Should not “smaller” be changed by “higher”?
698: Following CRAN, the canonical form to cite the CLIMATOL package is https://CRAN.R-project.org/package=climatol
787: “The Twentieth Century625”. “625” should be deleted.
817: “StateM eteorologicalAgency(AEM ET )” should be “State Meteorological Agency (AEMET)”.
Table S1. Corrections to CLIMATOL column:
Metadata usage: Optional.
Station distance constraints: None in the code. (But the user should avoid using reference stations so far away as to compromise a fair correlation with the tested station.)
Internal option for reference stations: Geographical.
Minimum number of series: 3 (2 if one is assumed to be homogeneous, such as from reanalysis).
Step 4: Missing data estimation.
Correlation threshold (Step 1, …): None (in any step).
(Correlation is only calculated for the initial cluster analysis, which is offered to the user to allow judgment on whether the area can be considered free of clear climatic boundaries. No further use of the clustering is applied to the selection of reference stations.)
Correction method: Missing data estimation.