the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Encoding-dependent verdicts and H_1 miscalibration in ensemble persistent homology of cyclone trajectories
Abstract. Ensemble topological data analysis (TDA) on multi-storm cyclone trajectories has been proposed as a tool to detect coherent perturbations — such as those associated with extreme geomagnetic events — that may not register in any single-storm intensity time series. We construct three principled longitude encodings for the cyclone point cloud (linear modular, unit-sphere, and lat-linear-plus-longitude-circle cylinder) and apply the same ensemble-TDA pipeline to the same data: all storms whose lifetime overlaps the ±15-day peak window of Halloween 2003, St Patrick’s Day 2015, or Gannon 2024 (25 event storms; 1,020 calendar-matched controls). Three findings emerge. First, the D1 pool H_1 permutation p-value depends on encoding choice in a way that flips the qualitative verdict: perm-p_H1 = 0.130 (linear), 0.009 (unit-sphere), 0.214 (cylinder), on identical event and control sets. Second, a 49-placebo calibration returns an H_0 false-positive rate close to the nominal 5 % in all three encodings (8.2 %, 4.1 %, 6.1 %), but an H_1 false-positive rate that is consistently above nominal in all three (8.2 %, 10.2 %, 14.3 %; directional consistency across encodings, not individually significant at n = 49 — the cylinder rate sits at the Wilson 95 % upper bound) — H_1 calibration is inflated by a factor 1.6×–2.9× regardless of encoding choice. Third, this dimension asymmetry between H_0 (calibrated) and H_1 (inflated) recurs in stratified attribution across two basin cells and three single-event cells, and in subsample-size sensitivity tests at N ∈ {400, 600, 800, 1000}. We interpret the pattern as observational evidence that ensemble TDA on cyclone-trajectory data carries an intrinsic H_1 inflation that is not curable by lon-encoding choice alone. The solar-perturbation hypothesis cannot be tested through this pipeline until an encoding-invariant H_1-calibrated protocol is in place. We frame the contribution as a methodological cautionary tale: pooled point-cloud TDA on geospatial trajectories with periodic coordinates is more fragile than its formal stability theorems suggest. We propose a minimum protocol of multi-encoding agreement testing plus dimension-resolved placebo calibration as a precondition for any positive ensemble-TDA claim in this class of problems.
- Preprint
(632 KB) - Metadata XML
-
Supplement
(123 KB) - BibTeX
- EndNote
Status: open (until 18 Aug 2026)
-
RC1: 'Comment on egusphere-2026-2975', Anonymous Referee #1, 18 Jul 2026
reply
-
AC1: 'Reply on RC1', Rongzhen Dai, 20 Jul 2026
reply
We thank the referee for three substantive comments. Since posting the preprint we have carried out a pre-registered repair-and-calibration campaign on the pipeline; it produced results that materially strengthen the paper and also correct one of its published claims. We summarise both, then respond point by point.
A correction we owe the record. While implementing the repair experiments (all criteria deposited before any run), we discovered that the published placebo-calibration figures were not bit-reproducible: the control-storm draw order inherited Python's per-process hash randomisation, so every run — including each published encoding — drew different control sets despite a fixed seed. After fixing this (deterministic ordering) and one latent crash-handling bug, we re-ran the full calibration under seven valid control-ordering implementations. The pooled H₁ false-positive count over 147 placebos ranges from 9 to 16 across implementations, with median 10 (binomial p = 0.34 against nominal 5 %); the published value, 16, is the single failing tail draw. The published claim "H₁ is inflated in every encoding" therefore does not survive: what survives — and what the revision will report — is stronger and more precise: the calibration verdict itself is implementation-dependent. A defensible re-ordering of the control draw flips the pipeline between "inflated" and "nominal". The revised §3.3.4 and §4.2 will state the correction explicitly, with all seven per-placebo tables archived.
On comment 1 (diagnosis without repair). We implemented the first two repair pathways of §5.6: an H₁ z-score normalisation and persistence-landscape L² statistics, each run through the full three-encoding placebo calibration. Neither improves on the permutation statistic (14/147 and 12/147 against the reproducible baseline's 9/147). The productive repair turned out to lie elsewhere: the corrected, deterministic pipeline at its median implementation is the first documented nominally-calibrated configuration of this class, and the revision adds a fourth requirement to the minimum protocol of §5.4 — calibration claims must be reported as a distribution over control-resampling implementations with a pre-registered summary statistic. We believe this answers the referee's underlying request: the revision no longer merely diagnoses; it delivers a calibrated configuration, a tested negative result on two proposed statistics, and a concrete protocol extension.
On comment 2 (physical interpretation; why encodings contradict). The revision consolidates the mechanism into a dedicated section with new quantitative content: the H₁ channel carries a 2–4× variance excess over H₀ in event distances and null means (all encodings); this produces an 11.6 % population of threshold-straddling placebos whose verdicts flip with the control draw, concentrated in the smallest-sample calendar window; a nuisance-dimension control shows ambient dimensionality is not a calibration driver; and a small robust false-positive core (two placebo dates under two encodings) persists across implementations and is reported as such. On the referee's conceptual question we will state the principle explicitly: a physical signal must be invariant under both re-encoding of coordinates and re-implementation of the resampling; a verdict that flips under either is thereby shown to be a property of the representation or the procedure, not of the physical system. A short forward-looking paragraph will indicate which mechanistic candidates of §1.2 would become testable once such invariance is achieved.
On comment 3 (condense the introduction). We accept; §§1.1–1.3 will be condensed by roughly 40 %, with the space reallocated to the mechanism and calibration-distribution sections above.
All per-implementation data, scripts, and the pre-registration files are archived and will accompany the revision.
Citation: https://doi.org/10.5194/egusphere-2026-2975-AC1
-
AC1: 'Reply on RC1', Rongzhen Dai, 20 Jul 2026
reply
-
RC2: 'Comment on egusphere-2026-2975', Anonymous Referee #2, 27 Jul 2026
reply
This manuscript applies ensemble topological data analysis (TDA) to tropical cyclone trajectories to test whether extreme geomagnetic events perturb cyclone activity at the ensemble-topology level. The authors construct three longitude encodings and report that H₁ permutation p-values depend critically on encoding choice (0.130, 0.009, 0.214), with H₁ placebo false-positive rates systematically inflated across all encodings (8.2%, 10.2%, 14.3%). While the experimental design is systematic, the manuscript suffers from fundamental weaknesses: the leap from "encoding-dependent results" to "tool cannot test the physical hypothesis" conflates representational robustness with causal robustness without justification; the H₁ inflation diagnosis remains descriptive rather than mechanistic; critical confounds of dimension versus periodicity and climatic variability versus methodological bias are acknowledged but not experimentally separated; the physical background is disproportionate to the explicitly non-physical conclusion; and the proposed minimum protocol is not itself followed. The study has potential value as a methodological diagnostic, but its current form risks being a well-executed yet inconclusive sensitivity study that discovers problems without advancing their solution or theoretical understanding. I recommend Major Revision.
1. The causal inference hierarchy of the core conclusion is confused. The paper equates "inconsistent results across three encodings" directly with "the tool cannot answer physical questions," but this leap requires an additional assumption—that a true effect should be detectable under all reasonable encodings—which is never justified. In fact, if a solar-cyclone effect exists but is encoding-specific, this might mean the effect couples to a particular geometric representation rather than indicating complete tool failure. The paper conflates three levels of robustness—statistical, representational, and causal—without clarifying which level it actually tests, and fails to rule out the possibility that a weak effect is amplified by only one encoding. The authors should clearly distinguish these three levels, justify why encoding inconsistency should be interpreted as tool failure rather than encoding-specific effect, or supplement with synthetic data experiments to calibrate the tool's detection capability.2. The explanation of H₁ false-positive inflation remains merely descriptive. The paper identifies elevated H₁ false-positive rates but inadequately probes their root causes: the "narrow null distribution" hypothesis lacks explanation for why it is narrow, "topological aggregation instability" is introduced as a working term rather than formal theory, and the mathematical properties of Wasserstein-2 distance itself are not excluded as an alternative explanation. W₂ allows noise features to be cheaply matched to the diagonal, potentially reducing the effective matching dimension for H₁; whether this bias is specific to W₂ choice rather than a general property of TDA remains untested. The authors should supplement with bottleneck distance control experiments, analyze the morphological differences in persistence distributions between H₀ and H₁, compute L² distances on persistence landscapes to test whether calibration improves, and clarify the differential predictions of topological aggregation instability versus W₂ mathematical bias.3. The placebo design lacks ecological validity, potentially confounding methodological defects with climatic variability. The 49 placebos use fixed dates with random years, but these dates correspond to specific seasonal phases with interannual variability in cyclone activity baselines, and no exclusion of ENSO phase, QBO phase, or other modulating factors is applied. If a random year happens to be a strong El Niño year, cyclone activity may be genuinely anomalous, causing placebo false positives to reflect climatic variability rather than methodological defects; the reported 8–14% H₁ false-positive rate may then overestimate or underestimate true methodological bias. The authors should apply climatological standardization to placebos, adopt a dual-placebo design separating seasonal from date effects, or report ENSO indices and sunspot number distributions for placebo years to demonstrate no systematic difference from event years.4. The three-encoding comparison suffers from dimensional confounding without isolating periodicity treatment from dimension increase effects. Linear modular is 3D and ignores periodicity, while unit-sphere and cylinder are both 4D with exact periodicity treatment; these two changes occur simultaneously, creating a confound. The critical control is missing: if linear modular were extended to 4D, would H₁ results still approximate the current 3D linear results? If 4D linear approximates sphere/cylinder, differences stem mainly from dimension rather than periodicity treatment; if it still approximates 3D linear, differences stem mainly from periodicity treatment. This control is essential for understanding the nature of encoding dependence. The authors should supplement with a 4D linear control experiment, or at minimum provide theoretical argumentation for why dimension increase should not affect the relative calibration pattern of H₀ versus H₁, and clarify the valid interpretive domain of the current encoding-dependence conclusion.5. The proposed minimum protocol lacks binding force, and the paper itself does not fully adhere to it. The three-encoding comparison was conducted post hoc, stratification cell choices are not explicitly pre-registered, and the placebo calibration threshold lacks power analysis. A deeper problem is the operational ambiguity of multi-encoding agreement: if two encodings are significant and one is not, what is the conclusion? If all three are significant but effect sizes point in opposite directions, what then? Does agreement require quantitative metrics? The authors should clarify whether their own analysis satisfies the proposed protocol, operationalize multi-encoding agreement into concrete decision rules, and provide a pre-registration template as supplementary material.6. The presentation of physical background is structurally at odds with the paper's positioning. The Introduction devotes substantial space to reviewing physical mechanisms of the solar-cyclone hypothesis, yet this discussion completely disappears in subsequent sections, while the conclusion explicitly refuses to make physical claims. This creates double reader frustration: physical-oriented readers are led to expect physical findings but receive methodological negation, while methodological readers must wade through irrelevant background to reach the core problem. The authors should fundamentally restructure: either compress physical background to half a page and reallocate space to expanding encoding geometry discussion and H₁ mechanism analysis, or retain physical background but add conditional physical discussion clearly distinguishing the current null (tool unreliability) from a physical null (effect nonexistence).Citation: https://doi.org/10.5194/egusphere-2026-2975-RC2 -
AC2: 'Reply on RC2', Rongzhen Dai, 03 Aug 2026
reply
RC2.1 (the causal-inference hierarchy is conflated; an effect could be encoding-specific rather than tool failure). We agree and have rewritten the core claim to keep the levels distinct (§5.1). We no longer equate encoding-inconsistency with tool failure; instead we (i) show the encoding verdict is representation-dependent, (ii) show the calibration is reproducibly nominal but order-sensitive, and (iii) implement a physically defined distance that removes the representational ambiguity. We map the three robustness levels explicitly: statistical (§4.2), representational (§4.1, §5.1), and causal (§4.5, §5.3). The synthetic-injection calibration of detection power the referee proposes is the right next instrument; we have not run it and say so (§5.3) — the present revision establishes the precondition (a physically defined, ordering-stable, calibrated pipeline) that makes such an injection test interpretable. Any surviving difference is framed as necessary-but-not-sufficient for attribution.RC2.2 (the H1 inflation is descriptive; is it a W2 artefact? Add bottleneck, landscapes, H0-vs-H1 morphology). Done (§4.2). Mechanism: the H1 channel carries a 2.0×–2.2× coefficient-of-variation excess over H0 in the event-to-null distance (and 3.9×–4.4× in the null mean), producing a narrow H1 null whose swing straddles α. Not W2-specific: an exact bottleneck distance (gudhi, ε = 0) is, if anything, more order-sensitive — a sign test on the paired per-placebo ranges gives p = 1.5 × 10⁻⁶ — so the diagonal matching of W2 is not the cause. An L2-persistence-landscape statistic gives 12/147 and does not restore calibration either. We also correct an earlier over-statement: in absolute terms the H0 permutation p swings marginally more than H1 across orderings; what is H1-specific is that its swing crosses α.RC2.3 (placebo ecological validity; ENSO/QBO could confound the false-positive rate). Two lines address this (§4.2, §4.5). First, the same placebo dates — hence identical climate state — change false-positive status purely when the control-draw order is permuted, so the variability across orderings cannot be climatic. Second, we no longer treat the placebo count as a pure methodological false-positive rate: we now describe it as the pipeline's empirical rejection frequency under realistic historical variability, which folds in climate, era and basin variability, and we probe how much of the event-side signal that variability can carry with the matched-null designs of §4.5 (per-window/seasonal count, basin, era, and joint basin×era). Third, on the specific ENSO/solar question, Supplementary §S5 reports the ENSO phase and solar state of the event and placebo windows: the three events fall in neutral-to-weak-El-Niño conditions and sit in the middle of the placebo ENSO distribution (which spans strong El Niño to strong La Niña years), so the event set is not ENSO-selected; the events are, by construction, more solar-active, but a climatic or solar driver cannot produce a verdict change at the fixed climate and solar state of a single placebo date.RC2.4 (dimension is confounded with periodicity; add a 4-D linear control). Done, with an honest split (§3.3.5). A 4-D linear encoding with a constant fourth coordinate is a code-path identity (it must reproduce the 3-D result bitwise, and does). A 4-D linear encoding with a Gaussian fourth coordinate leaves the calibration nominal (2/49) but does move the event verdict (p 0.037 → 0.072). We therefore separate two claims — the calibration behaviour is dimension-robust, but the event p-value is dimension-sensitive — and explicitly concede that the confound the referee raises is real at the level of the §4.1 event verdict.RC2.5 (the minimum protocol is not itself followed; operationalize agreement). We add a fourth, binding protocol requirement (report calibration as a distribution over orderings; §5.4) and, importantly, a paragraph that flags rather than hides where the present analysis does not meet its own protocol (n = 49 < 200 placebos; B = 100 permutations). A pre-registration template with an operational multi-encoding-agreement rule is added as Supplementary §S1.RC2.6 (physical background is disproportionate to a non-physical conclusion). Accepted; see RC1.3. The introduction now leads with the metric-choice problem and treats the solar–cyclone question as a case study.Citation: https://doi.org/
10.5194/egusphere-2026-2975-AC2
-
AC2: 'Reply on RC2', Rongzhen Dai, 03 Aug 2026
reply
-
RC3: 'Comment on egusphere-2026-2975', Anonymous Referee #3, 29 Jul 2026
reply
Recommendation: Major revision
This manuscript addresses an important methodological problem: whether persistent-homology results for cyclone trajectories depend on the representation of periodic longitude. The comparison among linear, spherical, and cylindrical encodings is potentially useful, and the manuscript is generally clear and appropriately cautious about making a physical solar–cyclone claim. The reported encoding-dependent H1 p-values are striking. However, several central conclusions are not yet supported by the experimental and statistical design. Overall, the methodological question is worthwhile, but the title, abstract, and conclusions should be substantially weakened unless the metric comparison, matched-null design, and calibration analysis are revised.
Major comments
- The encodings change not only longitude periodicity but also coordinate weighting, pairwise distances, ambient geometry, and in the cylindrical case the underlying H1 topology. The manuscript itself acknowledges that the cylinder can generate a longitude loop with no physical content. Therefore, disagreement among the three results does not by itself demonstrate “encoding fragility”; it may simply reflect comparison of three different geometric models. The main analysis should instead use a common physically defined distance matrix, such as great-circle distance combined with a clearly justified and sensitivity-tested intensity weight.
- The three events have strongly different seasonal and basin compositions, yet the controls are pooled across the three calendar windows and n storms are sampled uniformly from the complete pool. The null draws should preserve, at minimum, the number of storms from each event window and preferably basin composition. Otherwise, changes in basin and seasonal geometry may dominate the persistence diagrams. It should also be clarified whether all lifetime fixes are included for a storm that merely overlaps the ±15-day window. The methods appear to concatenate all fixes from selected storms, which could introduce observations outside the proposed exposure period and give longer-lived storms greater weight.
- Only 49 placebos are used, with just 100 null iterations per placebo. The manuscript reports a shared Wilson interval of 1.4–14.7%, but confidence intervals should be calculated around each observed proportion. For example, the Wilson 95% interval for 7/49 is approximately 7.1–26.7%, not the interval reported as the nominal acceptance range. More importantly, no paired comparison demonstrates that H1 rejects more frequently than H0; under the linear encoding both rates are exactly 4/49. The three encoding results are also correlated because they use the same placebo dates. A paired calibration analysis, more placebos, and substantially more permutation draws are required before using terms such as “intrinsic inflation.”
- Multiplying 0.009 by the ratio 10.2/5.0 does not produce a statistically calibrated p-value. Calibration should instead be based on an empirical null distribution of p-values or a nested permutation procedure. Monte Carlo p-values should also use the standard +1 correction and include uncertainty arising from the finite number of permutations.
- Each sample size appears to be represented by one seeded realization. Non-monotonic p-values across N=400–1000 may therefore reflect ordinary sampling variability rather than a structural noise mechanism. The experiment should be repeated over many independent seeds, using identical sampled point indices across encodings, and should report distributions or confidence intervals rather than single p-values.
Citation: https://doi.org/10.5194/egusphere-2026-2975-RC3 -
AC3: 'Reply on RC3', Rongzhen Dai, 03 Aug 2026
reply
RC3.1 (use a common, physically defined distance rather than three geometric models). This is the revision's main addition (§3.3.6, §4.5). The distance is d(i,j) = sqrt[(great-circle angle/π)² + (λ·|Δwind|/200)²]. It is not another arbitrary embedding: the great-circle term is the intrinsic geometry of the sphere the storms live on (an angle between unit position vectors, not a parameter choice), and the intensity term is a dimensional balance whose single weight λ is sensitivity-tested at {0, 0.5, 1.0}. The sensitivity test is itself informative: pure track geometry (λ = 0) is null (p = 0.157), and significance appears only as intensity is weighted up — so any event–control difference is carried by the intensity channel, not the track topology. Under this distance the verdict is single-valued and stable across seven control-draw orderings (per-ordering 2–4/49).RC3.2 (matched null preserving per-window count and basin; clarify fix inclusion). Done (§4.5, §2). The bootstrap control draw now preserves the event ensemble's per-window (seasonal) storm counts (11/9/5) — the referee's "at minimum" requirement — as well as its genesis-basin, era and joint basin×era composition. The event–control difference survives the per-window count matching (perm-p 0.004–0.010 across orderings) and basin matching (0.003–0.004), is most attenuated by era matching (era is the leading confounder), and none removes it at p < 0.05. On fix inclusion, §2 now states explicitly that the ensemble uses all lifetime fixes of each selected storm (the storm, not the fix, is the unit of exposure and resampling), that this includes out-of-window fixes and weights longer-lived storms more, and that a window-cropped variant is a natural sensitivity for a larger study.RC3.3 (per-proportion Wilson intervals; paired H0/H1; more placebos and permutations). Addressed (§4.2). We replace the shared Wilson interval with per-proportion Wilson intervals (e.g. 1/49 → [0.4, 10.7]%, 4/49 → [3.2, 19.2]%). We add a paired H0-vs-H1 comparison and report it honestly: the H1 excess over H0 is modest and encoding-specific — under the linear encoding H1 (1/49) is below H0 (2/49) — so we claim no general H1-over-H0 inflation. We also note that the three encodings in each pooled count of 147 share the same 49 placebo dates and are therefore correlated, so the pooled exact-binomial p-values understate the sampling uncertainty and are used only to place the counts on the nominal side; the per-proportion Wilson intervals independently support this. We agree B = 100 is small and flag it as a limitation (§5.4); a per-placebo decomposition shows much of the ordering spread is finite-B Monte-Carlo noise, so raising B would compress the spread rather than reveal an inflation. Consistent with all of the above, the term "intrinsic inflation" is withdrawn from the manuscript, and any statement that presupposed it no longer applies.RC3.4 (multiplying by a ratio is not a calibrated p-value; use an empirical null and the +1 correction). Accepted in full. The 0.009 × (10.2/5.0) ratio scaling is removed. The calibration is now the placebo empirical null — 49 placebo permutation tests per encoding compared to nominal with the exact binomial (§3.6). The permutation p-value is reported as the fraction #{null ≥ observed}/N; we verified that the finite-sample +1 correction shifts every value by < 0.001 and changes no false-positive count, so the verdicts do not depend on that choice.RC3.5 (single-seed sizes N = 400–1000; repeat over seeds and report distributions). We agree the non-monotonic p-values reflect ordinary sampling variability, not a structural size effect, and we no longer read them as the latter (§4.4). Table 4 is now explicitly presented as one seeded realization per cell, illustrating variability rather than a size-dependence curve; a full multi-seed distribution with encoding-matched sample indices is left to a larger study (stated in §3.6 and §4.4). Where calibration is the claim, we do report distributions — over six control-draw orderings for the encodings and seven for the physical distance.Citation: https://doi.org/
10.5194/egusphere-2026-2975-AC3
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 50 | 23 | 4 | 77 | 12 | 4 | 2 |
- HTML: 50
- PDF: 23
- XML: 4
- Total: 77
- Supplement: 12
- BibTeX: 4
- EndNote: 2
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
This manuscript tested three existing encoding strategies for ensemble TDA on cyclone trajectories to examine the effects of the solar and geomagnetic activity perturbation on the TC. It is revealed that the results are sensitive to different coordinate encoding methods and intrinsic H1 false positive inflation. Overall, this study is only a methodological cautionary one. It neither proposes a new method nor delivers new physical insights, which means that it has relatively weak novelty. Hence, the current version of this manuscript may be unsuitable for publication in current journal. The following are my main concerns.
1.This manuscript only verifies flaws embedded in the existing ensemble TDA. But there are no original methodological advancements proposed to solve the flaws. In other words, the manuscript only discovers methodological defects without fixing them. Whether the author can propose a new strategy or method to overcome the existing flaws? This will be able to improve the novelty of the manuscript.
2.This manuscript is too technical. It has not discussed the fundamental problems of the physical and methodological limitations. Could the author provide substantial physical interpretations regarding the linkage between solar storms, geomagnetic activity and TC? Additionally, from both methodological and physical perspectives, how to understand the phenomenon that different encoding schemes yield contradictory results?
3.If the manuscript does not aim to investigate the physical linkage between solar activity and tropical cyclones and only serves as a methodological diagnostic, the discussions of physical processes in the introduction should be substantially condensed and analyzing methodological flaws should be stressed.