the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Physics-Constrained Transfer Learning with a Spectral-Fidelity-Preserving Model for Satellite Remote Sensing Applications
Abstract. Accurate spectral transformation across satellite sensors with similar but different spectral response functions (SRFs) are essential for applying the same retrieval algorithms. A novel physics-constrained transfer learning (TL) framework is developed for transferring satellite radiance observations across different sensors while preserving physical consistency. It integrates a core Spectral-Fidelity-Preserving (SFP) model based on extensive radiative transfer simulations, allowing broad adaptability for radiance transformation under diverse satellite observational conditions. Sensitivity experiments demonstrate the robustness of the TL framework relating to radiometric calibration uncertainties, particularly in infrared (IR) channels, and further highlight the critical role of SRF similarity between sensors. Specifically, the scaling factor between the SRFs of the target and reference channels should be constrained within the range of 0.5 – 1.5. Meanwhile, the shift in central wavenumber should remain below 200 cm⁻¹ for visible or near IR channels, and more strictly below 20 cm⁻¹ for infrared window channels (e.g., 10.80 µm). Applying to radiance observations from Fengyun-4A/B (FY-4A/B) geostationary (GEO) satellites explicitly indicates that the TL approach improves retrieval accuracy for key geophysical parameters such as cloud amount profile and quantitative precipitation estimation, when compared those without applying TL. Thus, the TL approach enhances cross-satellite data consistency and provides a practical tool for operational satellite data applications (e.g., adopt algorithms of F-4A to FY-4B without operational interruption).
- Preprint
(7007 KB) - Metadata XML
-
Supplement
(2699 KB) - BibTeX
- EndNote
Status: final response (author comments only)
-
CC1: 'Comment on egusphere-2026-3959', Mengchu Tao, 16 Jul 2026
-
AC1: 'Reply on CC1', Min Min, 17 Jul 2026
Comment 1: From a practical application perspective, what do you think is the most important advantage of this spectral transfer framework compared with simply developing a new retrieval model for each new satellite sensor?
Response: Thanks for your question. In our view, the central advantage is that the framework decouples sensor characterization from retrieval-model development, so operational continuity can be achieved as soon as a new sensor's spectral response functions (SRFs) are known — without waiting to accumulate a new multi-year labeled dataset for that sensor, and without retraining or re-validating the retrieval model itself.
This is not a hypothetical benefit; it is illustrated directly by the FY-4B/CALIPSO case in Section 4.1. FY-4B did not enter operational service until December 2022, and by the time FY-4B/AGRI observations became routinely available, the overlap with CALIPSO had already narrowed to well under a year (June 2022–March 2023). This window would likely have been too short to independently develop and robustly validate a dedicated FY-4B-based CLANN model in the way the original FY-4A-based model was built (Lin et al., 2025). The physics-constrained TL framework instead allowed the existing, already-validated FY-4A-based CLANN model to be applied directly to FY-4B observations, using only the FY-4B/AGRI SRFs — normally characterized before or shortly after launch — rather than new satellite–CALIPSO matchups. The QPE case in Section 4.2 follows the same logic.
Two secondary benefits are also worth noting: avoiding the computational and engineering cost of retraining a deep-learning retrieval model for every new sensor generation, and preserving the internal consistency of long-term, multi-satellite data records, since independently retrained models risk introducing artificial discontinuities into an operational time series.
Comment 2: Your transfer model is trained based on MODTRAN simulations with 83 atmospheric profiles and different cloud, aerosol, surface and geometry conditions. How sensitive is the performance to the representativeness of these simulated atmospheric states? For example, could extreme conditions such as deep convection, polar regions, or unusual aerosol environments introduce additional uncertainties?
Response: Thanks so much. We would like to answer this honestly rather than overstate the current coverage. Table 1 does provide some relevant coverage: cloud optical depth spans from 0 (clear sky) to 216 (cumulus), intended to represent optically thick conditions broadly consistent with deep convective cores, and surface types include tundra and snow, which partially address high-albedo, cold-surface conditions relevant to polar regions. The 83 ECMWF profiles are chosen to span multiple climate regimes, and the resulting radiance distribution in Figure 2 is correspondingly broad.
That said, we agree the database has real, identifiable limits, and the reviewer's examples usefully expose them:
- Aerosol conditions are represented by only two discrete type/optical-depth combinations (ocean, τ = 0.17; desert, τ = 0.78). Unusual aerosol environments with different optical properties — biomass-burning smoke, volcanic ash, heavy anthropogenic haze — are not explicitly sampled, and we would expect these to introduce additional transfer uncertainty in solar-reflective channels beyond what is quantified in Section 3.3.
- We have not separately verified how many of the 83 profiles specifically represent polar/high-latitude conditions, nor conducted a stratified sensitivity analysis by latitude band.
- MODTRAN is used here under a plane-parallel treatment. We already note in the Discussion that this limits applicability to higher-spatial-resolution sensors; the same limitation applies to strongly three-dimensional scenes such as deep convective cores, where horizontal photon transport near cloud edges is not captured.
Citation: https://doi.org/10.5194/egusphere-2026-3959-AC1
-
AC1: 'Reply on CC1', Min Min, 17 Jul 2026
-
RC1: 'Comment on egusphere-2026-3959', Anonymous Referee #1, 17 Jul 2026
Overall Assessment
This manuscript presents a physics-constrained transfer learning framework centered on a Spectral-Fidelity-Preserving model for transferring satellite radiance observations between sensors with similar but non-identical spectral response functions. Its key innovation is to derive transfer coefficients from radiative-transfer-simulated and SRF-convolved radiance pairs, rather than relying solely on empirical regression using satellite matchups. The sensitivity analysis in Section 3 is comprehensive, combining distributional similarity metrics, fitting diagnostics, and analytical uncertainty propagation to assess the joint effects of SRF differences and calibration uncertainty. The two case studies in Section 4, cloud amount profile retrieval using CLANN and quantitative precipitation estimation using an UNet framework, further demonstrate that models trained on FY-4A data can be transferred to FY-4B observations without retraining. The proposed quantitative thresholds for SRF scaling and central-wavenumber shifts are also practically useful for satellite calibration and cross-sensor algorithm transfer. Overall, the manuscript is well designed and technically sound, but several issues should still be addressed before publication.
Specific Comments
- The direction implied by "target" and "reference" is not fully consistent across sections. Eq. (3) treats *y* (the output) as the target-sensor radiance and *x* (the input) as the reference-sensor radiance, whereas Section 4.1 calls the FY-4B input the "target" channel and the FY-4A-equivalent output "reference" — the reverse assignment. The authors themselves note in Section 3.3 that Figure 9 uses the opposite convention from Figure 3. Stating one consistent convention in Section 2.2, and flagging it explicitly wherever it is deliberately reversed, would resolve this.
- Readers may expect "physics-constrained" to mean a physical penalty term added to a loss function, but here physical consistency is instead built into the training data itself, via the radiative-transfer simulations (Section 2.1) and the SRF-convolution step (Eq. 4). Stating this explicitly would also clarify why the framework reduces SRF-driven bias between sensors. A brief note on why R² ≥ 0.95 and *n* ≤ 4 were chosen in Section 2.2, and how many simulated pairs enter each channel's fit, would also help.
- The manuscript touches on three related applications — spectral harmonization, retrieval-model transfer, and radiometric calibration transfer — without clearly distinguishing them. The FY-4A/FY-4B experiments here mainly demonstrate the first two, while the calibration-transfer claim relies more on the earlier FY-3D/MERSI-II study. A short clarification in the Abstract and Conclusions would make this distinction, and the manuscript's practical claims, more precise.
- Eqs. (17)–(21) and (A20)–(A21) propagate an assumed perturbation in the target channel through a fixed transformation, while uncertainty in the fitted coefficients, SRF characterization, and RTM simulations is implicitly treated as negligible. "Bias," "error," and "uncertainty" are also used somewhat interchangeably. A short paragraph near Eq. (17) or in Section 3.3 stating the scope of this derivation would be enough to resolve this, without a new analysis.
- The SRF-similarity guidance (scaling factor 0.5–1.5; wavenumber shift below 200 cm⁻¹ for VIS/NIR, below 20 cm⁻¹ for IR window channels) is repeated in near-identical wording in the Abstract, Section 3.2, and Discussion point (2). Keeping the full statement in Section 3.2 and shortening the other two to a brief cross-reference would reduce this redundancy.
- In Section 4.1, the sentence "Consequently, previous work did not attempt to develop an FY-4B/AGRI-based CLANN retrieval model" appears twice in immediate succession, each followed by a similar remark. This reads as a leftover from editing; removing the duplicate would tighten the paragraph.
- In Section 3.3, the sentence "Fig. 4 further indicates that when the radiometric calibration bias becomes large (e.g., reaching 6%)..." appears to reference the wrong figure: Figure 4 shows simulated SRF adjustments, whereas the 1%–6% error range described matches Figure 9's caption instead. The authors are encouraged to verify and correct this cross-reference.
- In Eq. (15), "Upper Fence" is defined as Q25 − 1.5×IQR and "Lower Fence" as Q75 + 1.5×IQR, which is the reverse of standard convention (upper fence = Q75 + 1.5×IQR; lower fence = Q25 − 1.5×IQR). The authors are encouraged to correct this labeling.
- There is a date discrepancy between the text and Figure 11: Section 4.2 gives Case 1 as "11 June 2024" and the extrapolation period as ending in "July 2024," while the Figure 11 caption gives "11 July 2024" and an end date of "August 2024." The authors are encouraged to reconcile these dates.
- In Section 4.2, the Multi-Task UNet model is attributed to "Jaegle et al., 2021," which in the reference list corresponds to the Perceiver architecture rather than a U-Net. The authors are encouraged to verify this citation.
- The same ~13.30 μm channel is labeled "Channel 15" (FY-4B convention) in the Figure 3 caption but "Channel 14" (FY-4A convention) in Section 3.3. A brief note clarifying the channel-number correspondence between the two sensors would prevent confusion.
- A few language points warrant a check: the Abstract's opening sentence has a subject–verb agreement issue ("transformation ... are essential" should be "is essential"); "compared those without applying TL" is missing a preposition; and "F-4A" appears to be a typo for "FY-4A."
Citation: https://doi.org/10.5194/egusphere-2026-3959-RC1 -
CC2: 'Reply on RC1', Min Min, 17 Jul 2026
Thank you very much for your feedback. We will carefully implement the necessary revisions based on your suggestions in the coming days and will provide you with point-by-point responses accordingly.
Citation: https://doi.org/10.5194/egusphere-2026-3959-CC2
-
RC2: 'Comment on egusphere-2026-3959', Anonymous Referee #2, 24 Aug 2026
Review of “Physics-Constrained Transfer Learning with a Spectral-Fidelity-Preserving Model for Satellite Remote Sensing Applications”
This study presents a physics-constrained transfer learning (TL) framework for transforming satellite radiance observations across sensors with different spectral response functions (SRFs). The core component is a Spectral-Fidelity-Preserving (SFP) model built on MODTRAN-based radiative transfer simulations, which derives polynomial transfer coefficients between spectrally similar channels on different satellite imagers. The framework is evaluated through sensitivity experiments on SRF discrepancies and radiometric calibration uncertainties, and demonstrated through cloud amount profile and quantitative precipitation estimation using FY-4A and FY-4B AGRI observations.
The concept of grounding cross-sensor transfer learning in radiative transfer physics rather than purely data-driven approaches is sound and practically relevant, particularly for operational satellite transitions (e.g., FY-4A to FY-4B). The sensitivity analysis quantifying the impact of SRF width scaling and central wavelength shifts on transfer performance is a useful contribution, and the empirical guidelines (scaling factor 0.5–1.5, JSD ≤ 0.3) provide actionable criteria for operational users.
Overall, the manuscript is well-structured, although several minor concerns and corrections are still needed. I therefore recommend minor revisions prior to acceptance, listed as follows.
- The FY-4B/AGRI against CALIPSO validation period for the cloud amount profile experiment is limited to approximately 10 months (June 2022–March 2023). While the TL model reduces the RMSE by 0.08 km and improves the correlation coefficient from 0.841 to 0.851 (lines 645–646), these improvements are relatively modest. Given the limited validation sample, the authors should discuss whether these differences are statistically significant. Reporting confidence intervals, a paired t-test, or a bootstrap significance test on the RMSE/R differences would substantially strengthen the claim that the TL framework yields a "measurable improvement" (line 648). Without such analysis, it is difficult to determine whether the observed gains exceed the natural sampling variability.
- The authors separate the theoretical framework into VIS/NIR (Eq. 1) and TIR (Eq. 2) regimes, and the error propagation analysis (Section 3.3) is similarly divided between reflectance-based (Eqs. 20–21) and BT-based (Eqs. A20–A21) formulations. However, the 3.75 µm channel included in the CLANN retrieval model (line 607) contains both reflected solar and thermal emissive contributions during daytime. This dual-source nature means that neither the pure VIS/NIR nor the pure TIR transfer model fully captures the channel's radiative behavior under all illumination conditions. It is suggested to briefly discuss how the SFP model and the transfer coefficients handle this mixed channel, particularly whether the MODTRAN simulations adequately capture the solar reflection component at 3.75 µm across the range of solar geometries in Table 1, and whether the transfer performance degrades for daytime observations at this wavelength.
- The manuscript provides several empirically derived guidelines for applying the TL framework, scattered across Sections 3.2, 3.3, and 5. These include: SRF scaling factor within 0.5–1.5, central wavenumber shift below 200 cm⁻¹ for VIS/NIR channels, below 20 cm⁻¹ for IR atmospheric window channels (e.g., 10.80 µm), and JSD ≤ 0.3 as a transferability threshold. I suggest providing a concise summary table consolidating these recommended thresholds, for example, organized by channel type (VIS/NIR, water vapor IR, atmospheric window IR) with the corresponding metric, threshold value, and consequence of violation. This would improve the manuscript's practical utility and would serve as a quick-reference guide for potential users.
- The manuscript states that "the classical MODTRAN (Version 4.2) model (Berk and Hawes, 2017)" is employed (line 296–297). However, the cited reference (Berk and Hawes, 2017) describes the validation of MODTRAN 6, not Version 4.2. Please make sure that the version of MODTRAN used is correctly cited.
- Lines 625–629 contain a near-identical repeated sentence: "Consequently, previous work did not attempt to develop an FY-4B/AGRI-based CLANN retrieval model" appears twice within five lines. Please clarify.
More technical issues:
- Line 92: "adopt algorithms of F-4A to FY-4B": "F-4A" should read "FY-4A."
- Line 93: The keyword list ("Radiative transfer, transfer learning, Physics-Constrained") uses inconsistent capitalization. Consider standardizing to either title case or sentence case throughout.
- Table 1: The surface type "Forest" uses "P" (Ρ = 0.08–0.83) instead of "ρ".
- The manuscript would benefit from a brief notation table or consistent first-use definitions for the numerous abbreviations (AGRI, SFP, SRF, BT, TL, FTL, MTL, FL, UDA, SSL, etc.), as the density of acronyms in the Introduction may challenge readers from adjacent fields.
Citation: https://doi.org/10.5194/egusphere-2026-3959-RC2 -
AC2: 'Reply on RC2', Min Min, 25 Aug 2026
Thank you very much for your suggestions. We will carefully implement the necessary revisions based on your suggestions in the coming days and will provide you with point-by-point responses accordingly.
Citation: https://doi.org/10.5194/egusphere-2026-3959-AC2
Viewed
| HTML | XML | Total | Supplement | BibTeX | EndNote | |
|---|---|---|---|---|---|---|
| 200 | 52 | 24 | 276 | 35 | 17 | 20 |
- HTML: 200
- PDF: 52
- XML: 24
- Total: 276
- Supplement: 35
- BibTeX: 17
- EndNote: 20
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
1.From a practical application perspective, what do you think is the most important advantage of this spectral transfer framework compared with simply developing a new retrieval model for each new satellite sensor?
2.Your transfer model is trained based on MODTRAN simulations with 83 atmospheric profiles and different cloud, aerosol, surface and geometry conditions. How sensitive is the performance to the representativeness of these simulated atmospheric states? For example, could extreme conditions such as deep convection, polar regions, or unusual aerosol environments introduce additional uncertainties?