Beyond Spectral Smile: Comprehensive Imaging Spectrometer Wavelength Calibration from Atmospheric Features
Abstract. Accurate spectral calibration is essential for quantitative imaging spectroscopy, as even sub-nanometer wavelength errors can propagate into atmospheric correction and surface-property retrievals. Existing approaches for in-flight wavelength calibration typically assume simplified spectral parameterizations, limiting the complexity of wavelength variations that can be recovered from atmospheric absorption features. In this study, we investigate the information content of visible-to-shortwave infrared (VSWIR) observations for constraining instrument wavelength calibration using a Bayesian maximum a posteriori retrieval framework. Using nine high signal-to-noise EMIT (Earth Surface Mineral Dust Source Investigation) scenes acquired over spectrally homogeneous desert targets, we systematically evaluate spline-based wavelength calibration models with varying numbers and placements of spline knot points. Model performance is assessed using agreement with Zemax optical simulations, solution consistency across independent scenes, and leave-one-out cross-validation. All three evaluation criteria identify a four-knot spline representation as the optimal balance between model flexibility and stability, whereas more complex parameterizations exhibit increased sensitivity to knot placement and reduced reproducibility. Applying this optimal configuration independently across the detector array reveals coherent cross-track wavelength-dispersion variations that cannot be adequately represented by traditional uniform or simple spectral-smile corrections. Although the retrieved spatial variations are small—typically on the order of 1 % of a spectral channel width—they exhibit a structured saddle-shaped pattern that is consistent across scenes and indicative of genuine instrument behavior. These results demonstrate that atmospheric absorption features provide sufficient information to retrieve spatially varying wavelength calibration for modern VSWIR imaging spectrometers, supporting more accurate radiometric processing and motivating spatially resolved spectral calibration strategies for current and future spaceborne missions.
This study focuses on the EMIT spectral calibration, specially the spectral dispersion. The novelty of this work mainly yields in extending the calibration simultaneously to both the spectral and spatial dimensions, providing a more specific calibration that better accounts for non-uniformities from the instrument, which ultimately improves applications such as atmospheric absorption characterization. The study is therefore well motivated and the science community can really benefit from these findings.
The work seems to be strongly based on previous publications from David R. Thompson. So much so, that I needed to read Thompson et al., (2024) to better understand this work and I noticed great redundancy. Along the text, I found that some important details regarding the methodology and results description were missing, making diffficult to follow the narrative of the manuscript. Moreover, I found important inconsistencies in the results that make me doubt about how reliable are the methods used here, specially in the results from Figure 5. Hopefully, the authors can address all these concerns.
Comments
L20 - In L4 this was expressed as 'visible-to-shortwave infrared'. Please, select just 1 way to express it.
L21 - There are more applications, such as gas retrievals. Please, rephrase as sth like 'a single instrument to address numerous applications such as...' or similar.
L26 - Any reference to back up this point?
L36 - This sentence can be misleading because 'microns' can also be used to refer to the nomial central wavelength. Please, clarify that this is a mechanical movement.
L40 - 'based on' instead of 'with'?
L40 - Do they only fit the atmosphere? and the surface? Please, clarify
L46 - and what about the distribution of sensitivity to the channel?
* Wavelength dispersion is a term that is not well described here. Please, clarify the meaning of this term.
L51 - Where is this analysis coming from? Please, cite if there is a good reference that backs this up. If this is something that has been found and shown later in this work, please consider to leave it for the results/discussions/conclusion part.
2.1. - Please, consider to make some comments here about the central wavelength and FWHM values that are provided in EMIT acquisitions. The point of this study is to study variations in reference to this data.
L69 - I'm not sure whether the citations are well inserted. Please, review this sentence.
L73 – ‚spatially exensive‘ - Then, are these scenes covering a wider area in comparison to most acquisitions? Please, clarify
L76 - One of my main concerns is that the instrument spectral calibration might be conditioned by the temperature, which could be dependent on the position in the orbit. What do you think of this effect and, if it's important, are the set of selected locations representative enough?
L76 - It'd be also interesting to analyze the same areas in different years in order to assess the temporal variability of the spectral calibration.
L87 - So the instrument is internally thermally regulated? Therefore, my previous comment about the position of the satellite in the orbit is not valid? Moreover, please, cite the source from which you obtain this information.
Figure 1. Are dunes and land changes such as those of Scene 5 breaking the homogeneity assumption?
L98 - Main concern: What about the striping effect? Making bins of 18 cross-track columns could in fact increase the dispersion. I think this is an important point. Please, clarify.
L100 - Where is this smooth function assumption coming from? Please, cite if needed.
L121 - I don't think this is true. The elements in the state vector related to the surface are typically the coefficients of the polynomial that is used to characterize the surface. This would be much clearer if the forward model were expressed in a equation. Please, clarify.
L122 - Wouldn't then this forward model only be valid for the 550 nm band? Please, clarify
L125 - Please, if not explained later, explain here a more detailed explanation of the process: why are these wavelengths selected for the four-knots cubic spline parameterization? how do you introduce the spectal calibration parameters into the state vector and the forward model? Now you assume a smooth/continuous perturbation in reference to the initial calibration... why? Is this assumption taken in the across-track direction shifts, or in the shifts from a particular down-track bin, or both?
Eq. 3 - This is the same equation as in Thompson et al., (2024) with the difference of the 1/2 factor, which can be removed since we are talking about a cost function that be minimized. Please, cite the formula correctly or just cite the reference and not inlcude the formula. On the other hand, I would also include the formulation of the forward model and see how you include the parameters to fit regarding the spectral calibration. Moreover, you included f(x) instead of F(x), which is how you formulated the forward model in a previous equation.
* So far, I'm seeing lots of similarities with the work of Thompson et al., (2024).
L141 - Is the nominal the one provided from EMIT metadata? Please, clarify.
L145 - I think that the previous selection of 4 knots from Thompson (2024) was the fact that the location of these 4 knots matchd the inflection points from the spectral dispersion. Which were the criteria followed here? Please, clarify.
L151 - Please, extend the explanation regarding the role of Zemax optical simulations. It's not clear for me whether the Zemax simulations set up really reproduce the satellite-based acquisitions from the EMIT instrument.
L155. Main concern: Why in this way and not the other way around? i.e. uniform in spectral dimension and linear in spatial dimension. Moreover, how is the linearity established? I guess that the only way to understand this is knowing how the Forward model depends on these parameterization and how are the paremeters set as priors next to the a priori covariance. For the moment, this setup is not well described, which makes this work fail for reproducibility. Some details that are key for reproducibility are missing.
L159 - Using less than 2 knots something that is worth it? Or is it just set to test?
L173 - What do you mean with single retrieval configuration? Do you mean just 1 iteration in the inversion? Does this explain why they differ so much? Please, explain.
L180 - I do not agree. You cannot say that this indicates a reliable retrieval of the sensor parameters. You could say that the offsets are showing a similar shape, but their magnitude changes significantly. Does this mean that the instrument spectral dispersion changes importantly with time or the fit is not making a good job or the selection of number and location of knots is not really good? Please, clarify.
L185 - 44 sequences are tested. However, how are 2 and 3 knots strategies working for the Strat 1, Strat 2, and On-orbit if these start from 4 pre-defined knots? Please, clarify. Moreover, which are these heuristic choices and which is the difference between Strat 1 and Strat 2? More detail to really understand what the authors tested is needed.
L193 - What does model selection mean? It's not clear. Vey confusing. Please, clarify.
L195 - Summing the errors or combining the discrepamcy from retrieval and Zemax? It's not clear. Please, clarify.
L195 - If you use Zemax as ~ground truth, why don't you just directly use the results from Zemax to set your calibration? I would prefer the calibration measured in laboratory as a reference to check consistency. Or at least, by-pass this and preserve your strategy by first checking that Zemax and laboratory results are consistent.
Figure 2. Which is the strategy followed in this specific example? Please, clarify. Moreover, ‚The suggested wavelength offsets of -1 to 2.5 nm are relatively large and reflect the simplifying assumption of a spectrally flat initial dispersion curve in the wavelength calibration.‘… This should be clearly included in the main text.
L200 - This is the same spline formulation that considers the same location and number of knots used in the retrieval, right? If not, there might be a problem of consistency. Please, clarify.
L202 - In Figure 2, you clearly showed that there is quite large discrepancy among scenes. I'm not sure whether the methodology is stable enough to rely on these results and assume that all of them fluctuate around a true value.
L209 - Major concern: You talk about the first performance metric, but only the plots were shown, not the metric itself. Please, discard talking about this metric if it's not going to be used later. Moreover, I wonder why a range of number of knots is considered instead a single number of knots. Finally, the results aren't really capturing the spectral variability. In Thompson et al., (2024), this variability seems to be better captured, at least in comparison to the laboratory-measured spectral dispersion. Which is the added value of these experiments then? Please, clarify
L217 - Wasn't this already the point discussed in Thompson et al., (2024)? Moreover, results from this metric are not shown. For transparency's sake, I'd include some table / plot putting together all the information provided in the text.
L227 - How can you be sure that the chi-2 is only reflecting how good the spectral calibration is? It's true that that changes are related to the different model configurations, which totally depend on the the number and position of the knots. But, maybe the lower chi-2 comes from setting forward models that can better handle differences in reference to the measurement by any other reason? Please, explain more thoroughly why you can really isolate the dispersion offset as the only cause for the variation in chi-2. This is key for your results.
Figure 3. B-Spline? Please, clarify.
Figure 5. Then, this should be expressed with other symbol, not Delta_lambda
L244 - Main concern: Why does this pattern being reproduced but the numbers are totally changed? For instance, at wavelength ~1000 nm there is a difference of the order of 0.1 nm in Figure 3, but there is a large difference of 2.5 nm in this figure. This difference is too large and shows lack of consistency with previous results. Please, clarify.
L258 - Main concern: It's hard to believe this claiming under the inconsistencies I pointed out before. Hopefully, I'm missing several points that I currently don't see.
L268 - I don't think this is true. There are other points in which there is similar discrepancy.
L275 - Here you should repeat that Scene 6 is the one that is the one situated more temporally far away in reference to the other scenes. Moreover, quantitavely which were the differences in reference for Scene 6. In Figure 2, there is not significant difference. Please, clarify.