the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Temporal evolution of anthropogenic carbon in the Irminger Sea between 2011–2021: from accumulation to stabilization
Abstract. The ocean mitigates anthropogenic climate change by absorbing roughly a quarter of anthropogenic CO2 (Cant), at the cost of ongoing ocean acidification, threatening marine life. The North Atlantic Ocean exhibits the highest ocean storage capacity of Cant per unit area. The subpolar North Atlantic is subject to a large seasonal to decadal variability that might impact this Cant storage. Here, we investigate the monthly Cant evolution over 2011–2021 and its relationship with regional ocean dynamics. We combined Argo-O2 observations, neural networks, and the back-calculation φCTO method to derive a monthly time series of Cant inventory in the Irminger Sea. In the top 2000 dbar, Cant inventories increased at a rate of 0.96±0.18 mol m-2 yr-1 (1.28±0.24 % yr-1) over 2011–2021. However, this increase was not consistent over time, with no net increase observed between 2016 and 2021. Over 2016–2021, when Deep-Argo data became available, the Cant inventory below 2000 dbar increased by 0.53±0.12 mol m-2 yr-1 but depicted a transition from positive to negative accumulation rates around mid- 2019. The monthly resolution of Argo-O2 data reveals that the largest changes in Cant occur during winter. It underscores the importance of wintertime processes such as deep convection that are not captured by summer cruises. This enhanced temporal resolution also enables robust detection of shifts in storage rates. This study highlights the ability of Argo-O2 data to resolve interannual to decadal Cant variability, complementing the interpretation ship-based measurements, which are more accurate but temporally sparse.
Status: open (until 23 Aug 2026)
- RC1: 'Comment on egusphere-2026-3477', Anonymous Referee #1, 02 Aug 2026 reply
-
RC2: 'Comment on egusphere-2026-3477', Anonymous Referee #2, 07 Aug 2026
reply
In this paper, DIC and TA are reconstructed from P, T, S, O2, time in situ data (mostly ARGO-O2 floats, but which miss the northern tip of the Irminger Sea based on Figure 1) by neural networks (the reconstruction is separately done for nutrients…) (Asselot et al., 2024), Then, back calculations are used to estimate Cant (Perez et al., 2008), similarly to what is done on the OVIDE cruise data. Nice and large use done on the Argo-O2 network of floats in the North Atlantic subpolar gyre to extend view in space and time that could be reached from repeated cruises (such as Ovide cruises, Bajon et al (2026)).
The main paper results, which extend on what is presented in Asselot et al (2024), are
1: a small negative trend in Cant near 3000 m but only from 2016 - 2021 (with a decrease of a positive trend peaking near 1400m t o3000m), but a good Cant of 23.5 at 3000m, 27.1 at 2000 m and up to 37.4 at 1400m. At the surface in 2021 it reaches 53.2 micromol/kg.
2: For the integrated content to 2000m, a peak in 2016, and nearly constant after. These seem also to be present in the Ovide cruises (+ early 2015 with a larger inventory (from Fröb et al, 2016, 2018), which actually is not well captured in Fig.3) with a decrease both in top and lower layer between 2018 and 2021.
3: There are differences between the local trends in CANT along the Ovide sections, and the Irminger-basin average estimate probably outlining some spatial variability at this fairly regional scale that is captured by the neural networks.
DIC (and AT) here is based on the neural networks of Bittig et al (2018). Thus, the time dependence for DIC in these neural networks comes from using the GLODAP2V2 data base with a decimal variable time, (and T, S, O2, which implicitly includes seasonal variability, not explicitly introduced) (CANYON-B + CONTENT, which homogenizes the 4 variables of the carbonate system). The data base GLODAP2V2 in recent years in the Irminger Sea is largely based on some of the OVIDE cruises (not all, as well as a few others including the early 2015 Norwegian cruise, also presented here). Thus, in some ways, I am not too surprised that the multi-annual time evolution of DIC is well captured during this decade (at least below the surface) by this neural network (and CONTENT). My main question is which are the cruises that were included, and what are the ones which are not used in this version of CANYON-B. I don’t know whether there have been any updates on the network since the first paper, and thus in particular whether the more recent OVIDE cruises are or not included . This is important as a large part of the time dependency and trends are based on the period that was fed. To a large extent, I expect (but does not know) that the mroe recent period (after the last data fed int oit) is pure extrapolation. There is a big question on how this time dependency works in the post-training period and whether it is (for the pure Cant part) just some extrapolation with likely larger errors than during the training period. I see no obvious reason why it would work better (for the post-data training part) than any other model, such as the stationary model with scaled forcing adopted, as a comparison. SO, alltogether, I believe there is a need of some discussions, tests, in order to convince the reader
On the other hand, I would be more easily convinced that the part of the variability associated t ochanges in the vertical structure, mixing, ocean circulation that will impact T, S, O2, would be nicely reproduced by the neural network of CANYON-B. Is there a way to separate that part of the response in DIC from the changes induced by the increase in atmospherici pCO2 and surface concentration, that are roughly modeled by the scaled forcing (at least for the vertical integrals).
There might be also an other issue related to the non-explicit consideration of the seasonal varaibility during the training phase (well as argued, it goes with the changes in near surface T, maybe also through the time variable in the model which is decimal). In particular, I worry about how the winter vertical mixing is modeled, as the GLODAP2V2 base is largely based on non-winter data in the Irminger Sea (with the exception of the early 2015 Norwegian cruise).
Minor comments:
- 43 ‘interpretation from ship-based measurements’
- 58: ‘to the ocean bottom’
- 82: what is meant by ‘intensified ocean acidification rates by 7-10% (Curbelo-Hernandez et al, 2024)? Was the ocean acidification rate estimated both in 2009 and 2019 (and over which period each time), or is it an average 2009-2019 compared to estimates from earlier periods (in this case, which one?)
- 126: ‘this study investigates…’
- 165: ‘… with a few additional profiles in the Gulf of Mexico and in the Mediterranean Sea’
- 190-193: uncertainties seem here to be based only on the ones on the T, S, O2 profile data base, whereas the uncertainties estimated in the CANYON-B + CONTENT tools are not taken into account (I realize that these might not be directly compatible, in the way the Bayesian approach is done…). Thus, it seems to me that the uncertainty reported is only part of the total uncertainty. On the other hand, this random uncertainty might be the most relevant here. L. 193 reports Suporting information, but which I did not find when downloading the preprint?
- 216: again, the use of the Monte Carlo method with no precise information on how it is implemented (how are the uncertainties correlated in the vertical? Maybe this is in the Supp Mat that I have not read?)
Section 2.6 l.238-254. I appreciate the use of the scaling method, as a way to weigh non-local effects, changes of circulation, convection, etc… compared to this very simple local accumulation assumption. I would add that the assumption of the exponential change is not only related to the 10 years investigated here, but has to be valid further in time, due to the average age of the different deep and intermediate water masses (but usually not that old in the Irminger Sea).
Fig. 3: I don’t understand the difference with OVIDE data in Panel c, which I would have expected to be the sum of the differences in panels a and b (for example, for the 2016 Ovide data: by eye I would have thought that the sum of a and b would yield a difference with blue curve on the order of 2-3 mol/m-2, much less than what is plotted on c and is close to 10). Is something else done?
Also, on Fig. 3, is there an explanation for the large difference in early 2015 with the Norwegian cruise data (I am aware that on top panel it is the TDD method used and not the f0CT one, but the lower panel for integrated value show a similar change, and this does use the TDD method?). Is it an issue of spatial locations of the April 2015 stations emphasizing the convection event, and not so much a wider average, or is it that the neural network does not well capture the increase in Cant during convection (and in which case, does one have good confidence in the increase in early 2016 with the neural network?). ALtogether, although the difference is mentioned, it is not always clear when and where the comparison is carried (separating the product averaged along the Ovide line, but then which stations are kept from the early 2015 cruise, or whether it is basin-averaged).
Citation: https://doi.org/10.5194/egusphere-2026-3477-RC2
Viewed
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 71 | 0 | 1 | 72 | 0 | 0 |
- HTML: 71
- PDF: 0
- XML: 1
- Total: 72
- BibTeX: 0
- EndNote: 0
Viewed (geographical distribution)
Since the preprint corresponding to this journal article was posted outside of Copernicus Publications, the preprint-related metrics are limited to HTML views.
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
Major comments
This manuscript creates a monthly dataset of anthropogenic carbon based on Argo float data using a method based on a previous paper (Asselot et al. 2024). While I think a monthly analysis and extension to full depth by utilising Deep Argo in the Irminger Sea is very useful and the Cant uptake analysis of the time range studied is sound, overall, I think this is a missed opportunity to further scrutinise and improve the Asselot et al. 2024 results (which were authored by almost the same team of researchers). Instead, it seems like the same algorithm was applied to the Irminger Sea. Some passages in the discussion, e.g. when the benefits of Argo float data are praised, seem like they belong more to the original paper first outlining the method, not in the re-application of the method. I think it would be helpful to discuss limitations and improvements of the method more. After all, even though uncertainty values are given as just being slightly higher than ship-based estimates (7.5 umol/kg vs. 5.2 umol/kg), Cant estimates are still being calculated from other estimates (i.e. from ESPER_NN and CONTENT), which would make a case for higher scrutiny. There are only very few statements regarding limitations, e.g. L343-344.
The introduction of the main text sounds too much like the previous Asselot et al. 2024 paper, with a few updated numbers and similar wording (e.g. “Since the beginning of the industrial revolution, human activities have emitted large amounts of carbon dioxide (CO2) in the atmosphere via fossil fuel burning” vs. “Since the beginning of the industrial revolution, human activities such as fossil fuel burning and changes in land-use have led to the emission of large amounts of carbon dioxide (CO2) into the atmosphere” in Asselot et al. 2024). Later in the introduction there are more phrases that closely resemble Asselot et al. 2024, for example “finest temporal resolution to date” in L127 (“finest spatio-temporal resolution to date” in Asselot et al. 2024). It would be good to rewrite it and check it carefully. Better yet, include some explanation what you are doing differently. This is only done much later, in L423. For the reader it would be better to make this clear early.
Section 2.1 lacks detail on how data was chosen and why those choices have been made. Why was “probably good” data included, given that ARGO floats already suffer from reduced data quality compared to ship-based data? Why was the data last accessed in 2021? Overall, this seems to be an update of Asselot et al. 2024 to include more data points, but if that’s the case it would be good to include more recent data – or at least explain why this wasn’t done. Why was the old GLODAPv2.2020 used and not a newer version (GLODAPv2.2021, GLODAPv2.2023 or more recently GLODAPv3), where more cruises are included and more corrections applied? Wouldn’t it be better to retrain an updated dataset? This is also noted in Asselot et al. 2024, where they state, “Therefore they must be trained with new data as the ocean changes”. Also, data accuracies are listed, but there is no discussion about whether this is acceptable for the task.
For section 2.2, I would rather that you shorten the neural network explanation and instead give more context on your workflow, i.e. the method of Asselot et al. 2024. Even if you reuse existing neural networks, the details about the neural network training are insufficient. It would be great to include some plot showing neural network performance – especially for the Irminger Sea. I believe it’s beneficial to be open about where the neural network is struggling and where we have the most confidence.
Minor comments:
L20: Are the amount of Deep Argo floats, compared to normal Argo floats, sufficient to simply expand from 2000 to full depth? Have you noticed any transition (of data or uncertainties) at 2000 m due to the changing data density?
L32: Perhaps you should make clear in the abstract that you’re using existing neural networks – some readers will only read the abstract first and may conclude that you have developed new ones. See also L155
L51-53: Despite this being a plain language summary, it would be more accurate to say that Cant can’t be measured directly and needs to be estimated from different methods. I know you use “studied” to create some ambiguity, but it still sounds like research ships go out and measure Cant directly. Just a slight rephrasing is enough
L63-65: This first sentence sounds a lot like Asselot et al. 2024
L67: You write about further consequences of increased Cant uptake further down (from L78), perhaps that’s where the sentence about the effective radiative forcing would fit better
L69: in the most recent version this estimate rises to 29% for the Global Carbon Budget 2025 - but you should still throw in an “approximately”
L76: I would specifically add “observation-based” to those products, to distinguish them from the models. Also, you can cite Friedlingstein here again as this is the main review (or “current state”) of models vs. observations and the difference between them (figure 10 in Global Carbon Budget 2025)
L87: I would say “estimate”, not compute. That is also what was done in Asselot et al. 2024 – at the same time, you again risk sounding a lot like Asselot et al. 2024
L89: You start with the abbreviation without introducing it first. It should be “Dissolved inorganic carbon (DIC)”. Also, perhaps it would be useful to include a bit more detail on the types of carbon in the ocean. First, only Cant is talked about, later, DIC is mentioned without explanation. If you use Cant, you should also make clear what Cnat would be and that DIC holds both. This can be done in a very short way, but would provide the reader with more context.
L89-91: Include a reference – especially because this wording could also be applied to the C* method (Gruber et al. 1996)
L99: there are more, e.g. the eMLR(C*) method of Clement and Gruber 2018
L110-113: Cant is not measured by ships directly, it’s been estimated using data from these time series
L122: You should make clear that this is what you’re attempting here – this sounds like getting Cant from neural networks is the default method
L127: Sounds a lot like “finest spatio-temporal resolution to date” in Asselot et al. 2024
L133: Why was the last access in 2021? Wouldn’t it be helpful to include more data in the training?
L133-134: Why is it relevant to say that real-time profiles have been replaced by delayed mode? This sounds like you were using real-time profiles at some point. Why? Wouldn’t it be better to have more QC prior to neural network estimation?
L136: Why do you include “probably good” QC flags?
L141: Not necessary to repeat the choice of years again
L155: It would be good to highlight in a sentence explaining that you’re using “existing neural networks (i.e., ESPER_NN, CANYON-B and CONTENT)” - like in Asselot et al. 2024, where the language around using existing algorithms is much more clear. If a reader hasn’t read Asselot et al. 2024 first, it initially seems like you are applying new neural networks.
L165: Why was GLODAPv2.2020 chosen? There are 3 more recent versions than this available. Also, was the complete, global dataset used, or just a subset?
L166-169: It would be helpful to add a few more error metrics like RMSE and bias, and more detail on how the training was evaluated – there are no plots or tables on this
L167-169: It is not immediately clear where those uncertainties apply to, how they were calculated and what they mean for your own estimates. Refer to the supplementary, if that’s where they come from.
L168: A bit more context on the workflow would be good. Reading Asselot et al. 2024 first lets you follow everything much better, but if a reader sees only this manuscript it is confusing why DIC and alkalinity is thrown in here
L193: Specify where in supporting information
L240-241: It would be helpful for the reader to explain a bit what this does, instead of just citing the associated manuscript
L307-308: Could this be an artifact from your method?
L325: Caption refers to Fröb et al. 2018, but figure refers to Fröb et al. 2016
L339: While the higher spatio-temporal resolution is an advantage, I would not necessarily say that this is more robust – after all, you are working more with neural network-generated estimates compared to ship-based data (moreover: estimates based partly on estimates)
L371-373: Yes, they align closely, but can you still say something about where they don't match? What may create differences? Is there anything your method has been lacking? I'm not doubting these curves, I just wish that there'd be more of a discussion regarding neural network performance and limitations
L384: How do the values compare if you do the same?
L423-424: These aims would’ve been good to state at the beginning
L450-451: How do you explain it? Again, could this be an artifact from your method? Was there anything suspicious in the individual components (e.g. the estimates from ESPER-NN and CONTENT?)
L465-475: This paragraph makes sense, and I can see the advantages – but it sounds like this belongs more in the original Asselot et al. 2024 paper. You are not creating a new method here
Supporting information
L34 This is partly repeated from the main text. More importantly, the main text states “accuracy of Argo data is 0.02°C for temperature” but in the supplementary it’s 0.002°C. Which is true?
L70-71 “The standard deviation fluctuates between ±2.7 µmol kg-1 and ±16.8 µmol kg-1“ it would be useful to know where the values have been higher/lower