Impact of ocean-forcing temporal resolution on biogeochemical dynamics and forecasting in the Mediterranean Sea
Abstract. The temporal forcing frequency of ocean dynamic fields used to drive transport–biogeochemical models in off-line coupling configurations can critically affect model performance, especially when mesoscale and high-frequency biogeochemical variability are resolved. This study demonstrates the benefits of using high-temporal-frequency ocean dynamics to force biogeochemical simulations within the Mediterranean Sea forecasting system of the EU Copernicus Marine Service. The OGSTM–BFM transport–biogeochemical model was forced with 6-hourly and 24-hourly averaged outputs from the Mediterranean Forecasting System (MedFS, based on NEMO–WW3-OceanVar) with particular attention to the role of vertical transport processes in regulating primary productivity. Validation against BGC-Argo and satellite observations revealed that the 6-hourly forcing improved the timing of phytoplankton blooms across several Mediterranean subbasins, as well as the spatial distribution of chlorophyll and nutrient concentrations—most notably in the western Mediterranean and Ionian Sea during winter. Enhanced sub-daily advection in winter and intensified mixing in summer were identified as the dominant processes driving these biogeochemical responses, leading to better representation of winter surface blooms and a broader deep chlorophyll maximum layer in summer. Although basin-scale primary productivity patterns remain broadly consistent, local productivity intensity varies in response to the altered vertical transport regime.
This manuscript seeks to evaluate how and whether higher frequency temporal forcing of a BGC-model in the Mediterranean Sea improves the representation of BGC variables, especially with respect to phytoplankton and the processes that affect their growth. This is overall a useful thing to do, and I appreciate the explicit investigation of links to the physics and how they are resolved. However, there are a few points, outlined below, that puzzle me with respect to the evaluation. The way the results are currently presented, I am not convinced that they support all the claims, especially since statistical significance is not reported. I hope that my suggestions for improvements are useful.
Written text:
The Introduction and Methods are overall well written, so I don’t have many comments for these. The discussion of physics results also makes sense to me (though I am not an expert in this field). However, the Results section is overall a bit repetitive and formulaic, and the Discussion is poorly organized and would benefit from sub-headers and a clearer focus. For the Results section, a structure along these lines could be considered:
Analysis:
Regarding float data: My main “sticky point” with the results is that the 3 BGC floats are under-utilized in their ability to be “points of truth” for the two model runs. For example for Figure 7, why not calculate the difference between float data and the 6H run, and also float data and the 24 H run? That appears to be an important step to me, because a priori I am not ready to accept that the 6H run is “better” than the 24H. The way the abstract and discussion are currently written, they really call for this kind of reality check. With such a figure you can pin-point where the improvements are, e.g. in the vertical distribution of properties. In my understanding (and based on my reading of the abstract), the ultimate question shouldn’t be “which run resolves more variability”, but “which run is more realistic”? These two concepts get a bit muddled in the current manuscript. I also suggest showing the actual float data the same way that model data is shown in Figure 6.
Regarding Table 2: I’m not sure that these numbers are meaningful. An “annual mean correlation” is likely largely driven by the seasonal progression, which you’d hope even crude models get right. The 6H run is after resolving small-scale processes (in time, and also space possibly), so showing such a crude correlation is not very meaningful. The observed differences between the 6H and 24H run are also rather small. Are they statistically significant? A more meaningful table would report correlations (between model runs and Argo float data) for the BGC metrics described in Table 1.
Regarding npp: I note that the npp estimates from the model are never directly compared to observational data (from satellites or BGC-Argo floats). This may be a deliberate choice, as the observational data are not without flaws, but it would be useful to be a bit more explicit about this. Direct comparisons are being made for chlorophyll and nitrate, so not taking this direct comparison step for npp is worthy of an explanation.
Statistic significance – is rarely mentioned when you make comparisons. More detailed comments in the below, but overall, this is something that should be reported when making comparisons between things like RMSD, model to model.
More specific comments:
Given how much the results and discussion rely on differences (between the two models, and between models and Argo data): it would make sense to define parameters for these differences (call them “delta” or something along those lines), that way it’s always clear what was subtracted from what. This information is currently lacking for quite a few figures.
Frequently, when “improvements” or changes are discussed (e.g. p19, l130), it would be informative to also mention the %change. Being told that npp increases in one run by 3 g C m-2 yr-1 doesn’t mean much to the reader if the context isn’t given. This is also very much true for Figure 9, where the changes appear miniscule to the eye. How big are the differences in %? And are any of them statistically significant?
Figure 6: npp is not a concentration, and the white and black lines in the section plots are not defined
Figure 7: Nitraclines are hard to see and don’t always seem to align with what is shown in Figure 6; they are also missing from the third row of section plots (oxygen)
Figure 11 and its discussion: the numbers discussed don’t align with what is in the Figure. So I can’t comment on this part, but this obviously needs tidying up.
All figures would benefit from bigger axes labels
Tables 3-5: suggest putting these in the supplementary and only pull out some meaningful metrics that are worth discussing in more detail (maybe that will be covered with my suggestions for Table 2 above). Also, are any of the RMSD comparisons statistically significant? I’m also unsure that “vertical correlations” between profiles are useful metrics, as these will be largely driven by profile shapes, which you’d hope are overall in agreement.