the Creative Commons Attribution 4.0 License.
the Creative Commons Attribution 4.0 License.
Review article: Global flood research across four decades: An analysis of 57,474 research articles based on a large language model
Abstract. The proliferation of research literature has resulted in increasingly fragmented information and diversified knowledge structures, making it increasingly difficult to develop a systematic understanding of disciplinary research systems. This study employed a large language model to screen and organize flood research articles from 1985 to 2024, provided visual spatiotemporal insights into the quantity and quality of publications, contributions of institutions and countries, evolution of keywords, research types and topics, as well as the geographic characteristics of river basin and urban units. Findings indicate that annual publications on flood studies now exceed 5,500, with the average number of authors and institutions per publication increasing from 2.1 to 4.8 and from 1.4 to 2.6, respectively. Chinese and American institutions lead in publication quantity, while European and American institutions dominate in research influence. The disciplines of climate and remote sensing have gradually gained an undeniable influence in flood research field. The evolution of keywords focus traces a conceptual progression from “ecological-hydrological system interactions,” through “hydrological and hydrodynamic processes analysis” and “flood disaster prevention and management,” to “flood prediction and risk assessment under climate change.” River flood remains a consistent focus (averaging 55 % of research), while urban flood has seen a notable rise in attention (increasing from 8 % to 15 % over the past decade). Research topics concentrate on management, simulation, monitoring, and risk assessment. River basin flood research has evolved from Mississippi River Basin leadership to bipolar dominance with the Yangtze River Basin. Urban flood research has shifted from leadership by Texas in the United States to a dual-core structure with Guangdong Province in China. This study aims to advance the application of large language models in flood research literature review, thereby enabling a more systematic and efficient understanding of the comprehensive characteristics of the field’s development.
- Preprint
(5271 KB) - Metadata XML
- BibTeX
- EndNote
Status: open (until 16 Sep 2026)
-
RC1: 'Comment on egusphere-2026-3343', Anonymous Referee #1, 29 Jul 2026
reply
-
AC1: 'Reply on RC1', Hao Wang, 14 Aug 2026
reply
Thank you very much for your valuable suggestions. We have answered your questions one by one in the reply below (block color). However, the current status of the journal does not allow us to modify the manuscript or submit the manuscript as supplementary material in the discussion. Therefore, we will directly mention or show our changes, such as modified pictures, text to be added and comparison with the old text, etc. For more details, please refer to our the PDF attachment.
-
AC1: 'Reply on RC1', Hao Wang, 14 Aug 2026
reply
-
RC2: 'Comment on egusphere-2026-3343', Bernard Twaróg, 21 Aug 2026
reply
The study appears valuable, very comprehensive, and innovative. An interesting element is the use of LLMs to classify tens of thousands of publications. The study has several important strengths: the scale of the analysis is impressive; combining LLMs with bibliometrics represents a valuable methodological direction; the multidimensional nature of the analysis is a strong point; dividing the 40-year study period into five-year intervals is practical; the authors provide the prompts used in the analysis; the integration of semantic location extraction with GIS analysis is particularly interesting; and the study effectively demonstrates the increasingly interdisciplinary nature of flood research. The main innovation of the study lies in its scale and in the use of LLMs to semantically enrich bibliometric analysis. The figures and tables are clear and well described. The language is appropriate and generally correct.
However, the manuscript contains several inconsistencies. The first concerns the number of publications. The authors state that they retrieved 150,000 publications from WoS and subsequently used an LLM to select 57,474 articles closely related to flood research. However, the wording in Section 2.2 suggests that the 57,474 records were obtained “after deduplication.” This would imply that no records were excluded by the LLM-based screening process. This discrepancy requires clarification.
The study resembles an LLM-assisted bibliometric analysis more than a conventional systematic literature review. Several elements typically expected in a systematic review are missing, including a detailed description of the publication screening and selection process, exclusion criteria applied at successive stages, and an assessment of the quality of the included studies.
Another concern relates to the concept of an “objective knowledge structure.” The adopted structure is, at least to some extent, imposed by the authors. For example, the LLM is required to assign publications to predefined categories: management, simulation, monitoring, and risk assessment. Similarly, nine flood types are defined in advance. Therefore, the resulting knowledge structure cannot be regarded as entirely objective, as it is partly determined by the classification framework established by the authors.
The authors specify the Gemma 3-27B model and provide the prompts used; however, several details necessary for full reproducibility are missing. For example, the inference parameters of the model are not reported. Such information is important when using LLMs because, even with identical prompts, model outputs may not be completely deterministic.
There is also an issue with the mutual exclusivity and conceptual consistency of the flood-type definitions. For example, a tsunami should not be treated as a cause of a storm surge. Moreover, tsunami is simultaneously defined as a separate category, creating a logical classification conflict. Similarly, mountain flood is defined using mudslide as an example. A mudflow or landslide is not simply a type of mountain flood.
The definition of urban flood refers to flooding caused by insufficient drainage capacity during rainfall in urban areas. However, the authors subsequently analyse “urban flood research” in which 71% of the cases are classified as river floods, 18% as mountain floods, and 10% as storm surges. This suggests that the authors may be conflating two different variables: the geographical setting of the study and the flood mechanism/type. These concepts should be clearly distinguished.
The proposed influence index, defined as the mean number of citations per article, does not appear to be a sufficiently robust measure. One major problem is that older publications have had considerably more time to accumulate citations. Moreover, comparisons among countries, journals, or institutions without normalization for publication year, research field, and publication age may lead to misleading conclusions. It should also be emphasized that citation impact is not equivalent to scientific quality. An article may be highly cited because it is fundamental, controversial, methodological, widely used, or even because it contains an important error. Citation counts measure bibliometric impact rather than scientific quality directly.
The study relies exclusively on Web of Science. Consequently, the results primarily describe global flood research as represented in Web of Science, rather than the entirety of global flood research. This limitation should be explicitly acknowledged when interpreting and generalizing the findings.
There are also errors in the prompts. For example, the location prompt contains the example “North America, USA, London,” which is geographically incorrect if it refers to London in the United Kingdom. Furthermore, in the prompt concerning countries, the authors provide the examples “China, American, England, Japan.” Here, “American” is not a country name, while “England” is not equivalent to the sovereign state of the United Kingdom. Such inconsistencies may affect the standardization and reliability of the extracted data and should be corrected.
Overall, the study has considerable scientific potential. After addressing the above comments, clarifying the identified inconsistencies, and providing the necessary methodological details, the manuscript should meet the requirements for scientific publication.
Citation: https://doi.org/10.5194/egusphere-2026-3343-RC2
Viewed
| HTML | XML | Total | BibTeX | EndNote | |
|---|---|---|---|---|---|
| 181 | 48 | 18 | 247 | 23 | 20 |
- HTML: 181
- PDF: 48
- XML: 18
- Total: 247
- BibTeX: 23
- EndNote: 20
Viewed (geographical distribution)
| Country | # | Views | % |
|---|
| Total: | 0 |
| HTML: | 0 |
| PDF: | 0 |
| XML: | 0 |
- 1
The study relies on Gemma 3-27B to identify relevant articles and extract study locations, flood types, topics, methods, institutions, and countries. Because these outputs form the basis of the results, the authors should explain how their reliability was verified.
The Introduction discusses earlier LLM-assisted reviews, including a study analysing 310,000 hydrology publications. The authors should explain more clearly how the present manuscript differs from previous large-scale hydrological literature reviews.
The definitions of several flood categories require clarification. For example, coastal storm surge, tsunami, ice flood, snowmelt flood, and mountain flood may overlap. The prompt also describes the task as multi-label classification but instructs the model to avoid multiple categories whenever possible.
Some interpretations connect publication patterns directly to policies, disasters, economic development, and major infrastructure projects. For example, the manuscript links research changes to the Rio Earth Summit, the EU Water Framework Directive, the Three Gorges Project, Hurricane Katrina, and the Sponge City initiative. These explanations may be reasonable, but the analysis mainly demonstrates temporal associations. The wording should therefore be more cautious. Expressions such as “may have contributed,” “coincided with,” or “could partly explain” would be more appropriate than direct causal statements. The terms “economically driven basins” and “religiously driven basins” should also be reconsidered. The description of the Indus and Ganges as “religiously driven” is broad and may oversimplify complex hydrological, political, economic, and cultural conditions. More neutral and evidence-based terminology is recommended.
The conclusion mainly repeats the descriptive findings. It should more clearly explain how the results can guide future flood research.
The manuscript contains valuable maps, networks, word clouds, and multi-panel figures, but several labels are small and difficult to read. The geographical figures in particular contain many panels, abbreviated locations, small legends, and crowded labels.