Document Type : Research Article
Authors
1
MSc graduated, Department of Water Engineering, Faculty of Agriculture, Ferdowsi University of Mashhad, Mashhad, Iran
2
Associate Professor, Department of Irrigation and Drainage Engineering, Faculty of Agriculture, University of Tehran, Karaj, Iran
Abstract
Introduction
Floods, as one of the most destructive natural disasters, cause significant human and financial losses worldwide every year. With climate change and urban development, the need for accurate and efficient tools for zoning flood-prone areas is increasingly felt. Accurate flood zoning plays a vital role in crisis management, land use planning, and reducing the vulnerability of communities (Saber et al., 2023). This study aims to evaluate and compare the performance of two advanced machine learning algorithms, namely Random Forest and Support Vector Machine, in zoning flood-prone areas in the Navrud watershed in Guilan province(Fig.1).
Methodology
In this study, a set of factors affecting flood occurrence, including distance from the watercourse, drainage density, flow direction, slope, slope direction, precipitation, elevation, land use, and geology, were used as input variables(Fig.2& Fig.3)Historical flood occurrence data were also used to train and validate the models. To prepare the data, different information layers were first prepared in a geographic information system (GIS) environment and then used in an integrated manner to analyze machine learning models. Random Forest (RF) Algorithm: This algorithm is an ensemble learning method that provides stronger predictions by building a large number of decision trees and combining their results. RF model ability to handle outliers, nonlinear features, and the lack of need for specific statistical assumptions make it suitable for flood zoning problems. Support Vector Machine (SVM) Algorithm: SVM is a powerful algorithm for classification problems that works by finding an optimal hyperplane that best separates the data. The use of different kernel functions (such as RBF) allows SVM to model complex nonlinear relationships between input variables and flood occurrence. To validate the models, the data was divided into two parts: training (70%) and testing (30%). The performance of the models was evaluated and compared using metrics such as Accuracy, Sensitivity, Specificity, and Area Under the ROC Curve (AUC).
Results and Discussion
The results of this study showed that both RF and SVM algorithms have a high ability in zoning flood-prone areas. However, a more detailed comparison of the models performance indicated that: drainage density, geology, land use, and elevation are the most important parameters affecting flood occurrence in the Navrod basin, respectively(Fig4). The SVM model with an AUC value of 0.97 is more accurate than the RF model with an AUC value of 0.93 in determining flood-prone areas. Also, the SVM model, with RMSE and R2 evaluation statistics, is more accurate than the RF model with values of 0.26 and 0.89, respectively (Table 1). The downstream areas and near the outlet of the watershed are more sensitive to floods due to their lower elevation and slope and the location of the rivers confluence. The flood zoning maps produced by both models showed a good agreement with the historical flood occurrence data(Fig7&Fig8 & Fig9).
Conclusion
The findings of this study confirm the effectiveness of Random Forest and Support Vector Machine algorithms as powerful tools in the zoning of flood-prone areas. Both models can be considered as suitable alternatives to traditional methods in flood zoning studies and can greatly assist decision-makers in adopting preventive measures and reducing the destructive effects of floods. However, the SVM support vector machine model can outperform Random Forest in scenarios with high-dimensional data, smaller data with a specific structure, and the need for models with optimal separation margins.
Keywords
Subjects