Phishing Email Detection With Machine Learning Algorrithms
DOI:
https://doi.org/10.66021/Keywords:
cybersecurity, machine learning, email classification, phishing email detection, and zero-day phishingAbstract
Phishing emails are considered one of the greatest serious cybersecurity threats due to the exploitation of human vulnerabilities to leak private data or make compromises to an online infrastructure. Four conventional algorithm include NB, LR, DT, and RF are used in the suggested lightweight ML architecture for phishing email identification. TF-IDF based feature extraction, text normalization, and methodical data pretreatment are all included. A collection of phishing emails with both positive and negative class labels was used to train and evaluate individual algorithms. The accuracy, precision, recall, F1 score, ROC curve, and confusion matrix will all be used to measure each model's performance. With accuracy levels exceeding 98%, outstanding precision, and recall, experimental results demonstrate that RF and LR are the best evaluated models. While the NB model performed good at a very cheap computing cost, the Decision Tree showed good scores and a modest accuracy that was still interpretable. The comparison analysis further demonstrates that the proposed lightweight models performance competitiveness with deep learning based approaches with significantly lower computational resource consumption. Overall, this research highlights the fact that classical machine learning algorithms, when combined with effective preprocessing and TF-IDF representation, form an efficient and practical solution for phishing email detection to suit real time cyber-security applications.