Javascript must be enabled to continue!
Machine-Learning Classification Model and Tools for Real-time URL Phishing Detection
View through CrossRef
Abstract
Phishing attacks are considered a significant cybersecurity concern, employing deceptive tactics to entice individuals into engaging with counterfeit websites. These malicious pages are skillfully designed replicas of legitimate platforms, aiming to collect sensitive data like usernames, passwords, banking credentials, and other personal details. This study focuses on phishing via Uniform Resource Locators (URLs) and investigates the potential of machine learning to identify such deceptive websites based on their behavior and URL attributes. To accomplish this, the work introduces and demonstrates two key tools; one for dataset creation and the other for URL classification.Machine learning has already shown its effectiveness in identifying phishing attacks from URLs, though there are still some obstacles to be overcome, such as the need for vast quantities of high-quality training data and the requirement to keep up with the constantly changing tactics employed by phishing attackers. The integration of the proposed tools in a web browser plugin is supposed to enable real-time URL analysis within web browsers, enhancing the system's effectiveness against phishing attacks and hence improving user experience.Using a self-collected dataset of 46,000 URLs, several machine learning algorithms were trained and tested including support vector machine (SVM), XGBoost, decision tree, and random forest algorithms. Among these, XGBoost model achieved an impressive classification accuracy of 96%, F1-Score of 96.7%, Recall of 96.6% and Precision 96.9% after assessing various permutations of hyperparameter values using the grid search procedure. This success underscores the potency of machine learning techniques in bolstering cyber defenses and mitigating the impact of phishing attacks.
Title: Machine-Learning Classification Model and Tools for Real-time URL Phishing Detection
Description:
Abstract
Phishing attacks are considered a significant cybersecurity concern, employing deceptive tactics to entice individuals into engaging with counterfeit websites.
These malicious pages are skillfully designed replicas of legitimate platforms, aiming to collect sensitive data like usernames, passwords, banking credentials, and other personal details.
This study focuses on phishing via Uniform Resource Locators (URLs) and investigates the potential of machine learning to identify such deceptive websites based on their behavior and URL attributes.
To accomplish this, the work introduces and demonstrates two key tools; one for dataset creation and the other for URL classification.
Machine learning has already shown its effectiveness in identifying phishing attacks from URLs, though there are still some obstacles to be overcome, such as the need for vast quantities of high-quality training data and the requirement to keep up with the constantly changing tactics employed by phishing attackers.
The integration of the proposed tools in a web browser plugin is supposed to enable real-time URL analysis within web browsers, enhancing the system's effectiveness against phishing attacks and hence improving user experience.
Using a self-collected dataset of 46,000 URLs, several machine learning algorithms were trained and tested including support vector machine (SVM), XGBoost, decision tree, and random forest algorithms.
Among these, XGBoost model achieved an impressive classification accuracy of 96%, F1-Score of 96.
7%, Recall of 96.
6% and Precision 96.
9% after assessing various permutations of hyperparameter values using the grid search procedure.
This success underscores the potency of machine learning techniques in bolstering cyber defenses and mitigating the impact of phishing attacks.
Related Results
PUMMP: Phishing URL Detection using Machine Learning with Monomorphic and Polymorphic Treatment of Features
PUMMP: Phishing URL Detection using Machine Learning with Monomorphic and Polymorphic Treatment of Features
Phishing scams are increasing drastically, which affects Internet users in compromising personal credentials. This paper proposes a novel feature utilization method for phishing UR...
Anti-Phishing Technologies and Tools
Anti-Phishing Technologies and Tools
Phishing continues to be one of the most common and effective forms of cyber security threats and involve deception of users thereby getting them provide unauthorized individuals w...
Phishing Cyber Security Threats
Phishing Cyber Security Threats
Phishing is a growing threat in the realm of cybersecurity, where cybercriminals use various phishing techniques to steal sensitive information from individuals and organizations. ...
Spear-Phishing in the Wild: A Real-World Study of Personality, Phishing Self-Efficacy and Vulnerability to Spear-Phishing Attacks
Spear-Phishing in the Wild: A Real-World Study of Personality, Phishing Self-Efficacy and Vulnerability to Spear-Phishing Attacks
Recent research has begun to focus on the factors that cause people to respond to phishing attacks. In this study a real-world spear-phishing attack was performed on employees in o...
Intelligent Detection Designs of HTML URL Phishing Attacks
Intelligent Detection Designs of HTML URL Phishing Attacks
Phishing attacks are a type of cybercrime that has grown in recent years. It is part of social engineering attacks where an attacker deceives users by sending fake messages using s...
Woningcorporaties en Vastgoedontwikkeling
Woningcorporaties en Vastgoedontwikkeling
This summary highlights the findings of the PhD-thesis ‘Woningcorporaties en Vastgoedontwikkeling: Fit for Use’ (‘Housing associations and Real Estate Development: Fit for Use?’). ...
AI-powered phishing detection: Integrating natural language processing and deep learning for email security
AI-powered phishing detection: Integrating natural language processing and deep learning for email security
Phishing attacks are major threats to email security and pose challenges, while cyber attackers utilize increasingly sophisticated means to deceive the user and steal away importan...
Knowledge-Grounded LLM-Driven Augmentation via Graph RAG for Phishing URL Detection
Knowledge-Grounded LLM-Driven Augmentation via Graph RAG for Phishing URL Detection
Integrating prior rule knowledge into generative data augmentation remains an open problem when detectors must generalize under distribution shift. URL-based phishing illustrates t...

