Javascript must be enabled to continue!
Concept and Assessment of Machine Learning Models for the Classification of Malicious URLs
View through CrossRef
E-mails: 1ogunjimi.olalekan@phoenixuniversity.edu.ng 2econwubiko@cstp.nasrda.gov 3negedumoses@gmail.com 4stmaster777@gmail.com 5aremo_emmanuel@yahoo.com 6martbell4@gmail.com
ABSTRACT
Identifying safe websites can be challenging as millions of new websites are created daily. Cybersecurity is the practice of safeguarding individuals and organizations against online threats. Phishing attacks are one of the many tactics used by cybercriminals to deceive people into revealing their personal information. In 2022, over 74,000 phishing attacks were reported in Australia alone, resulting in financial losses exceeding $24 million. In various domains such as the detection of financial fraud, and cancer, and the development of chatbots, artificial intelligence (AI) and machine learning have proven to be useful tools. Support Vector Machines and Random Forest are often employed as machine learning models for classification tasks. With the rise in cybercrime, machine learning is crucial for identifying fraudulent URLs, both new and old. The study comparing different machine learning models for classifying malicious URLs found that the Random Forest model achieved the highest accuracy, followed by the Support Vector Machine. Key findings emphasized the importance of balanced datasets, appropriate instance selection methods, and the feature "has HTTP" in achieving accurate results. Further research is suggested to explore additional models and categories of malicious URLs for enhanced cybersecurity measures. In summary, Random Forest outperforms other models in classifying malicious URLs, and balanced datasets and relevant features are crucial for optimal performance. Around 650,000 URLs from Kaggle were used in the dataset for this investigation. Malware, benign URLs, defacement, and phishing were the four categories that made up the dataset. Instance selection techniques (DRLSH, BPLSH, and random selection) written in MATLAB were used to construct three datasets, each including approximately 170,000 URLs. We used SVM, DT, KNNs, and RF as examples of machine-learning models. To train the machine learning models on the malicious URL datasets, the study employed these instance selection techniques. Then, using 16 features and one output feature, it assessed the models' performance. During the hyperparameter tuning procedure, four models with various hyperparameter settings were trained using the training dataset. For each model, the ideal hyperparameter was determined via Bayesian optimization. The system of classification
Keywords: Machine learning, Cyber security, Classification, Malicious URL, Instance selection
Ogunjimi Olalekan L. A., Emmanuel C. Onwubiko. Negedu Moses Ugbedeojo, Elumelu Azuka Andrew, Aremo Emmanuel Akindele & Abhulimen, Martins Enehireba (2025): Concept and Assessment of Machine Learning Models for the Classification of Malicious URLs. Journal of Advances in Mathematical & Computational Science. Vol. 13, No. 3. Pp 1-23
Available online at www.isteams.net/mathematics-computationaljournal. dx.doi.org/10.22624/AIMS/MATHS/V13N3P1
Title: Concept and Assessment of Machine Learning Models for the Classification of Malicious URLs
Description:
E-mails: 1ogunjimi.
olalekan@phoenixuniversity.
edu.
ng 2econwubiko@cstp.
nasrda.
gov 3negedumoses@gmail.
com 4stmaster777@gmail.
com 5aremo_emmanuel@yahoo.
com 6martbell4@gmail.
com
ABSTRACT
Identifying safe websites can be challenging as millions of new websites are created daily.
Cybersecurity is the practice of safeguarding individuals and organizations against online threats.
Phishing attacks are one of the many tactics used by cybercriminals to deceive people into revealing their personal information.
In 2022, over 74,000 phishing attacks were reported in Australia alone, resulting in financial losses exceeding $24 million.
In various domains such as the detection of financial fraud, and cancer, and the development of chatbots, artificial intelligence (AI) and machine learning have proven to be useful tools.
Support Vector Machines and Random Forest are often employed as machine learning models for classification tasks.
With the rise in cybercrime, machine learning is crucial for identifying fraudulent URLs, both new and old.
The study comparing different machine learning models for classifying malicious URLs found that the Random Forest model achieved the highest accuracy, followed by the Support Vector Machine.
Key findings emphasized the importance of balanced datasets, appropriate instance selection methods, and the feature "has HTTP" in achieving accurate results.
Further research is suggested to explore additional models and categories of malicious URLs for enhanced cybersecurity measures.
In summary, Random Forest outperforms other models in classifying malicious URLs, and balanced datasets and relevant features are crucial for optimal performance.
Around 650,000 URLs from Kaggle were used in the dataset for this investigation.
Malware, benign URLs, defacement, and phishing were the four categories that made up the dataset.
Instance selection techniques (DRLSH, BPLSH, and random selection) written in MATLAB were used to construct three datasets, each including approximately 170,000 URLs.
We used SVM, DT, KNNs, and RF as examples of machine-learning models.
To train the machine learning models on the malicious URL datasets, the study employed these instance selection techniques.
Then, using 16 features and one output feature, it assessed the models' performance.
During the hyperparameter tuning procedure, four models with various hyperparameter settings were trained using the training dataset.
For each model, the ideal hyperparameter was determined via Bayesian optimization.
The system of classification
Keywords: Machine learning, Cyber security, Classification, Malicious URL, Instance selection
Ogunjimi Olalekan L.
A.
, Emmanuel C.
Onwubiko.
Negedu Moses Ugbedeojo, Elumelu Azuka Andrew, Aremo Emmanuel Akindele & Abhulimen, Martins Enehireba (2025): Concept and Assessment of Machine Learning Models for the Classification of Malicious URLs.
Journal of Advances in Mathematical & Computational Science.
Vol.
13, No.
3.
Pp 1-23
Available online at www.
isteams.
net/mathematics-computationaljournal.
dx.
doi.
org/10.
22624/AIMS/MATHS/V13N3P1.
Related Results
Prediction and Prevention of Malicious URL Using ML and LR Techniques for Network Security
Prediction and Prevention of Malicious URL Using ML and LR Techniques for Network Security
Understandable URLs are utilized to recognize billions of websites hosted over the present-day internet. Opposition who tries to get illegal admittance to the classified data may u...
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
Selection of Injectable Drug Product Composition using Machine Learning Models (Preprint)
BACKGROUND
As of July 2020, a Web of Science search of “machine learning (ML)” nested within the search of “pharmacokinetics or pharmacodynamics” yielded over 100...
A Machine Learning Based Three-Step Framework for Malicious URL Detection
A Machine Learning Based Three-Step Framework for Malicious URL Detection
AbstractIn order to solve the shortcomings of using blacklist method to detect malicious URLs, such as slow update speed, the research of using machine learning to detect malicious...
LEXICAL PATTERN INTELLIGENCE: A MACHINE LEARNING SYSTEM FOR PREEMPTIVE DETECTION OF MALICIOUS URLS
LEXICAL PATTERN INTELLIGENCE: A MACHINE LEARNING SYSTEM FOR PREEMPTIVE DETECTION OF MALICIOUS URLS
The rapid increase in malicious Uniform Resource Locators (URLs) presents serious cybersecurity risks, enabling phishing attacks, malware dissemination, and financial fraud. Tradit...
Detection of the Malicious URL Using the Language Models
Detection of the Malicious URL Using the Language Models
Abstract
Today, the internet has become an indispensable aspect of modern life, with many organizations providing services through web applications. At the same tim...
Construction of a Cybersecurity Behavior Knowledge Base for Malicious Behavior Analysis
Construction of a Cybersecurity Behavior Knowledge Base for Malicious Behavior Analysis
Facing the surge in malicious behaviors in the network environment, the existing cybersecurity knowledge graph suffers from fragmented security knowledge and limited application sc...
Construction of a Cybersecurity Behavior Knowledge Base for Malicious Behavior Analysis
Construction of a Cybersecurity Behavior Knowledge Base for Malicious Behavior Analysis
Facing the surge in malicious behaviors in the network environment, the existing cybersecurity knowledge graph suffers from fragmented security knowledge and limited application sc...
An Analysis of Malicious URL Detection Using Deep Learning
An Analysis of Malicious URL Detection Using Deep Learning
Considerable progress has been achieved in the digital domain, particularly in the online realm where a multitude of activities are being conducted. Cyberattacks, particularly mali...

