Ziauddin Ursani

dblp:44/10683 · DBLP profile ↗
← Back
2ranked-venue papers in the field
1as first author
2since 2021 · last 2025
0000-0002-2972-4024ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2 (1 first)
YearPublicationVenuePosition
2025 Evolutionary Train-Test Split for Hierarchical Monte Carlo Ensemble
abstract
In machine learning, splitting data into training and test sets is usually achieved using random stratified sampling, in which classes are proportionally divided into two subsets. Other methods also consider feature-aware criteria, and some of those methods claim to have achieved optimal split of minimised variance. We do not advocate aiming to achieve an optimal split or minimise variance since this would be counterproductive for ensemble methods, where the diversity of the training set is desired. Ensemble methods achieve diversity through bagging and boosting schemes. In the recently introduced Monte Carlo ensemble approach, diversity can be maintained through random stratified sampling without using bagging or boosting methods. This work introduces a feature-aware split that retains the diversity of the ensemble. To this end, we propose an evolutionary algorithm that starts with an entirely random population and aims at objectives of proportional class-representation and minimisation of the normalized mean error rather than minimisation of variance. The proposed data-split method is tested on three different models used within a hierarchical Monte Carlo ensemble. The results show that the method positively affects the predictability performance when applied on two domain-specific material science datasets and a collection of 38 general machine learning datasets.
Ziauddin Ursani, Dmytro Antypov, Katie Atkinson, Matthew S. Dyer, Matthew J. Rosseinsky, Sven Schewe, Ahsan Ahmad Ursani, Andrij Vasylenko
BDCAT1
2021 AMoC: A Multifaceted Machine Learning-based Toolkit for Analysing Cybercriminal Communities on the Darknet
abstract
There is an increasing demand for expert analysis of cybercriminal communities. Cybercrime is continually becoming more complex due to the rapid development of digital technologies, on the one hand, in new types of criminal activity, such as hacking, distributing malware and DDoS attacks, and on the other hand, in digitised forms of more traditional crimes, such as email scams, phishing, identity theft, and cryptographically secured black markets. Tackling this broad array of behaviour requires tool support for multi-disciplinary investigations, and a connecting framework that can adjust flexibly to changes in the populations being studied. In this work, we present AMoC, a multi-faceted machine learning toolkit that combines structured queries, anomaly detection, social network analysis, topic modelling and accounts recognition to enable comprehensive analysis of cybercriminal communities and users. The toolkit enables the extraction of findings regarding the motivations, behaviour and characteristics of offenders, and how cybercriminal communities react to interventions such as arrests and take-downs. In our demonstration, the toolkit is deployed to analyse over 150,000 accounts from 35 underground marketplaces.
Claudia Peersman, Matthew Edwards 0001, Ziauddin Ursani, Awais Rashid
IEEE BigData4