Sunil Aryal

dblp:129/2634 · DBLP profile ↗
← Back
20ranked-venue papers in the field
7as first author
10since 2021 · last 2025
0000-0002-6639-6824ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 15 (7 first)Information Retrieval & Web Search · 4Database Systems & Data Management · 1
YearPublicationVenuePosition
2025 MissDDIM: Deterministic and Efficient Conditional Diffusion for Tabular Data Imputation
abstract
Diffusion models have recently emerged as powerful tools for missing data imputation by modeling the joint distribution of observed and unobserved variables. However, existing methods, typically based on stochastic denoising diffusion probabilistic models (DDPMs), suffer from high inference latency and variable outputs, limiting their applicability in real-world tabular settings. To address these deficiencies, we present in this paper MissDDIM, a conditional diffusion framework that adapts Denoising Diffusion Implicit Models (DDIM) for tabular imputation. While stochastic sampling enables diverse completions, it also introduces output variability that complicates downstream processing. MissDDIM replaces this with a deterministic, non-Markovian sampling path, yielding faster and more consistent imputations. To better leverage incomplete inputs during training, we introduce a self-masking strategy that dynamically constructs imputation targets from observed features-enabling robust conditioning without requiring fully observed data. Experiments on five benchmark datasets demonstrate that MissDDIM matches or exceeds the accuracy of state-of-the-art diffusion models, while significantly improving inference speed and stability. These results highlight the practical value of deterministic diffusion for real-world imputation tasks.
Youran Zhou, Mohamed Reda Bouadjenek, Sunil Aryal
CIKM3
2025 Bias vs Bias Dawn of Justice: A Fair Fight in Recommendation Systems
Tahsin Alamgir Kheya, Mohamed Reda Bouadjenek, Sunil Aryal
ECML/PKDD (1)3
2025 Unmasking Gender Bias in Recommendation Systems and Enhancing Category-Aware Fairness
abstract
Recommendation systems are now an integral part of our daily lives. We rely on them for tasks such as discovering new movies, finding friends on social media, and connecting job seekers with relevant opportunities. Given their vital role, we must ensure these recommendations are free from societal stereotypes. Therefore, evaluating and addressing such biases in recommendation systems is crucial. Previous work evaluating the fairness of recommended items fails to capture certain nuances as they mainly focus on comparing performance metrics for different sensitive groups. In this paper, we introduce a set of comprehensive metrics for quantifying gender bias in recommendations. Specifically, we show the importance of evaluating fairness on a more granular level, which can be achieved using our metrics to capture gender bias using categories of recommended items like genres for movies. Furthermore, we show that employing a category-aware fairness metric as a regularization term along with the main recommendation loss during training can help effectively minimize bias in the models' output. We experiment on three real-world datasets, using five baseline models alongside two popular fairness-aware models, to show the effectiveness of our metrics in evaluating gender bias. Our metrics help provide an enhanced insight into bias in recommended items compared to previous metrics. Additionally, our results demonstrate how incorporating our regularization term significantly improves the fairness in recommendations for different categories without substantial degradation in overall recommendation performance.
Tahsin Alamgir Kheya, Mohamed Reda Bouadjenek, Sunil Aryal
WWW3
2025 Improving out-of-distribution detection by enforcing confidence margin
abstract
Abstract In many critical machine learning applications, such as autonomous driving and medical image diagnosis, the detection of out-of-distribution (OOD) samples is as crucial as accurately classifying in-distribution (ID) inputs. Recently, outlier exposure (OE)-based methods have shown promising results in detecting OOD inputs via model fine-tuning with auxiliary outlier data. However, most of the previous OE-based approaches emphasize more on synthesizing extra outlier samples or introducing regularization to diversify OOD sample space, which is rather unquantifiable in practice. In this work, we propose a novel and straightforward method called Margin-bounded Confidence Scores (MaCS) to address the nontrivial OOD detection problem by enlarging the disparity between ID and OOD scores, which in turn makes the decision boundary more compact facilitating effective segregation with a simple threshold. Specifically, we augment the learning objective of an OE regularized classifier with a supplementary constraint, which penalizes high confidence scores for OOD inputs compared to that of ID and significantly enhances the OOD detection performance while maintaining the ID classification accuracy. Extensive experiments on various benchmark datasets for image classification tasks demonstrate the effectiveness of the proposed method by significantly outperforming state-of-the-art methods on various benchmarking metrics. The code is publicly available at https://github.com/lakpa-tamang9/margin_ood/tree/kais
Lakpa Dorje Tamang, Mohamed Reda Bouadjenek, Richard Dazeley, Sunil Aryal
Knowl. Inf. Syst.4
2025 Handling Out-of-Distribution Data: A Survey
abstract
In the field of Machine Learning (ML) and data-driven applications, one of the significant challenge is the change in data distribution between the training and deployment stages, commonly known as distribution shift. This paper outlines different mechanisms for handling two main types of distribution shifts: (i)Covariate shift:where the value of features or covariates change between train and test data, and (ii)Concept/Semantic-shift:where model experiences shift in the concept learned during training due to emergence of novel classes in the test phase. We sum up our contributions in three folds. First, we formalize distribution shifts, recite on how the conventional method fails to handle them adequately and urge for a model that can simultaneously perform better in all types of distribution shifts. Second, we discuss why handling distribution shifts is important and provide an extensive review of the methods and techniques that have been developed to detect, measure, and mitigate the effects of these shifts. Third, we discuss the current state of distribution shift handling mechanisms and propose future research directions in this area. Overall, we provide a retrospective synopsis of the literature in the distribution shift, focusing on OOD data that had been overlooked in the existing surveys.
Lakpa Dorje Tamang, Mohamed Reda Bouadjenek, Richard Dazeley, Sunil Aryal
IEEE Trans. Knowl. Data Eng.4
2024 Margin-Bounded Confidence Scores for Out-of-Distribution Detection
abstract
In many critical Machine Learning applications, such as autonomous driving and medical image diagnosis, the detection of out-of-distribution (OOD) samples is as crucial as accurately classifying in-distribution (ID) inputs. Recently Outlier Exposure (OE) based methods have shown promising results in detecting OOD inputs via model fine-tuning with auxiliary outlier data. However, most of the previous OE-based approaches emphasize more on synthesizing extra outlier samples or introducing regularization to diversify OOD sample space, which is rather unquantifiable in practice. In this work, we propose a novel and straightforward method called Margin bounded Confidence Scores (MaCS) to address the nontrivial OOD detection problem by enlarging the disparity between ID and OOD scores, which in turn makes the decision boundary more compact facilitating effective segregation with a simple threshold. Specifically, we augment the learning objective of an OE regularized classifier with a supplementary constraint, which penalizes high confidence scores for OOD inputs compared to that of ID and significantly enhances the OOD detection performance while maintaining the ID classification accuracy. Extensive experiments on various benchmark datasets for image classification tasks demonstrate the effectiveness of the proposed method by significantly outperforming state-of-the-art (S.O.T.A) methods on various benchmarking metrics. The code is publicly available at https://github.com/lakpa-tamang9/margin_ood
Lakpa Dorje Tamang, Mohamed Reda Bouadjenek, Richard Dazeley, Sunil Aryal
ICDM4
2024 Missing Data Imputation: Do Advanced ML/DL Techniques Outperform Traditional Approaches?
Youran Zhou, Mohamed Reda Bouadjenek, Sunil Aryal
ECML/PKDD (10)3
2023 Data-dependent and Scale-Invariant Kernel for Support Vector Machine Classification
Vinayaka Vivekananda Malgi, Sunil Aryal, Zafaryab Rasool, David Tay
PAKDD (1)2
2022 sGrid++: Revising Simple Grid Based Density Estimator for Mining Outlying Aspect
Durgesh Samariya, Jiangang Ma, Sunil Aryal
WISE3
2021 Ensemble of Local Decision Trees for Anomaly Detection in Mixed Data
Sunil Aryal, Jonathan R. Wells
ECML/PKDD (1)1
2020 A New Effective and Efficient Measure for Outlying Aspect Mining
Durgesh Samariya, Sunil Aryal, Kai Ming Ting, Jiangang Ma
WISE (2)2
2020 A comparative study of data-dependent approaches without learning in measuring similarities of data objects
Sunil Aryal, Kai Ming Ting, Takashi Washio, Gholamreza Haffari
Data Min. Knowl. Discov.1
2020 Simple supervised dissimilarity measure: Bolstering iForest-induced similarity with class information without learning
Jonathan R. Wells, Sunil Aryal, Kai Ming Ting
Knowl. Inf. Syst.2
2018 Which Outlier Detector Should I use?
abstract
This tutorial has four aims: (1) Providing the current comparative works on different outlier detectors, and analysing the strengths and weaknesses of these works and their recommendations. (2) Presenting non-obvious applications of outlier detectors. This provides examples of how outlier detectors are used in areas which are not normally considered to be the domains of outlier detection. (3) Inviting the research community to explore future research directions, in terms of both comparative study and outlier detection in general. (4) Giving an advice on the factors to consider when choosing an outlier detector, and strengths and weaknesses of some "top" recommended algorithms based on the current understanding in the literature.
Kai Ming Ting, Sunil Aryal, Takashi Washio
ICDM2
2018 Anomaly Detection Technique Robust to Units and Scales of Measurement
Sunil Aryal
PAKDD (1)1
2017 Data-dependent dissimilarity measure: an effective alternative to geometric distance measures
Sunil Aryal, Kai Ming Ting, Takashi Washio, Gholamreza Haffari
Knowl. Inf. Syst.1
2014 Mp-Dissimilarity: A Data Dependent Dissimilarity Measure
abstract
Nearest neighbour search is a core process in many data mining algorithms. Finding reliable closest matches of a query in a high dimensional space is still a challenging task. This is because the effectiveness of many dissimilarity measures, that are based on a geometric model, such as lp-norm, decreases as the number of dimensions increases. In this paper, we examine how the data distribution can be exploited to measure dissimilarity between two instances and propose a new data dependent dissimilarity measure called 'mp-dissimilarity'. Rather than relying on geometric distance, it measures the dissimilarity between two instances in each dimension as a probability mass in a region that encloses the two instances. It deems the two instances in a sparse region to be more similar than two instances in a dense region, though these two pairs of instances have the same geometric distance. Our empirical results show that the proposed dissimilarity measure indeed provides a reliable nearest neighbour search in high dimensional spaces, particularly in sparse data. Mp-dissimilarity produced better task specific performance than lp-norm and cosine distance in classification and information retrieval tasks.
Sunil Aryal, Kai Ming Ting, Gholamreza Haffari, Takashi Washio
ICDM1
2014 Improving iForest with Relative Mass
Sunil Aryal, Kai Ming Ting, Jonathan R. Wells, Takashi Washio
PAKDD (2)1
2013 MassBayes: A New Generative Classifier with Multi-dimensional Likelihood Estimation
Sunil Aryal, Kai Ming Ting
PAKDD (1)1
2013 DEMass: a new density estimator for big data
Kai Ming Ting, Takashi Washio, Jonathan R. Wells, Fei Tony Liu, Sunil Aryal
Knowl. Inf. Syst.5