VLDB 2026 Research / reviewers in the wild / expert
Kashaf Khan
dblp:167/8081
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automating mixture model fitting of task durations for process conformance checkingabstractAbstract Process task duration data often exhibit multiple peaks, indicating differences in, for example, customer ages and preferences, resource capabilities or the day/hour of a week. This heterogeneous data, which captures diverse customer patterns, should be represented using different models, resulting in an overall mixture model. This paper introduces gamma mixture models to represent various customer patterns in task duration data, with a focus on automating the fitting process. The approach involves a two-stage procedure: first, divide-and-conquer using peak-, equidistance- and cluster-based techniques to partition data, and automatically fit gamma distributions to each subset. The second stage then improves the fitted mixture model by directly searching the log-likelihood surface. The method is compared with the expectation–maximization (EM) algorithm and an open tool (HyperStar), using both artificially generated datasets and a publicly available hospital billing dataset, demonstrating its effectiveness and time efficiency in modelling heterogeneous process duration data. Furthermore, a case study on process conformance checking is conducted using the hospital billing dataset, highlighting a potential application area for the method in process mining. Lingkai Yang, Sally I. McClean, Malcolm J. Faddy, Mark P. Donnelly, Kashaf Khan, Kevin Burke |
Data Min. Knowl. Discov. | 5 |
| 2025 | Modelling process durations with gamma mixtures for right-censored data: Applications in customer clustering, pattern recognition, drift detection, and rationalisation
Lingkai Yang, Sally I. McClean, Kevin Burke, Mark P. Donnelly, Kashaf Khan |
Data Knowl. Eng. | 5 |
| 2024 | Detecting Process Duration Drift Using Gamma Mixture Models in a Left-Truncated and Right-Censored EnvironmentabstractWithin the realm of business context, process duration signifies time spent by customers between successive activities. This temporal perspective offers important insight to customer behavior, highlighting potential bottlenecks, and influencing business management decisions. The distribution of these process duration often changes over time due to factors such as seasonality, emerging legislation, changes to supply chains, and customer demand. Referred to as concept drift, these variations pose challenges for robust process modeling, understanding, and refinement. Subsequently, gamma mixture models are widely employed to model durations. These source data can, however, become left-truncated and right-censored within any specific observation window thereby necessitating a (well-known) modification to the likelihood function. The approach reported in this article leveraged this adapted likelihood across a series of observation windows, applying the likelihood ratio test to identify duration changes/concept drift. Due to its flexibility in modelling any duration distribution, the gamma mixture model was used with Nelder–Mead optimized likelihood for the left-truncated and right-censored data. The number of gamma components was determined by the Bayesian information criterion. The proposed framework underwent validation through simulated exponential samples, leading to recommendations for its practical application. Subsequently, we applied the methodology to three real-life event logs exhibiting diverse characteristics. Experimental results showcase the effectiveness of our approach in terms of data fitting, as compared to Kaplan–Meier curves, and in detecting instances of drift. This comprehensive validation underscores the practical utility and reliability of our framework for dynamic business scenarios. Lingkai Yang, Sally I. McClean, Mark P. Donnelly, Kashaf Khan, Kevin Burke |
ACM Trans. Knowl. Discov. Data | 4 |
| 2022 | A multi-components approach to monitoring process structure and customer behaviour concept drift
Lingkai Yang, Sally I. McClean, Mark P. Donnelly, Kevin Burke, Kashaf Khan |
Expert Syst. Appl. | 5 |
| 2021 | H-FFMRA: A Multi Resource Fully Fair Resources Allocation Algorithm in Heterogeneous Cloud ComputingabstractThe allocation of multiple types of resources fairly and efficiently has become a substantial concern in state-of-the-art computing systems. Accordingly, the rapid growth of cloud computing has highlighted the importance of resource management as a complicated and NP-hard problem. Unlike traditional frameworks, in modern data centers, incoming jobs pose demand profiles, including diverse sets of resources such as CPU, memory, and bandwidth across multiple servers. Accordingly, the fair distribution of resources, respecting such heterogeneity appears to be a challenging issue. Furthermore, the efficient use of resources as well as fairness, establish trade-off that renders a higher degree of satisfaction for both users and providers. Dominant Resource Fairness (DRF) has been introduced as an initial attempt to address fair resource allocation in multi-resource cloud computing infrastructures. Dozens of approaches have been proposed to overcome existing shortcomings associated with DRF. Although all those developments have satisfied several desirable fairness features, there are still substantial gaps. Firstly, it is not clear how to measure the fair allocation of resources among users. Secondly, no particular trade-off considers non-dominant resources in allocation decisions. Thirdly, those allocations are not intuitively fair as some users are not able to maximize their allocations. In particular, the recent approaches have not considered the aggregate resource demands concerning dominant and non-dominant resources across multiple servers. These issues lead to an uneven allocation of resources over numerous servers which is an obstacle against utility maximization for some users with dominant resources. Correspondingly, in this paper, a resource allocation algorithm called H-FFMRA is proposed to distribute resources with fairness across servers and users, considering dominant and non-dominant resources. The experiments show that H-FFMRA achieves approximately %20 improvements on fairness as well as full utilization of resources compared to DRF in multi-server settings. Hamed Hamzeh, Sofia Meacham, Kashaf Khan, Angelos Stefanidis, Keith Phalp |
COMPSAC | 3 |
| 2020 | MRFS: A Multi-resource Fair Scheduling Algorithm in Heterogeneous Cloud ComputingabstractTask scheduling in cloud computing is considered as a significant issue that has attracted much attention over the last decade. In cloud environments, users expose considerable interest in submitting tasks on multiple Resource types. Subsequently, finding an optimal and most efficient server to host users' tasks seems a fundamental concern. Several attempts have suggested various algorithms, employing Swarm optimization and heuristics methods to solve the scheduling issues associated with cloud in a multi-resource perspective. However, these approaches have not considered the equalization of dominant resources on each specific resource type. This substantial gap leads to unfair allocation, SLA degradation and resource contention. To deal with this problem, in this paper we propose a novel task scheduling mechanism called MRFS. MRFS employs Lagrangian multipliers to locate tasks in suitable servers with respect to the number of dominant resources and maximum resource availability. To evaluate MRFS, we conduct time-series experiments in the cloudsim driven by randomly generated workloads. The results show that MRFS maximizes per-user utility function by %15-20 in FFMRA compared to FFMRA in absence of MRFS. Furthermore, the mathematical proofs confirm that the sharingincentive, and Pareto-efficiency properties are improved under MRFS. Hamed Hamzeh, Sofia Meacham, Kashaf Khan, Keith Phalp, Angelos Stefanidis |
COMPSAC | 3 |