VLDB 2026 Research / reviewers in the wild / expert
Danial Dervovic
dblp:203/8299
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-6135-561XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 43% Graph learning · 37% Learning theory · 11% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 35% Algorithmic game theory and mechanism design · 35% Approximation and online algorithms · 30% | |
| Network and information security
1 paper |
Privacy and data protection · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning › graph pre-training
cross-domain graph pre-training |
0.9 | 1 | 2025 | Cross-Domain Graph Data Scaling: A Showcase with Diffusion Models · NeurIPS 2025 |
Machine learning › Graph learning
graph pre-training |
0.9 | 1 | 2025 | Cross-Domain Graph Data Scaling: A Showcase with Diffusion Models · NeurIPS 2025 |
Machine learning › Learning theory
excess risk bounds |
0.8 | 1 | 2024 | Bounding the Excess Risk for Linear Models Trained on Marginal-Preserving, Differentially-Private, Synthetic Data · ICML 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | A Canonical Data Transformation for Achieving Inter- and Within-Group Fairness · IEEE Trans. Inf. Forensics Secur. 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | Are Logistic Models Really Interpretable? · IJCAI 2024 |
Machine learning › Trustworthy machine learning › privacy
privacy-preserving machine learning |
0.8 | 1 | 2024 | Bounding the Excess Risk for Linear Models Trained on Marginal-Preserving, Differentially-Private, Synthetic Data · ICML 2024 |
Machine learning › Trustworthy machine learning › fairness
within-group fairness |
0.8 | 1 | 2024 | A Canonical Data Transformation for Achieving Inter- and Within-Group Fairness · IEEE Trans. Inf. Forensics Secur. 2024 |
Privacy and data protection
differential privacy |
0.8 | 1 | 2024 | Bounding the Excess Risk for Linear Models Trained on Marginal-Preserving, Differentially-Private, Synthetic Data · ICML 2024 |
Privacy and data protection › differential privacy
synthetic data generation |
0.8 | 1 | 2024 | Bounding the Excess Risk for Linear Models Trained on Marginal-Preserving, Differentially-Private, Synthetic Data · ICML 2024 |
Algorithmic game theory and mechanism design › resource allocation › online resource allocation
admission control |
0.6 | 1 | 2022 | Optimal Admission Control for Multiclass Queues with Time-Varying Arrival Rates via State Abstraction · AAAI 2022 |
Mathematical optimization › sequential decision making
markov decision processes |
0.6 | 1 | 2022 | Optimal Admission Control for Multiclass Queues with Time-Varying Arrival Rates via State Abstraction · AAAI 2022 |
Approximation and online algorithms
online algorithms |
0.5 | 1 | 2021 | Non-Parametric Stochastic Sequential Assignment With Random Arrival Times · IJCAI 2021 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2025 | Cross-Domain Graph Data Scaling: A Showcase with Diffusion Models · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
discrete diffusion model |
0.3 | 1 | 2025 | Cross-Domain Graph Data Scaling: A Showcase with Diffusion Models · NeurIPS 2025 |
Data mining › knowledge discovery process
preprocessing |
0.2 | 1 | 2024 | A Canonical Data Transformation for Achieving Inter- and Within-Group Fairness · IEEE Trans. Inf. Forensics Secur. 2024 |
Machine learning › Reinforcement learning › function approximation › representation learning for reinforcement learning
state abstraction |
0.2 | 1 | 2022 | Optimal Admission Control for Multiclass Queues with Time-Varying Arrival Rates via State Abstraction · AAAI 2022 |
Methods — techniques the papers use, named apart from their topics
regularization · 1.5marginal-preserving synthetic data · 1.5lipschitz loss · 1.5empirical risk minimization · 1.5canonical data transformation · 1.5state abstraction · 1.1dynamic programming · 1.1guided generation · 0.9discrete diffusion model · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Model Evaluation in the Dark: Robust Classifier Metrics with Missing LabelsabstractMissing data in supervised learning is well-studied, but the specific issue of missing labels during model evaluation has been overlooked. Ignoring samples with missing values, a common solution, can introduce bias, especially when data is Missing Not At Random (MNAR). We propose a multiple imputation technique for evaluating classifiers using metrics such as precision, recall, and ROC-AUC. This method not only offers point estimates but also a predictive distribution for these quantities when labels are missing. We empirically show that the predictive distribution’s location and shape are generally correct, even in the MNAR regime. Moreover, we establish that this distribution is approximately Gaussian and provide finite-sample convergence bounds. Additionally, a robustness proof is presented, confirming the validity of the approximation under a realistic error model. Danial Dervovic, Michael Cashmore |
AISTATS | 1 |
| 2025 | Cross-Domain Graph Data Scaling: A Showcase with Diffusion ModelsabstractModels for natural language and images benefit from data scaling behavior: the more data fed into the model, the better they perform. This 'better with more' phenomenon enables the effectiveness of large-scale pre-training on vast amounts of data. However, current graph pre-training methods struggle to scale up data due to heterogeneity across graphs. To achieve effective data scaling, we aim to develop a general model that is able to capture diverse data patterns of graphs and can be utilized to adaptively help the downstream tasks. To this end, we propose UniAug, a universal graph structure augmentor built on a diffusion model. We first pre-train a discrete diffusion model on thousands of graphs across domains to learn the graph structural patterns. In the downstream phase, we provide adaptive enhancement by conducting graph structure augmentation with the help of the pre-trained diffusion model via guided generation. By leveraging the pre-trained diffusion model for structure augmentation, we consistently achieve performance improvements across various downstream tasks in a plug-and-play manner. To the best of our knowledge, this study represents the first demonstration of a data-scaling graph structure augmentor on graphs across domains. Wenzhuo Tang, Haitao Mao, Danial Dervovic, Ivan Brugere, Saumitra Mishra, Yuying Xie 0001, Jiliang Tang |
NeurIPS | 3 |
| 2025 | Surrogate-Assisted Monte-Carlo Tree Search in Facility Location and Beyond (Extended Abstract)abstractCombinatorial problems abound in industry. A persistent issue encountered using search-based solutions is that evaluating particular nodes may be expensive. As an example, organisations frequently adjust their facilities network by opening new branches in promising areas and closing branches in areas where they expect low profits, which may be formulated as a combinatorial search problem. In this extended abstract, we examine a particular class of facility location problems, where the objective is to minimize the loss of sales resulting from the removal of several retail stores. However, estimating sales accurately is expensive and time-consuming. To overcome this challenge, we leverage Monte-Carlo Tree Search assisted by a surrogate model that computes evaluations faster. Initial results suggest that MCTS supported by a fast surrogate function can generate solutions faster while maintaining a solution consistent with non-assisted MCTS. Saeid Amiri, Danial Dervovic, Parisa Zehtabi, Michael Cashmore |
SOCS | 2 |
| 2025 | Balancing Fairness and Accuracy in Data-Restricted Binary ClassificationabstractFair decision-making in Machine Learning (ML) remains a critical challenge, particularly when access to sensitive information is restricted due to legal, ethical, or organizational constraints. These limitations affect both accuracy and fairness, creating tradeoffs central to the deployment of ML systems in the real world. While prior work has studied fairness-accuracy tradeoffs, most approaches focus on model outputs rather than directly examining how restricted data access impacts fairness. This leaves an important gap: understanding how fairness constraints affect model performance under real-world data restrictions . To address this gap, we propose a framework that explicitly models fairness-accuracy tradeoffs in data-restricted environments. Unlike prior work, our approach analyzes the behavior of the optimal Bayesian classifier using a discrete approximation of the data distribution, allowing us to systematically isolate the effects of fairness constraints. We evaluate our framework on three benchmark datasets—Adult, Law, and Dutch Census—revealing key insights: (1) enforcing equal accuracy on imbalanced datasets can substantially degrade performance under additional fairness constraints, (2) individual and group fairness often impose conflicting constraints, and (3) decorrelating sensitive attributes from features does not usually reduce accuracy. These findings demonstrate that our framework provides an effective, structured approach for practitioners to assess fairness constraints in decision-making pipelines. Zachary McBride Lazri, Danial Dervovic, Antigoni Polychroniadou, Ivan Brugere, Dana Dachman-Soled, Furong Huang, Min Wu 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | Bounding the Excess Risk for Linear Models Trained on Marginal-Preserving, Differentially-Private, Synthetic DataabstractThe growing use of machine learning (ML) has raised concerns that an ML model may reveal private information about an individual who has contributed to the training dataset. To prevent leakage of sensitive data, we consider using differentially- private (DP), synthetic training data instead of real training data to train an ML model. A key desirable property of synthetic data is its ability to preserve the low-order marginals of the original distribution. Our main contribution comprises novel upper and lower bounds on the excess empirical risk of linear models trained on such synthetic data, for continuous and Lipschitz loss functions. We perform extensive experimentation alongside our theoretical results. Yvonne Zhou, Mingyu Liang, Ivan Brugere, Danial Dervovic, Antigoni Polychroniadou, Min Wu 0001, Dana Dachman-Soled |
ICML | 4 |
| 2024 | Are Logistic Models Really Interpretable?
Danial Dervovic, Freddy Lécué, Nicolas Marchesotti, Daniele Magazzeni |
IJCAI | 1 |
| 2024 | A Canonical Data Transformation for Achieving Inter- and Within-Group FairnessabstractIncreases in the deployment of machine learning algorithms for applications that deal with sensitive data have brought attention to the issue of fairness in machine learning. Many works have been devoted to applications that require different demographic groups to be treated fairly. However, algorithms that aim to satisfy inter-group fairness (also called group fairness) may inadvertently treat individuals within the same demographic group unfairly. To address this issue, this article introduces a formal definition of within-group fairness that maintains fairness among individuals from within the same group. A pre-processing framework is proposed to meet both inter- and within-group fairness criteria with little compromise in performance. The framework maps the feature vectors of members from different groups to an inter-group fair canonical domain before feeding them into a scoring function. The mapping is constructed to preserve the relative relationship between the scores obtained from the unprocessed feature vectors of individuals from the same demographic group, guaranteeing within-group fairness. This framework has been applied to the Adult, COMPAS risk assessment, and Law School datasets, and its performance is demonstrated and compared with two regularization-based methods in achieving inter-group and within-group fairness. Zachary McBride Lazri, Ivan Brugere, Xin Tian 0018, Dana Dachman-Soled, Antigoni Polychroniadou, Danial Dervovic, Min Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2023 | On the Connection between Game-Theoretic Feature Attributions and Counterfactual ExplanationsabstractExplainable Artificial Intelligence (XAI) has received widespread interest in recent years, and two of the most popular types of explanations are feature attributions, and counterfactual explanations. These classes of approaches have been largely studied independently and the few attempts at reconciling them have been primarily empirical. This work establishes a clear theoretical connection between game-theoretic feature attributions, focusing on but not limited to SHAP, and counterfactuals explanations. After motivating operative changes to Shapley values based feature attributions and counterfactual explanations, we prove that, under conditions, they are in fact equivalent. We then extend the equivalency result to game-theoretic solution concepts beyond Shapley values. Moreover, through the analysis of the conditions of such equivalence, we shed light on the limitations of naively using counterfactual explanations to provide feature importances. Experiments on three datasets quantitatively show the difference in explanations at every stage of the connection between the two approaches and corroborate the theoretical findings. Emanuele Albini, Saumitra Mishra, Danial Dervovic, Daniele Magazzeni |
AIES | 4 |
| 2022 | Optimal Admission Control for Multiclass Queues with Time-Varying Arrival Rates via State AbstractionabstractWe consider a novel queuing problem where the decision-maker must choose to accept or reject randomly arriving tasks into a no buffer queue which are processed by N identical servers. Each task has a price, which is a positive real number, and a class. Each class of task has a different price distribution, service rate, and arrives according to an inhomogenous Poisson process. The objective is to decide which tasks to accept so that the total price of tasks processed is maximised over a finite horizon. We formulate the problem using a discrete time Markov Decision Process (MDP) with a hybrid state space. We show that the optimal value function has a specific structure, which enables us to solve the hybrid MDP exactly. Moreover, we rigorously prove that as the gap between successive decision epochs grows smaller, the discrete time solution approaches the optimal solution to the original continuous time problem. To improve the scalability of our approach to a greater number of servers and task classes, we present an approximation based on state abstraction. We validate our approach on synthetic data, as well as a real financial fraud data set, which is the motivating application for this work. Marc Rigter, Danial Dervovic, Parisa Hassanzadeh, Jason Long, Parisa Zehtabi, Daniele Magazzeni |
AAAI | 2 |
| 2021 | Non-Parametric Stochastic Sequential Assignment With Random Arrival TimesabstractWe consider a problem wherein jobs arrive at random times and assume random values. Upon each job arrival, the decision-maker must decide immediately whether or not to accept the job and gain the value on offer as a reward, with the constraint that they may only accept at most n jobs over some reference time period. The decision-maker only has access to M independent realisations of the job arrival process. We propose an algorithm, Non-Parametric Sequential Allocation (NPSA), for solving this problem. Moreover, we prove that the expected reward returned by the NPSA algorithm converges in probability to optimality as M grows large. We demonstrate the effectiveness of the algorithm empirically on synthetic data and on public fraud-detection datasets, from where the motivation for this work is derived. Danial Dervovic, Parisa Hassanzadeh, Samuel A. Assefa, Prashant Reddy |
IJCAI | 1 |