EDBT 2026 Demo / reviewers in the wild / expert
Faisal Hamman
dblp:332/3468
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Trustworthy machine learning · 53% Efficient and distributed learning · 33% Language models and text generation · 14% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
1.2 | 3 | 2025 | Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees · ICML 2023 Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations · NeurIPS 2025 Quantifying Prediction Consistency Under Fine-tuning Multiplicity in Tabular LLMs · ICML 2025 |
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation |
0.9 | 2 | 2025 | Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees · ICML 2023 Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
data selection |
0.9 | 1 | 2025 | T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction Tuning · NeurIPS 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction Tuning · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.9 | 1 | 2025 | Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › robustness
prediction consistency |
0.9 | 1 | 2025 | Quantifying Prediction Consistency Under Fine-tuning Multiplicity in Tabular LLMs · ICML 2025 |
Machine learning › Trustworthy machine learning
uncertainty and reliability |
0.9 | 1 | 2025 | Quantifying Prediction Consistency Under Fine-tuning Multiplicity in Tabular LLMs · ICML 2025 |
Machine learning › Efficient and distributed learning › federated learning › trustworthy federated learning
fair federated learning |
0.8 | 1 | 2024 | Demystifying Local & Global Fairness Trade-offs in Federated Learning Using Partial Information Decomposition · ICLR 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Demystifying Local & Global Fairness Trade-offs in Federated Learning Using Partial Information Decomposition · ICLR 2024 |
Machine learning › Efficient and distributed learning
federated learning |
0.8 | 1 | 2024 | Demystifying Local & Global Fairness Trade-offs in Federated Learning Using Partial Information Decomposition · ICLR 2024 |
Machine learning › Trustworthy machine learning › fairness
group fairness |
0.8 | 1 | 2024 | Demystifying Local & Global Fairness Trade-offs in Federated Learning Using Partial Information Decomposition · ICLR 2024 |
Machine learning › Trustworthy machine learning › interpretability › counterfactual explanation
robust counterfactual explanation |
0.7 | 1 | 2023 | Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees · ICML 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | Robust Counterfactual Explanations for Neural Networks With Probabilistic Guarantees · ICML 2023 |
Methods — techniques the papers use, named apart from their topics
local stability sampling · 0.9instruction-following difficulty scoring · 0.9embedding space analysis · 0.9decision boundary analysis · 0.9counterfactual explanation · 0.9partial information decomposition · 0.8convex optimization · 0.8lipschitz concentration bounds · 0.7gaussian sampling · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Quantifying Knowledge Distillation using Partial Information DecompositionabstractKnowledge distillation deploys complex machine learning models in resource-constrained environments by training a smaller student model to emulate internal representations of a complex teacher model. However, the teacher’s representations can also encode nuisance or additional information not relevant to the downstream task. Distilling such irrelevant information can actually impede the performance of a capacity-limited student model. This observation motivates our primary question: What are the information-theoretic limits of knowledge distillation? To this end, we leverage Partial Information Decomposition to quantify and explain the transferred knowledge and knowledge left to distill for a downstream task. We theoretically demonstrate that the task-relevant transferred knowledge is succinctly captured by the measure of redundant information about the task between the teacher and student. We propose a novel multi-level optimization to incorporate redundant information as a regularizer, leading to our framework of Redundant Information Distillation (RID). RID leads to more resilient and effective distillation under nuisance teachers as it succinctly quantifies task-relevant knowledge rather than simply aligning student and teacher representations. Pasan Dissanayake, Faisal Hamman, Barproda Halder, Ilia Sucholutsky, Qiuyi Zhang 0001, Sanghamitra Dutta |
AISTATS | 2 |
| 2025 | Quantifying Prediction Consistency Under Fine-tuning Multiplicity in Tabular LLMsabstractFine-tuning LLMs on tabular classification tasks can lead to the phenomenon of *fine-tuning multiplicity* where equally well-performing models make conflicting predictions on the same input. Fine-tuning multiplicity can arise due to variations in the training process, e.g., seed, weight initialization, minor changes to training data, etc., raising concerns about the reliability of Tabular LLMs in high-stakes applications such as finance, hiring, education, healthcare. Our work formalizes this unique challenge of fine-tuning multiplicity in Tabular LLMs and proposes a novel measure to quantify the consistency of individual predictions without expensive model retraining. Our measure quantifies a prediction's consistency by analyzing (sampling) the model's local behavior around that input in the embedding space. Interestingly, we show that sampling in the local neighborhood can be leveraged to provide probabilistic guarantees on prediction consistency under a broad class of fine-tuned models, i.e., inputs with sufficiently high local stability (as defined by our measure) also remain consistent across several fine-tuned models with high probability. We perform experiments on multiple real-world datasets to show that our local stability measure preemptively captures consistency under actual multiplicity across several fine-tuned models, outperforming competing measures. Faisal Hamman, Pasan Dissanayake, Saumitra Mishra, Freddy Lécué, Sanghamitra Dutta |
ICML | 1 |
| 2025 | Counterfactual Explanations for Model Ensembles Using Entropic Risk Measures
Erfaun Noorani, Pasan Dissanayake, Faisal Hamman, Sanghamitra Dutta |
AAMAS | 3 |
| 2025 | T-SHIRT: Token-Selective Hierarchical Data Selection for Instruction TuningabstractInstruction tuning is essential for Large Language Models (LLMs) to effectively follow user instructions. To improve training efficiency and reduce data redundancy, recent works use LLM-based scoring functions, e.g., Instruction-Following Difficulty (IFD), to select high–quality instruction-tuning data with scores above a threshold. While these data selection methods often lead to models that can match or even exceed the performance of models trained on the full datasets, we identify two key limitations: (i) they assess quality at the sample level, ignoring token-level informativeness; and (ii) they overlook the robustness of the scoring method, often selecting a sample due to superficial lexical features instead of its true quality. In this work, we propose Token-Selective HIeRarchical Data Selection for Instruction Tuning (T-SHIRT), a novel data selection framework that introduces a new scoring method to include only informative tokens in quality evaluation and also promote robust and reliable samples whose neighbors also show high quality with less local inconsistencies. We demonstrate that models instruction-tuned on a curated dataset (only 5% of the original size) using T-SHIRT can outperform those trained on the entire large-scale dataset by up to 5.48 points on average across eight benchmarks. Across various LLMs and training set scales, our method consistently surpasses existing state-of-the-art data selection techniques, while also remaining both cost-effective and highly efficient. For instance, by using GPT-2 for score computation, we are able to process a dataset of 52k samples in 40 minutes on a single GPU. Our code is available at https://github.com/Dynamite321/T-SHIRT. Yanjun Fu, Faisal Hamman, Sanghamitra Dutta |
NeurIPS | 2 |
| 2025 | Few-Shot Knowledge Distillation of LLMs With Counterfactual ExplanationsabstractKnowledge distillation is a promising approach to transfer capabilities from complex teacher models to smaller, resource-efficient student models that can be deployed easily, particularly in task-aware scenarios. However, existing methods of task-aware distillation typically require substantial quantities of data which may be unavailable or expensive to obtain in many practical scenarios. In this paper, we address this challenge by introducing a novel strategy called **Co**unterfactual-explanation-infused **D**istillation CoD for *few-shot task-aware knowledge distillation by systematically infusing counterfactual explanations*. Counterfactual explanations (CFEs) refer to inputs that can flip the output prediction of the teacher model with minimum perturbation. Our strategy CoD leverages these CFEs to precisely map the teacher's decision boundary with significantly fewer samples. We provide theoretical guarantees for motivating the role of CFEs in distillation, from both statistical and geometric perspectives. We mathematically show that CFEs can improve parameter estimation by providing more informative examples near the teacher’s decision boundary. We also derive geometric insights on how CFEs effectively act as knowledge probes, helping the students mimic the teacher's decision boundaries more effectively than standard data. We perform experiments across various datasets and LLMs to show that CoD outperforms standard distillation approaches in few-shot regimes (as low as 8 - 512 samples). Notably, CoD only uses half of the original samples used by the baselines, paired with their corresponding CFEs and still improves performance. Faisal Hamman, Pasan Dissanayake, Yanjun Fu, Sanghamitra Dutta |
NeurIPS | 1 |
| 2024 | Demystifying Local & Global Fairness Trade-offs in Federated Learning Using Partial Information DecompositionabstractThis work presents an information-theoretic perspective to group fairness trade-offs in federated learning (FL) with respect to sensitive attributes, such as gender, race, etc. Existing works often focus on either $\textit{global fairness}$ (overall disparity of the model across all clients) or $\textit{local fairness}$ (disparity of the model at each client), without always considering their trade-offs. There is a lack of understanding regarding the interplay between global and local fairness in FL, particularly under data heterogeneity, and if and when one implies the other. To address this gap, we leverage a body of work in information theory called partial information decomposition (PID), which first identifies three sources of unfairness in FL, namely, $\textit{Unique Disparity}$, $\textit{Redundant Disparity}$, and $\textit{Masked Disparity}$. We demonstrate how these three disparities contribute to global and local fairness using canonical examples. This decomposition helps us derive fundamental limits on the trade-off between global and local fairness, highlighting where they agree or disagree. We introduce the $\textit{Accuracy and Global-Local Fairness Optimality Problem}$ (AGLFOP), a convex optimization that defines the theoretical limits of accuracy and fairness trade-offs, identifying the best possible performance any FL strategy can attain given a dataset and client distribution. We also present experimental results on synthetic datasets and the ADULT dataset to support our theoretical findings. Faisal Hamman, Sanghamitra Dutta |
ICLR | 1 |
| 2024 | A Unified View of Group Fairness Tradeoffs Using Partial Information DecompositionabstractThis paper introduces a novel information-theoretic perspective on the relationship between prominent group fairness notions in machine learning, namely statistical parity, equalized odds, and predictive parity. It is well known that simultaneous satisfiability of these three fairness notions is usually impossible, motivating practitioners to resort to approximate fairness solutions rather than stringent satisfiability of these definitions. How-ever, a comprehensive analysis of their interrelations, particularly when they are not exactly satisfied, remains largely unexplored. Our main contribution lies in elucidating an exact relationship between these three measures of (un)fairness by leveraging a body of work in information theory called partial information decomposition (PID). In this work, we leverage PID to identify the granular regions where these three measures of (un)fairness overlap and where they disagree with each other leading to potential tradeoffs. We also include numerical simulations to complement our results. Faisal Hamman, Sanghamitra Dutta |
ISIT | 1 |
| 2023 | Robust Counterfactual Explanations for Neural Networks With Probabilistic GuaranteesabstractThere is an emerging interest in generating robust counterfactual explanations that would remain valid if the model is updated or changed even slightly. Towards finding robust counterfactuals, existing literature often assumes that the original model $m$ and the new model $M$ are bounded in the parameter space, i.e., $\|\text{Params}(M){-}\text{Params}(m)\|{<}\Delta$. However, models can often change significantly in the parameter space with little to no change in their predictions or accuracy on the given dataset. In this work, we introduce a mathematical abstraction termed naturally-occurring model change, which allows for arbitrary changes in the parameter space such that the change in predictions on points that lie on the data manifold is limited. Next, we propose a measure – that we call Stability – to quantify the robustness of counterfactuals to potential model changes for differentiable models, e.g., neural networks. Our main contribution is to show that counterfactuals with sufficiently high value of Stability as defined by our measure will remain valid after potential “naturally-occurring” model changes with high probability (leveraging concentration bounds for Lipschitz function of independent Gaussians). Since our quantification depends on the local Lipschitz constant around a data point which is not always available, we also examine practical relaxations of our proposed measure and demonstrate experimentally how they can be incorporated to find robust counterfactuals for neural networks that are close, realistic, and remain valid after potential model changes. Faisal Hamman, Erfaun Noorani, Saumitra Mishra, Daniele Magazzeni, Sanghamitra Dutta |
ICML | 1 |