VLDB 2026 Research / reviewers in the wild / expert
Pasan Dissanayake
dblp:292/8397
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0003-0997-332XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Quantifying Knowledge Distillation using Partial Information DecompositionabstractKnowledge distillation deploys complex machine learning models in resource-constrained environments by training a smaller student model to emulate internal representations of a complex teacher model. However, the teacher’s representations can also encode nuisance or additional information not relevant to the downstream task. Distilling such irrelevant information can actually impede the performance of a capacity-limited student model. This observation motivates our primary question: What are the information-theoretic limits of knowledge distillation? To this end, we leverage Partial Information Decomposition to quantify and explain the transferred knowledge and knowledge left to distill for a downstream task. We theoretically demonstrate that the task-relevant transferred knowledge is succinctly captured by the measure of redundant information about the task between the teacher and student. We propose a novel multi-level optimization to incorporate redundant information as a regularizer, leading to our framework of Redundant Information Distillation (RID). RID leads to more resilient and effective distillation under nuisance teachers as it succinctly quantifies task-relevant knowledge rather than simply aligning student and teacher representations. Pasan Dissanayake, Faisal Hamman, Barproda Halder, Ilia Sucholutsky, Qiuyi Zhang 0001, Sanghamitra Dutta |
AISTATS | 1 |
| 2025 | Quantifying Prediction Consistency Under Fine-tuning Multiplicity in Tabular LLMsabstractFine-tuning LLMs on tabular classification tasks can lead to the phenomenon of *fine-tuning multiplicity* where equally well-performing models make conflicting predictions on the same input. Fine-tuning multiplicity can arise due to variations in the training process, e.g., seed, weight initialization, minor changes to training data, etc., raising concerns about the reliability of Tabular LLMs in high-stakes applications such as finance, hiring, education, healthcare. Our work formalizes this unique challenge of fine-tuning multiplicity in Tabular LLMs and proposes a novel measure to quantify the consistency of individual predictions without expensive model retraining. Our measure quantifies a prediction's consistency by analyzing (sampling) the model's local behavior around that input in the embedding space. Interestingly, we show that sampling in the local neighborhood can be leveraged to provide probabilistic guarantees on prediction consistency under a broad class of fine-tuned models, i.e., inputs with sufficiently high local stability (as defined by our measure) also remain consistent across several fine-tuned models with high probability. We perform experiments on multiple real-world datasets to show that our local stability measure preemptively captures consistency under actual multiplicity across several fine-tuned models, outperforming competing measures. Faisal Hamman, Pasan Dissanayake, Saumitra Mishra, Freddy Lécué, Sanghamitra Dutta |
ICML | 2 |
| 2025 | Counterfactual Explanations for Model Ensembles Using Entropic Risk Measures
Erfaun Noorani, Pasan Dissanayake, Faisal Hamman, Sanghamitra Dutta |
AAMAS | 2 |
| 2025 | Private Counterfactual Retrieval with Immutable FeaturesabstractIn a classification task, counterfactual explanations provide the minimum change needed for an input to be classified into a favorable class. We consider the problem of privately retrieving the exact closest counterfactual from a database of accepted samples while enforcing that certain features of the input sample cannot be changed, i.e., they are immutable. An applicant (user) whose feature vector is rejected by a machine learning model wants to retrieve the sample closest to them in the database without altering a private subset of their features, which constitutes the immutable set. While doing this, the user should keep their feature vector, immutable set and the resulting counterfactual index information-theoretically private from the institution. We refer to this as immutable private counterfactual retrieval (I-PCR) problem which generalizes PCR to a more practical setting. In this paper, we propose two I-PCR schemes by leveraging techniques from private information retrieval (PIR) and characterize their communication costs. Further, we quantify the information that the user learns about the database and compare it for the proposed schemes. Shreya Meel, Pasan Dissanayake, Mohamed W. Nomeir, Sanghamitra Dutta, Sennur Ulukus |
ISIT | 2 |
| 2025 | Few-Shot Knowledge Distillation of LLMs With Counterfactual ExplanationsabstractKnowledge distillation is a promising approach to transfer capabilities from complex teacher models to smaller, resource-efficient student models that can be deployed easily, particularly in task-aware scenarios. However, existing methods of task-aware distillation typically require substantial quantities of data which may be unavailable or expensive to obtain in many practical scenarios. In this paper, we address this challenge by introducing a novel strategy called **Co**unterfactual-explanation-infused **D**istillation CoD for *few-shot task-aware knowledge distillation by systematically infusing counterfactual explanations*. Counterfactual explanations (CFEs) refer to inputs that can flip the output prediction of the teacher model with minimum perturbation. Our strategy CoD leverages these CFEs to precisely map the teacher's decision boundary with significantly fewer samples. We provide theoretical guarantees for motivating the role of CFEs in distillation, from both statistical and geometric perspectives. We mathematically show that CFEs can improve parameter estimation by providing more informative examples near the teacher’s decision boundary. We also derive geometric insights on how CFEs effectively act as knowledge probes, helping the students mimic the teacher's decision boundaries more effectively than standard data. We perform experiments across various datasets and LLMs to show that CoD outperforms standard distillation approaches in few-shot regimes (as low as 8 - 512 samples). Notably, CoD only uses half of the original samples used by the baselines, paired with their corresponding CFEs and still improves performance. Faisal Hamman, Pasan Dissanayake, Yanjun Fu, Sanghamitra Dutta |
NeurIPS | 2 |
| 2024 | Model Reconstruction Using Counterfactual Explanations: A Perspective From Polytope TheoryabstractCounterfactual explanations provide ways of achieving a favorable model outcome with minimum input perturbation. However, counterfactual explanations can also be leveraged to reconstruct the model by strategically training a surrogate model to give similar predictions as the original (target) model. In this work, we analyze how model reconstruction using counterfactuals can be improved by
further leveraging the fact that the counterfactuals also lie quite close to the decision boundary. Our main contribution is to derive novel theoretical relationships between the error in model reconstruction and the number of counterfactual queries required using polytope theory. Our theoretical analysis leads us to propose a strategy for model reconstruction that we call Counterfactual Clamping Attack (CCA) which trains a surrogate model using a unique loss function that treats counterfactuals differently than ordinary instances. Our approach also alleviates the related problem of decision boundary shift that arises in existing model reconstruction approaches when counterfactuals are treated as ordinary instances. Experimental results demonstrate that our strategy improves fidelity between the target and surrogate model predictions on several datasets. Pasan Dissanayake, Sanghamitra Dutta |
NeurIPS | 1 |
| 2022 | The Eigenvectors of Single-Spiked Complex Wishart Matrices: Finite and Asymptotic AnalysesabstractLet$\mathrm {W}\in \mathbb {C}^{n\times n}$be a single-spiked Wishart matrix in the class$\mathrm {W}\sim \mathcal {CW}_{n}(m,\mathrm {I}_{n}+ \theta \mathrm {v}\mathrm {v}^{\dagger}) $with$m\geq n$, where${\mathrm {I}}_{n}$is the$n\times n$identity matrix,$\mathrm {v}\in \mathbb {C}^{n\times 1}$is an arbitrary vector with unit Euclidean norm,$\theta \geq 0$is a non-random parameter, and$(\cdot)^{\dagger} $represents the conjugate-transpose operator. Let u1 and${\mathrm {u}}_{n}$denote the eigenvectors corresponding to the smallest and the largest eigenvalues of W, respectively. This paper investigates the probability density function (p.d.f.) of the random quantity$Z_{\ell }^{(n)}=\left |{\mathrm {v}^{\dagger} \mathrm {u}_\ell }\right |^{2}\in (0,1)$for$\ell =1,n$. In particular, we derive a finite dimensional closed-form p.d.f. for$Z_{1}^{(n)}$which is amenable to asymptotic analysis as$m,n$diverges with$m-n$fixed. It turns out that, in this asymptotic regime, the scaled random variable$nZ_{1}^{(n)}$converges in distribution to$\chi ^{2}_{2}/2(1+\theta)$, where$\chi _{2}^{2}$denotes a chi-squared random variable with two degrees of freedom. This reveals that u1 can be used to infer information about the spike. On the other hand, the finite dimensional p.d.f. of$Z_{n}^{(n)}$is expressed as a double integral in which the integrand contains a determinant of a square matrix of dimension$(n-2)$. Although a simple solution to this double integral seems intractable, for special configurations of$n=2,3$, and 4, we obtain closed-form expressions. Prathapasinghe Dharmawansa, Pasan Dissanayake, Yang Chen 0002 |
IEEE Trans. Inf. Theory | 2 |
| 2022 | Distribution of the Scaled Condition Number of Single-Spiked Complex Wishart MatricesabstractLet$\mathbf {X}\in \mathbb {C}^{n\times m}$($m\geq n$) be a random matrix with independent columns each distributed as complex multivariate Gaussian with zero mean andsingle-spikedcovariance matrix$\mathbf {I}_{n}+ \eta \mathbf {u}\mathbf {u}^{*}$, where$\mathbf {I}_{n}$is the$n\times n$identity matrix,$\mathbf {u}\in \mathbb {C}^{n\times 1}$is an arbitrary vector with unit Euclidean norm,$\eta \geq 0$is a non-random parameter, and$(\cdot)^{*}$represents the conjugate-transpose. This paper investigates the distribution of the random quantity$\kappa _{\text {SC}}^{2}(\mathbf {X})=\sum _{k=1}^{n} \lambda _{k}/\lambda _{1}$, where$0\le \lambda _{1}\le \lambda _{2}\le \ldots \leq \lambda _{n} < \infty $are the ordered eigenvalues of$\mathbf {X}\mathbf {X}^{*}$(i.e., single-spiked Wishart matrix). This random quantity is intimately related to the so calledscaled condition numberor the Demmel condition number (i.e.,$\kappa _{\text {SC}}(\mathbf {X})$) and the minimum eigenvalue of the fixed trace Wishart-Laguerre ensemble (i.e.,$\kappa _{\text {SC}}^{-2}(\mathbf {X})$). In particular, we use an orthogonal polynomial approach to derive an exact expression for the probability density function of$\kappa _{\text {SC}}^{2}(\mathbf {X})$which is amenable to asymptotic analysis as matrix dimensions grow large. Our asymptotic results reveal that, as$m,n\to \infty $such that$m-n$is fixed and when$\eta $scales on the order of$1/n$,$\kappa _{\text {SC}}^{2}(\mathbf {X})$scales on the order of$n^{3}$. In this respect we establish simple closed-form expressions for the limiting distributions. It turns out that, as$m,n\to \infty $such that$n/m\to c\in (0,1)$, properly centered$\kappa _{\text {SC}}^{2}(\mathbf {X})$fluctuates on the scale$m^{\frac {1}{3}}$. Pasan Dissanayake, Prathapasinghe Dharmawansa, Yang Chen 0002 |
IEEE Trans. Inf. Theory | 1 |