VLDB 2026 Research / reviewers in the wild / expert
Pramuditha Perera
dblp:192/4889
· DBLP profile ↗
22ranked-venue papers
12as first author
9since 2021 · last 2025
0000-0003-2821-6367ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 10 first-author · 5 since 2021Security and privacy · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Approximately Aligned DecodingabstractIt is common to reject undesired outputs of Large Language Models (LLMs); however, current methods to do so require an excessive amount of computation to re-sample after a rejection, or distort the distribution of outputs by constraining the output to highly improbable tokens.
We present a method, Approximately Aligned Decoding (AprAD), to balance the distortion of the output distribution with computational efficiency, inspired by algorithms from the speculative decoding literature.
AprAD allows for the generation of long sequences of text with difficult-to-satisfy constraints, while amplifying low probability outputs much less compared to existing methods.
We show through a series of experiments that the task-specific performance of AprAD is comparable to methods that do not distort the output distribution, while being much more computationally efficient. Daniel Melcer, Sujan K. Gonugondla, Pramuditha Perera, Haifeng Qian, Wen-Hao Chiang, Nihal Jain, Pranav Garg 0001, Xiaofei Ma 0001, Anoop Deoras |
NeurIPS | 3 |
| 2024 | Multi-Modal Hallucination Control by Visual Information GroundingabstractGenerative Vision-Language Models (VLMs) are prone to generate plausible-sounding textual answers that, however, are not always grounded in the input image. We investigate this phenomenon, usually referred to as “hallucination” and show that it stems from an excessive reliance on the language prior. In particular, we show that as more tokens are generated, the reliance on the visual prompt decreases, and this behavior strongly correlates with the emergence of hallucinations. To reduce hallucinations, we introduce Multi-Modal Mutual-Information Decoding (M3ID), a new sampling method for prompt amplification. M3ID amplifies the influence of the reference image over the language prior, hence favoring the generation of tokens with higher mutual information with the visual prompt. M3ID can be applied to any pre-trained autoregressive VLM at inference time without necessitating further training and with minimal computational overhead. If training is an option, we show that M3ID can be paired with Direct Preference Optimization (DPO) to improve the model's reliance on the prompt image without requiring any labels. Our empirical findings show that our algorithms maintain the fluency and linguistic capabilities of pre-trained VLMs while reducing hallucinations by mitigating visually ungrounded answers. Specifically, for the LLaVA 13B model, M3ID and M3ID+DPO reduce the percentage of hallucinated objects in captioning tasks by 25% and 28%, respectively, and improve the accuracy on VQA benchmarks such as POPE by 21% and 24%. Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, Stefano Soatto |
CVPR | 5 |
| 2024 | Meaning Representations from Trajectories in Autoregressive ModelsabstractWe propose to extract meaning representations from autoregressive language models by considering the distribution of all possible trajectories extending an input text. This strategy is prompt-free, does not require fine-tuning, and is applicable to any pre-trained autoregressive model. Moreover, unlike vector-based representations, distribution-based representations can also model asymmetric relations (e.g., direction of logical entailment, hypernym/hyponym relations) by using algebraic operations between likelihood functions. These ideas are grounded in distributional perspectives on semantics and are connected to standard constructions in automata theory, but to our knowledge they have not been applied to modern language models. We empirically show that the representations obtained from large models align well with human annotations, outperform other zero-shot and prompt-free methods on semantic similarity tasks, and can be used to solve more complex entailment and containment tasks that standard embeddings cannot handle. Finally, we extend our method to represent data from different modalities (e.g., image and text) using multimodal autoregressive models. Our code is available at: https://github.com/tianyu139/meaning-as-trajectories Tian Yu Liu, Matthew Trager, Alessandro Achille, Pramuditha Perera, Luca Zancato, Stefano Soatto |
ICLR | 4 |
| 2024 | Federated Generalized Face Presentation Attack DetectionabstractFace presentation attack detection (fPAD) plays a critical role in the modern face recognition pipeline. An fPAD model with good generalization can be obtained when it is trained with face images from different input distributions and different types of spoof attacks. In reality, training data (both real face images and spoof images) are not directly shared between data owners due to legal and privacy issues. In this article, with the motivation of circumventing this challenge, we propose a federated face presentation attack detection (FedPAD) framework that simultaneously takes advantage of rich fPAD information available at different data owners while preserving data privacy. In the proposed framework, each data owner (referred to as data centers) locally trains its own fPAD model. A server learns a global fPAD model by iteratively aggregating model updates from all data centers without accessing private data in each of them. Once the learned global model converges, it is used for fPAD inference. To equip the aggregated fPAD model in the server with better generalization ability to unseen attacks from users, following the basic idea of FedPAD, we further propose a federated generalized face presentation attack detection (FedGPAD) framework. A federated domain disentanglement strategy is introduced in FedGPAD, which treats each data center as one domain and decomposes the fPAD model into domain-invariant and domain-specific parts in each data center. Two parts disentangle the domain-invariant and domain-specific features from images in each local data center. A server learns a global fPAD model by only aggregating domain-invariant parts of the fPAD models from data centers, and thus, a more generalized fPAD model can be aggregated in server. We introduce the experimental setting to evaluate the proposed FedPAD and FedGPAD frameworks and carry out extensive experiments to provide various insights about federated learning for fPAD. Rui Shao 0001, Pramuditha Perera, Pong C. Yuen, Vishal M. Patel |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | À-la-carte Prompt Tuning (APT): Combining Distinct Data Via Composable PromptingabstractWe introduce À-la-carte Prompt Tuning (APT), a transformer-based scheme to tune prompts on distinct data so that they can be arbitrarily composed at inference time. The individual prompts can be trained in isolation, possibly on different devices, at different times, and on different distributions or domains. Furthermore each prompt only contains information about the subset of data it was exposed to during training. During inference, models can be assembled based on arbitrary selections of data sources, which we call à-la-carte learning. À-la-carte learning enables constructing bespoke models specific to each user's individual access rights and preferences. We can add or remove information from the model by simply adding or removing the corresponding prompts without retraining from scratch. We demonstrate that à-la-carte built models achieve accuracy within 5% of models trained on the union of the respective sources, with comparable cost in terms of training and inference time. For the continual learning benchmarks Split CIFAR- 100 and CORe50, we achieve state-of-the-art performance. Benjamin Bowman, Alessandro Achille, Luca Zancato, Matthew Trager, Pramuditha Perera, Giovanni Paolini, Stefano Soatto |
CVPR | 5 |
| 2023 | Train/Test-Time Adaptation with RetrievalabstractWe introduce Train/Test-Time Adaptation with Retrieval (T3AR), a method to adapt models both at train and test time by means of a retrieval module and a searchable pool of external samples. Before inference, T3AR adapts a given model to the downstream task using refined pseudo-labels and a self-supervised contrastive objective function whose noise distribution leverages retrieved real samples to improve feature adaptation on the target data manifold. The retrieval of real images is key to T3AR since it does not rely solely on synthetic data augmentations to compensate for the lack of adaptation data, as typically done by other adaptation algorithms. Furthermore, thanks to the retrieval module, our method gives the user or service provider the possibility to improve model adaptation on the downstream task by incorporating further relevant data or to fully remove samples that may no longer be available due to changes in user preference after deployment. First, we show that T3AR can be used at training time to improve downstream fine-grained classification over standard fine-tuning baselines, and the fewer the adaptation data the higher the relative improvement (up to 13%). Second, we apply T3ARfor test-time adaptation and show that exploiting a pool of external images at test-time leads to more robust representations over existing methods on DomainNet-126 and VISDA-C, especially when few adaptation data are available (up to 8%). Luca Zancato, Alessandro Achille, Tian Yu Liu, Matthew Trager, Pramuditha Perera, Stefano Soatto |
CVPR | 5 |
| 2023 | Linear Spaces of Meanings: Compositional Structures in Vision-Language ModelsabstractWe investigate compositional structures in data embeddings from pre-trained vision-language models (VLMs). Traditionally, compositionality has been associated with algebraic operations on embeddings of words from a preexisting vocabulary. In contrast, we seek to approximate representations from an encoder as combinations of a smaller set of vectors in the embedding space. These vectors can be seen as "ideal words" for generating concepts directly within embedding space of the model. We first present a framework for understanding compositional structures from a geometric perspective. We then explain what these compositional structures entail probabilistically in the case of VLM embeddings, providing intuitions for why they arise in practice. Finally, we empirically explore these structures in CLIP’s embeddings and we evaluate their usefulness for solving different vision-language tasks such as classification, debiasing, and retrieval. Our results show that simple linear algebraic operations on embedding vectors can be used as compositional and interpretable methods for regulating the behavior of VLMs. Matthew Trager, Pramuditha Perera, Luca Zancato, Alessandro Achille, Parminder Bhatia, Stefano Soatto |
ICCV | 2 |
| 2022 | Open-Set Adversarial Defense with Clean-Adversarial Mutual Learning
Rui Shao 0001, Pramuditha Perera, Pong C. Yuen, Vishal M. Patel |
Int. J. Comput. Vis. | 2 |
| 2021 | Geometric Transformation-Based Network Ensemble for Open-Set RecognitionabstractOpen-set recognition focuses on the problem of determining whether a given query image belongs to one of the classes known to the network. Recent works in open-set recognition have attempted to solve this problem using an external deep network by either modeling known class samples or simulating open-set samples at the expense of more parameters. In this work, we propose a modified network structure and an inference rule that leads to better open-set recognition performance. First, we show that networks learn different representations when they are trained on datasets subjected to extreme geometric transformations. By exploiting this fact, the proposed mechanism learns multiple representations from the same set of image data using a set of parallel networks with identical structures. During inference, decisions of all independent networks are fused using majority voting to arrive at predictions. The proposed method obtains state-of-the-art open-set detection performance on multiple object recognition datasets. Pramuditha Perera, Vishal M. Patel |
ICME | 1 |
| 2020 | Generative-Discriminative Feature Representations for Open-Set RecognitionabstractWe address the problem of open-set recognition, where the goal is to determine if a given sample belongs to one of the classes used for training a model (known classes). The main challenge in open-set recognition is to disentangle open-set samples that produce high class activations from known-set samples. We propose two techniques to force class activations of open-set samples to be low. First, we train a generative model for all known classes and then augment the input with the representation obtained from the generative model to learn a classifier. This network learns to associate high classification probabilities both when image content is from the correct class as well as when the input and the reconstructed image are consistent with each other. Second, we use self-supervision to force the network to learn more informative featues when assigning class scores to improve separation of classes from each other and from open-set samples. We evaluate the performance of the proposed method with recent open-set recognition works across three datasets, where we obtain state-of-the-art results. Pramuditha Perera, Vlad I. Morariu, Rajiv Jain, Varun Manjunatha, Curtis Wigington, Vicente Ordonez, Vishal M. Patel |
CVPR | 1 |
| 2020 | Open-Set Adversarial Defense
Rui Shao 0001, Pramuditha Perera, Pong C. Yuen, Vishal M. Patel |
ECCV (17) | 2 |
| 2020 | Anomaly Detection-Based Unknown Face Presentation Attack DetectionabstractAnomaly detection-based spoof attack detection is a recent development in face Presentation Attack Detection (fPAD), where a spoof detector is learned using only non-attacked images of users. These detectors are of practical importance as they are shown to generalize well to new attack types. In this paper, we present a deep-learning solution for anomaly detection-based spoof attack detection where both classifier and feature representations are learned together end-to-end. First, we introduce a pseudo-negative class during training in the absence of attacked images. The pseudo-negative class is modeled using a Gaussian distribution whose mean is calculated by a weighted running mean. Secondly, we use pairwise confusion loss to further regularize the training process. The proposed approach benefits from the representation learning power of the CNNs and learns better features for fPAD task as shown in our ablation study. We perform extensive experiments on four publicly available datasets: Replay-Attack, Rose-Youtu, OULU-NPU and Spoof in Wild to show the effectiveness of the proposed approach over the previous methods. Code is available at: https://github.com/yashasvi97/IJCB2020_anomaly. Yashasvi Baweja, Poojan Oza, Pramuditha Perera, Vishal M. Patel |
IJCB | 3 |
| 2020 | Quickest Intruder Detection For Multiple User Active AuthenticationabstractIn this paper, we investigate how to detect intruders with low latency for Active Authentication (AA) systems with multiple-users. We extend the Quickest Change Detection (QCD) framework to the multiple-user case and formulate the Multiple-user Quickest Intruder Detection (MQID) algorithm. Furthermore, we extend the algorithm to the data-efficient scenario where intruder detection is carried out with fewer observation samples. We evaluate the effectiveness of the proposed method on two publicly available AA datasets on the face modality. Pramuditha Perera, Julian Fierrez, Vishal M. Patel |
ICIP | 1 |
| 2020 | A Joint Representation Learning and Feature Modeling Approach for One-class RecognitionabstractOne-class recognition is traditionally approached either as a representation learning problem or a feature modelling problem. In this work, we argue that both of these approaches have their own limitations; and a more effective solution can be obtained by combining the two. The proposed approach is based on the combination of a generative framework and a one-class classification method. First, we learn generative features using the one-class data with a generative framework. We augment the learned features with the corresponding reconstruction errors to obtain augmented features. Then, we qualitatively identify a suitable feature distribution that reduces the redundancy in the chosen classifier space. Finally, we force the augmented features to take the form of this distribution using an adversarial framework. We test the effectiveness of the proposed method on three one-class classification tasks and obtain state-of-the-art results. Pramuditha Perera, Vishal M. Patel |
ICPR | 1 |
| 2019 | OCGAN: One-Class Novelty Detection Using GANs With Constrained Latent RepresentationsabstractWe present a novel model called OCGAN for the classical problem of one-class novelty detection, where, given a set of examples from a particular class, the goal is to determine if a query example is from the same class. Our solution is based on learning latent representations of in-class examples using a de-noising auto-encoder network. The key contribution of our work is our proposal to explicitly constrain the latent space to exclusively represent the given class. In order to accomplish this goal, firstly, we force the latent space to have bounded support by introducing a tanh activation in the encoder's output layer. Secondly, using a discriminator in the latent space that is trained adversarially, we ensure that encoded representations of in-class examples resemble uniform random samples drawn from the same bounded space. Thirdly, using a second adversarial discriminator in the input space, we ensure all randomly drawn latent samples generate examples that look real. Finally, we introduce a gradient-descent based sampling technique that explores points in the latent space that generate potential out-of-class examples, which are fed back to the network to further train it to generate in-class examples from those points. The effectiveness of the proposed method is measured across four publicly available datasets using two one-class novelty detection protocols where we achieve state-of-the-art results. Pramuditha Perera, Ramesh Nallapati, Bing Xiang |
CVPR | 1 |
| 2019 | Deep Transfer Learning for Multiple Class Novelty DetectionabstractWe propose a transfer learning-based solution for the problem of multiple class novelty detection. In particular, we propose an end-to-end deep-learning based approach in which we investigate how the knowledge contained in an external, out-of-distributional dataset can be used to improve the performance of a deep network for visual novelty detection. Our solution differs from the standard deep classification networks on two accounts. First, we use a novel loss function, membership loss, in addition to the classical cross-entropy loss for training networks. Secondly, we use the knowledge from the external dataset more effectively to learn globally negative filters, filters that respond to generic objects outside the known class set. We show that thresholding the maximal activation of the proposed network can be used to identify novel objects effectively. Extensive experiments on four publicly available novelty detection datasets show that the proposed method achieves significant improvements over the state-of-the-art methods. Pramuditha Perera, Vishal M. Patel |
CVPR | 1 |
| 2019 | Face-Based Multiple User Active Authentication on Mobile DevicesabstractMultiple user active authentications, in contrast with the single user active authentication, require the verification of identity of multiple subjects. Both traditional verification and identification-based solutions fail to address the specific challenges presented in this problem. We introduce Extremal Openset Rejection, a two-fold mechanism with a sparse representation-based identification step and a verification step for this purpose. In the verification step, concentration of the sparsity vector and the overlap between matched and non-matched distributions are considered for decision making. We introduce a semi-parametric model based on Extreme Value Theory for modeling the distributions, and an algorithm to estimate the parameters of extreme value distributions. Effectiveness of the proposed method is demonstrated using three publicly available face-based mobile active authentication data sets. Pramuditha Perera, Vishal M. Patel |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2019 | Learning Deep Features for One-Class ClassificationabstractWe present a novel deep-learning-based approach for one-class transfer learning in which labeled data from an unrelated task is used for feature learning in one-class classification. The proposed method operates on top of a convolutional neural network (CNN) of choice and produces descriptive features while maintaining a low intra-class variance in the feature space for the given class. For this purpose two loss functions, compactness loss and descriptiveness loss, are proposed along with a parallel CNN architecture. A template matching-based framework is introduced to facilitate the testing process. Extensive experiments on publicly available anomaly detection, novelty detection, and mobile active authentication datasets show that the proposed deep one-class (DOC) classification method achieves significant improvements over the state-of-the-art. Pramuditha Perera, Vishal M. Patel |
IEEE Trans. Image Process. | 1 |
| 2018 | In2I: Unsupervised Multi-Image-to-Image Translation Using Generative Adversarial NetworksabstractIn unsupervised image-to-image translation, the goal is to learn the mapping between an input image and an output image using a set of unpaired training images. In this paper, we propose an extension of the unsupervised image-to-image translation problem to multiple input setting. Given a set of paired images from multiple modalities, a transformation is learned to translate the input into a specified domain. For this purpose, we introduce a Generative Adversarial Network (GAN) based framework along with a multi-modal generator structure and a new loss term, latent consistency loss. Through various experiments we show that leveraging multiple inputs generally improves the visual quality of the translated images. Moreover, we show that the proposed method outperforms current state-of-the-art unsupervised image-to-image translation methods. Pramuditha Perera, Mahdi Abavisani, Vishal M. Patel |
ICPR | 1 |
| 2018 | Efficient and Low Latency Detection of Intruders in Mobile Active AuthenticationabstractActive authentication (AA) refers to the problem of continuously verifying the identity of a mobile device user for the purpose of securing the device. We address the problem of quickly detecting intrusions with lower false detection rates in mobile AA systems with higher resource efficiency. Bayesian and MiniMax versions of the quickest change detection (QCD) algorithms are introduced to quickly detect intrusions in mobile AA systems. These algorithms are extended with an update rule to facilitate low-frequency sensing which leads to low utilization of resources. Effectiveness of the proposed framework is demonstrated using three publicly available unconstrained face and touch gesture-based AA datasets. It is shown that the proposed QCD-based intrusion detection methods can perform better than many state-of-the-art AA methods in terms of latency and low false detection rates. Furthermore, it is shown that employing the proposed resource-efficient extension further improves the performance of the QCD-based setup. Pramuditha Perera, Vishal M. Patel |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Extreme Value Analysis for Mobile Active User AuthenticationabstractIn this paper, we propose to improve the performance of mobile Active Authentication (AA) systems in the low false alarm region using the statistical Extreme Value Theory (EVT). The problem is studied under a Bayesian framework where extremal observations that contribute to mis-verification are given more prominence. We propose modeling the tail of the match distribution using a Generalized Pareto Distribution (GPD) in order to make better inferences about the extremal observations. A method based on the mean excess function is introduced for parameter estimation of the GPD. Effectiveness of the proposed framework is demonstrated using publicly available unconstrained mobile active authentication datasets. It is shown that the proposed EVT-based method can significantly enhance the performance of traditional AA systems in the low false alarm rate region. Pramuditha Perera, Vishal M. Patel |
FG | 1 |
| 2017 | Towards Multiple User Active Authentication in Mobile DevicesabstractTraditionally, practical authentication systems have considered only a single enrolled subject for verification. However, with the advent of mobile devices this paradigm has changed since a mobile device may be accessed by more than a single enrolled user. In this context, verification of multiple enrolled users has a practical importance. We address the issue of perfonnance degradation associated with multiple user authentication as compared to single user authentication. We interpret this problem in an open-set framework and introduce the notion of probability of negativity to alleviate the effect of multiple users in authentication. We further introduce a simple fusion scheme with the existing authentication methods to increase the intruder detection accuracy. Effectiveness of the proposed method is demonstrated using three publicly available face and touch gesture-based mobile active authentication datasets. Pramuditha Perera, Vishal M. Patel |
FG | 1 |