EDBT 2026 Demo / reviewers in the wild / expert
Aniket Roy
dblp:173/0075
· DBLP profile ↗
15ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0002-0241-4165ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning More from Less: Resource-Constrained Generative AI for Classification, Generation, and PersonalizationabstractThe rapid advancement of generative models has created new opportunities for addressing core challenges in computer vision, including data scarcity, image quality, and efficient personalization. My research develops principled, resource- aware methods that enable models to generalize effectively from limited supervision, adapt efficiently to new concepts, and generate high-fidelity visual content. I first address few-shot learning through augmentation-driven uncertainty- guided mixup, improving robustness in data-constrained regimes. Building on this, I propose caption-guided multi-modal augmentation techniques that enrich visual diversity while mitigating real-to-synthetic domain gaps. To enhance the quality and realism of generated images, I introduce diffusion models grounded in natural image statistics, yielding perceptually aligned outputs suitable for downstream tasks. To advance personalization, I develop parameter-efficient mechanisms for combining low-rank adapters, enabling fine-grained control over content and style without retraining. I further extend personalization to a zero-shot setting through a training-free textual-inversion-based method that customizes arbitrary objects directly within the diffusion process. Finally, I present a frequency-guided multi-LoRA fusion framework that leverages wavelet-domain cues and timestep-aware weighting for accurate, training-free concept composition. Collectively, these contributions move toward a unified vision of generative models that are efficient, adaptive, and capable of high-quality, customizable image synthesis. Aniket Roy |
AAAI | 1 |
| 2025 | AeroGen: Ground-to-Air Generalization for Action RecognitionabstractWe address the problem of action recognition from aerial views using only ground-based videos for training. Due to the viewpoint-induced domain shift, models trained solely on ground videos exhibit significant performance degradation when naively applied to aerial videos. To mitigate this performance gap, we introduce a domain generalization technique that addresses the viewpoint-induced domain shift. Our method uses available real ground videos to generate additional synthetic training data from both ground and air viewpoints, for improving generalization to aerial video-based action recognition. Specifically, we perform 3D human mesh estimation from the ground videos, then render synthetic videos from alternate viewpoints with additional appearance randomizations. To further align the ground-air and syntheticreal domains, we propose a Dual Domain Alignment loss by enforcing consistency in predictions between the original ground videos and augmented videos from each domain. In order to facilitate research on the problem of ground-to-air generalization for human action recognition, we also create a benchmark by combining parts of the NTU-60, UAV-Human and NEC-DRONE datasets. We demonstrate the effectiveness of our approach on these new benchmarks along with the existing RoCoG-Ground $\rightarrow$ RoCoG-Air benchmark, and also perform extensive ablations. Ketul Shah, Anshul Shah 0001, Arun V. Reddy, Aniket Roy, Celso de Melo, Rama Chellappa |
FG | 4 |
| 2025 | DuoLoRA: Cycle-Consistent and Rank-Disentangled Content-Style Personalization
Aniket Roy, Shubhankar Borse, Shreya Kadambi, Debasmit Das, Shweta Mahajan, Risheek Garrepalli, Hyojin Park 0004, Ankita Nayak, Rama Chellappa, Munawar Hayat, Fatih Porikli |
ICCV | 1 |
| 2025 | DIFFUSE2ADAPT: Controlled Diffusion for Synthetic-to-Real Domain AdaptationabstractSynthetic data generated from graphics engines has been shown to be effective for learning, while also being a cost-effective alternative to annotating real-world data. However, models trained on synthetic data often suffer from performance degradation when applied to real-world data due to the domain gap. In this paper, we propose Diffuse2Adapt, a novel unsupervised domain adaptation (UDA) approach that leverages controlled diffusion models to bridge the synthetic-to-real domain gap. Our method utilizes text-to-image generative models to translate synthetic images to the target domain while preserving class semantics. We introduce two methods to reduce the domain gap: (1) incorporating target domain context extracted from multimodal language models, and (2) capturing target domain style via learned textual tokens. Extensive experiments are performed on three synthetic-to-real domain adaptation benchmarks, VisDA-2017, S2RDA-49, and S2RDA-MS-39. Diffuse2Adapt outperforms state-of-the-art methods by +3.00% on VisDA-2017, +2.55% on S2RDA-49 and +1.01% on S2RDA-MS-39. Code will be released. Ketul Shah, Arushi Sinha, Arun V. Reddy, Aniket Roy, Rama Chellappa |
ICIP | 4 |
| 2025 | Cap2Aug: Caption Guided Image data AugmentationabstractVisual recognition in a low-data regime is challenging and often prone to overfitting. To mitigate this issue, several data augmentation strategies have been proposed. However, standard transformations, e.g., rotation, cropping, and flip-ping provide limited semantic variations. To this end, we propose Cap2Aug, an image-to-image diffusion model-based data augmentation strategy using image captions to condition the image synthesis step. We generate a caption for an image and use this caption as an additional input for an image-to-image diffusion model. This increases the semantic diversity of the augmented images due to caption conditioning compared to the usual data augmentation techniques. We show that Cap2Aug is particularly effective where only a few samples are available for an object class. However, naively generating the synthetic images is not adequate due to the domain gap between real and synthetic images. Thus, we employ a maximum mean discrepancy loss to align the synthetic images to the real images to minimize the domain gap. We evaluate our method on few-shot classification and image classification with long-tail class distribution tasks. Cap2Aug achieves state-of-the-art performance on both tasks while evaluated on eleven benchmarks. Code: https://github.com/aniket004/Cap_2_Aug.git Aniket Roy, Anshul Shah 0001, Ketul Shah, Rama Chellappa |
WACV | 1 |
| 2024 | DiversiNet: Mitigating Bias in Deep Classification Networks across Sensitive Attributes through Diffusion-Generated DataabstractDeep learning models trained on sensitive data often show biases towards certain demographics, posing fairness challenges, especially with limited datasets. Diffusion generated data effectively supplement the underrepresented dataset, serving as a regularization technique to enhance feature learning. In addition to the original balanced dataset, we incorporate synthetic data generated by the diffusion model to train classifiers and subsequently assess their performance. Experimental results demonstrate a reduction in bias across all target attributes along with an increase in overall accuracy. For instance, for gender classification in the FFHQ dataset, the overall accuracy rises to 94.44% from 93.92% after including data generated from a diffusion model. Simultaneously, the bias, measured as the absolute difference between the true positive rates of young and old individuals, decreases from 0.0340 to 0.0204 (reduction of 40%). Moreover, we extend our analysis to multi-attribute scenarios, successfully mitigating bias with respect to multiple sensitive attributes simultaneously in sensitive attribute classification as well as in other downstream tasks. To the best of our knowledge, this study introduces a novel approach to bias mitigation, highlighting the versatility of diffusion-based data augmentation in addressing biases concerning age, gender, and race. Basudha Pal, Aniket Roy, Ram Prabhakar Kathirvel, Alice J. O'Toole, Rama Chellappa |
IJCB | 2 |
| 2024 | Bri3L: A Brightness Illusion Image Dataset for Identification and Localization of Regions of Illusory PerceptionabstractVisual illusions play a significant role in understanding visual perception. Current methods in understanding and evaluating visual illusions are mostly deterministic filtering based approach and they evaluate on a handful of visual illusions, and the conclusions therefore, are not generic. To this end, we generate a large-scale dataset of 22,366 images (BRI3L: BRightness Illusion Image dataset for Identification and Localization of illusory perception) of the five types of brightness illusions and benchmark the dataset using data-driven neural network based approaches. The dataset contains label information - (1) whether a particular image is illusory/nonillusory, (2) the segmentation mask of the illusory region of the image. Hence, both the classification and segmentation task can be evaluated using this dataset. We follow the standard psychophysical experiments involving human subjects to validate the dataset. To the best of our knowledge, this is the first attempt to develop a dataset of visual illusions and benchmark using data-driven approach for illusion classification and localization. We consider five well-studied types of brightness illusions: 1) Hermann grid, 2) Simultaneous Brightness Contrast, 3) White illusion, 4) Grid illusion, and 5) Induced Grating illusion. Benchmarking on the dataset achieves $99.56 \%$ accuracy in illusion identification and $84.37 \%$ pixel accuracy in illusion localization. The application of deep learning model, it is shown, also generalizes over unseen brightness illusions like brightness assimilation to contrast transitions. We also test the ability of state-of-the-art diffusion models to generate brightness illusions. We have provided all the code, dataset, instructions etc in the github repo: https://github.com/aniket004/BRI3L Aniket Roy, Soma Mitra, Kuntal Ghosh |
ICIP | 1 |
| 2023 | HaLP: Hallucinating Latent Positives for Skeleton-based Self-Supervised Learning of ActionsabstractSupervised learning of skeleton sequence encoders for action recognition has received significant attention in recent times. However, learning such encoders without labels continues to be a challenging problem. While prior works have shown promising results by applying contrastive learning to pose sequences, the quality of the learned representations is often observed to be closely tied to data augmentations that are used to craft the positives. However, augmenting pose sequences is a difficult task as the geometric constraints among the skeleton joints need to be enforced to make the augmentations realistic for that action. In this work, we propose a new contrastive learning approach to train models for skeleton-based action recognition without labels. Our key contribution is a simple module, HaLP - to Hallucinate Latent Positives for contrastive learning. Specifically, HaLP explores the latent space of poses in suitable directions to generate new positives. To this end, we present a novel optimization formulation to solve for the synthetic positives with an explicit control on their hardness. We propose approximations to the objective, making them solvable in closed form with minimal overhead. We show via experiments that using these generated positives within a standard contrastive learning framework leads to consistent improvements across benchmarks such as NTU-60, NTU-120, and PKU-II on tasks like linear evaluation, transfer learning, and kNN evaluation. Our code can be found at https://github.com/anshulbshah/HaLP. Anshul Shah 0001, Aniket Roy, Ketul Shah, Shlok Kumar Mishra, David Jacobs 0001, Anoop Cherian, Rama Chellappa |
CVPR | 2 |
| 2023 | Certified Robustness via Dynamic Margin Maximization and Improved Lipschitz RegularizationabstractTo improve the robustness of deep classifiers against adversarial perturbations, many approaches have been proposed, such as designing new architectures with better robustness properties (e.g., Lipschitz-capped networks), or modifying the training process itself (e.g., min-max optimization, constrained learning, or regularization). These approaches, however, might not be effective at increasing the margin in the input (feature) space. In this paper, we propose a differentiable regularizer that is a lower bound on the distance of the data points to the classification boundary. The proposed regularizer requires knowledge of the model's Lipschitz constant along certain directions. To this end, we develop a scalable method for calculating guaranteed differentiable upper bounds on the Lipschitz constant of neural networks accurately and efficiently. The relative accuracy of the bounds prevents excessive regularization and allows for more direct manipulation of the decision boundary. Furthermore, our Lipschitz bounding algorithm exploits the monotonicity and Lipschitz continuity of the activation layers, and the resulting bounds can be used to design new layers with controllable bounds on their Lipschitz constant. Experiments on the MNIST, CIFAR-10, and Tiny-ImageNet data sets verify that our proposed algorithm obtains competitively improved results compared to the state-of-the-art. Mahyar Fazlyab, Taha Entesari, Aniket Roy, Rama Chellappa |
NeurIPS | 3 |
| 2022 | FeLMi : Few shot Learning with hard MixupabstractLearning from a few examples is a challenging computer vision task. Traditionally,meta-learning-based methods have shown promise towards solving this problem.Recent approaches show benefits by learning a feature extractor on the abundantbase examples and transferring these to the fewer novel examples. However, thefinetuning stage is often prone to overfitting due to the small size of the noveldataset. To this end, we propose Few shot Learning with hard Mixup (FeLMi)using manifold mixup to synthetically generate samples that helps in mitigatingthe data scarcity issue. Different from a naïve mixup, our approach selects the hardmixup samples using an uncertainty-based criteria. To the best of our knowledge,we are the first to use hard-mixup for the few-shot learning problem. Our approachallows better use of the pseudo-labeled base examples through base-novel mixupand entropy-based filtering. We evaluate our approach on several common few-shotbenchmarks - FC-100, CIFAR-FS, miniImageNet and tieredImageNet and obtainimprovements in both 1-shot and 5-shot settings. Additionally, we experimented onthe cross-domain few-shot setting (miniImageNet → CUB) and obtain significantimprovements. Aniket Roy, Anshul Shah 0001, Ketul Shah, Prithviraj Dhar, Anoop Cherian, Rama Chellappa |
NeurIPS | 1 |
| 2022 | Multimodal Learning using Optimal Transport for Sarcasm and Humor DetectionabstractMultimodal learning is an emerging yet challenging research area. In this paper, we deal with multimodal sarcasm and humor detection from conversational videos and image-text pairs. Being a fleeting action, which is reflected across the modalities, sarcasm detection is challenging since large datasets are not available for this task in the literature. Therefore, we primarily focus on resource-constrained training, where the number of training samples is limited. To this end, we propose a novel multimodal learning system, MuLOT (Multimodal Learning using Optimal Transport), which utilizes self-attention to exploit intra-modal correspondence and optimal transport for cross-modal correspondence. Finally, the modalities are combined with multimodal attention fusion to capture the inter-dependencies across modalities. We test our approach for multimodal sarcasm and humor detection on three benchmark datasets - MUStARD [7] (video, audio, text), UR-FUNNY [20] (video, audio, text), MST [3] (image, text) and obtain 2.1%, 1.54% and 2.34% accuracy improvements over state-of-the-art. Shraman Pramanick, Aniket Roy, Vishal M. Patel |
WACV | 2 |
| 2021 | PASS: Protected Attribute Suppression System for Mitigating Bias in Face RecognitionabstractFace recognition networks encode information about sensitive attributes while being trained for identity classification. Such encoding has two major issues: (a) it makes the face representations susceptible to privacy leakage (b) it appears to contribute to bias in face recognition. However, existing bias mitigation approaches generally require end-to-end training and are unable to achieve high verification accuracy. Therefore, we present a descriptor-based adversarial de-biasing approach called ‘Protected Attribute Suppression System (PASS)’. PASS can be trained on top of descriptors obtained from any previously trained high-performing network to classify identities and simultaneously reduce encoding of sensitive attributes. This eliminates the need for end-to-end training. As a component of PASS, we present a novel discriminator training strategy that discourages a network from encoding protected attribute information. We show the efficacy of PASS to reduce gender and skintone information in descriptors from SOTA face recognition networks like Arcface. As a result, PASS descriptors outperform existing baselines in reducing gender and skintone bias on the IJB-C dataset, while maintaining a high verification accuracy. Prithviraj Dhar, Joshua Gleason, Aniket Roy, Carlos Domingo Castillo, Rama Chellappa |
ICCV | 3 |
| 2020 | Toward Optimal Prediction Error Expansion-Based Reversible Image WatermarkingabstractReversible image watermarking is a technique that allows the cover image to remain unmodified after watermark extraction. Prediction error expansion-based schemes are currently the most efficient and widely used class of reversible image watermarking techniques. In this paper, first, we prove that the bounded capacity distortion minimization problem for prediction error expansion-based reversible watermarking schemes is NP-hard, and the corresponding decision version of the problem is NP-complete. Then, we prove that the dual problem of bounded distortion capacity maximization problem for prediction error expansion-based reversible watermarking schemes is NP-hard, and the corresponding decision problem is NP-complete. Furthermore, taking advantage of the integer linear programming formulations of the optimization problems, we find the optimal performance metric values for a given image, using concepts from the optimal linear prediction theory. Our technique allows the calculation of these performance metric limit without assuming any particular prediction scheme. The experimental results for several common benchmark images are consistent with the calculated performance limits validate our approach. Aniket Roy, Rajat Subhra Chakraborty |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Copy move forgery detection with similar but genuine objectsabstractCopy-Move Forgery Detection (CMFD) is a well-studied image forensics problem. However, CMFD with Similar but Genuine Objects (SGO) has received relatively less attention. Recently, it has been found that current state-of-the-art CFMD techniques are mostly inadequate in satisfactorily solving this important problem variant. In this paper, we have addressed this issue by using Rotated Local Binary Pattern (RLBP) based rotation-invariant texture features, followed by Generalized Two Nearest Neighbourhood (g2NN) based feature matching, hierarchical clustering and geometric transformation estimation. Experimental results show that our technique outperforms the state-of-the-art CFMD techniques for forged images having similar but genuine objects, and matches the accuracy of state-of-the-art techniques for other copy-move forgery types. Our method is also robust with respect to filtering and compression based post-processing. Aniket Roy, Akhil Konda, Rajat Subhra Chakraborty |
ICIP | 1 |
| 2016 | Optimal Distortion Estimation for Prediction Error Expansion Based Reversible Watermarking
Aniket Roy, Rajat Subhra Chakraborty |
IWDW | 1 |