EDBT 2026 Demo / reviewers in the wild / expert
Kartik Thakral
dblp:256/4230
· DBLP profile ↗
12ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-2528-9950ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation ModelsabstractExisting unlearning algorithms in text-to-image generative models often fail to preserve the knowledge of semantically related concepts when removing specific target concepts—a challenge known as adjacency. To address this, we propose FADE (Fine-Grained Attenuation for Diffusion Erasure), introducing adjacency-aware unlearning in diffusion models. FADE comprises two components: (1) the Concept Neighborhood, which identifies an adjacency set of related concepts, and (2) Mesh Modules, employing a structured combination of Expungement, Adjacency, and Guidance loss components. These enable precise erasure of target concepts while preserving fidelity across related and unrelated concepts. Evaluated on datasets like Stanford Dogs, Oxford Flowers, CUB, I2P, Imagenette, and ImageNet-1k, FADE effectively removes target concepts with minimal impact on correlated concepts, achieving at least a 12% improvement in retention performance over state-of-the-art methods. Our code and models are available on the project page: iab-rubric/unlearning/FG-Un. Kartik Thakral, Tamar Glaser, Tal Hassner, Mayank Vatsa, Richa Singh 0001 |
CVPR | 1 |
| 2025 | ILLUSION: Unveiling Truth with a Comprehensive Multi-Modal, Multi-Lingual Deepfake DatasetabstractThe proliferation of deepfakes and AI-generated content has led to a surge in media forgeries and misinformation, necessitating robust detection systems. However, current datasets lack diversity across modalities, languages, and real-world scenarios. To address this gap, we present ILLUSION (Integration of Life-Like Unique Synthetic Identities and Objects from Neural Networks), a large-scale, multi-modal
deepfake dataset comprising 1.3 million samples spanning audio-visual forgeries, 26 languages, challenging noisy environments, and various manipulation protocols. Generated using 28 state-of-the-art generative techniques, ILLUSION includes
faceswaps, audio spoofing, synchronized audio-video manipulations, and synthetic media while ensuring a balanced representation of gender and skin tone for unbiased evaluation. Using Jaccard Index and UpSet plot analysis, we demonstrate ILLUSION’s distinctiveness and minimal overlap with existing datasets, emphasizing its novel generative coverage. We benchmarked image, audio, video, and multi-modal detection models, revealing key challenges such as performance degradation in multilingual and multi-modal contexts, vulnerability to real-world distortions, and limited generalization to zero-day attacks. By bridging synthetic and real-world complexities, ILLUSION provides a challenging yet essential platform for advancing deepfake detection research. The dataset is publicly available at https://www.iab-rubric.org/illusion-database. Kartik Thakral, Rishabh Ranjan, Akshat Jain, Mayank Vatsa, Richa Singh 0001 |
ICLR | 1 |
| 2025 | Harmonizing Geometry and Uncertainty: Diffusion with HyperspheresabstractDo contemporary diffusion models preserve the class geometry of hyperspherical data? Standard diffusion models rely on isotropic Gaussian noise in the forward process, inherently favoring Euclidean spaces. However, many real-world problems involve non-Euclidean distributions, such as hyperspherical manifolds, where class-specific patterns are governed by angular geometry within hypercones. When modeled in Euclidean space, these angular subtleties are lost, leading to suboptimal generative performance. To address this limitation, we introduce \textbf{HyperSphereDiff} to align hyperspherical structures with directional noise, preserving class geometry and effectively capturing angular uncertainty. We demonstrate both theoretically and empirically that this approach aligns the generative process with the intrinsic geometry of hyperspherical data, resulting in more accurate and geometry-aware generative models. We evaluate our framework on four object datasets and two face datasets, showing that incorporating angular uncertainty better preserves the underlying hyperspherical manifold. Muskan Dosi, Chiranjeev Chiranjeev, Kartik Thakral, Mayank Vatsa, Richa Singh 0001 |
ICML | 3 |
| 2025 | LitMAS: A Lightweight and Generalized Multi-Modal Anti-Spoofing Framework for Biometric Security
Nidheesh Gorthi, Kartik Thakral, Rishabh Ranjan, Richa Singh 0001, Mayank Vatsa |
INTERSPEECH | 2 |
| 2024 | ToonerGAN: Reinforcing GANs for Obfuscating Automated Facial IndexingabstractThe rapid evolution of automatic facial indexing technologies increases the risk of compromising personal and sensitive information. To mitigate the issue, we propose creating cartoon avatars, or ‘toon avatars', designed to effectively obscure identity features. The primary objective is to deceive current AI systems, preventing them from accurately identifying individuals while making minimal modifications to their facial features. Moreover, we aim to ensure that a human observer can still recognize the person depicted in these altered avatar images. To achieve this, we introduce ‘ToonerGAN’, a novel approach that utilizes Generative Adversarial Networks (GANs) to craft personalized cartoon avatars. The ToonerGAN framework consists of a style and a de-identification module that work together to produce high-resolution, realistic cartoon images. For the efficient training of our network, we have developed ‘ToonSet’ dataset, consisting of around 23,000 facial images and their cartoon renditions. Through comprehensive experiments and benchmarking against existing datasets, including CelebA-HQ, our method demonstrates superior performance in obfuscating identity while preserving the utility of data. Additionally, a user-centric study exploring the effectiveness of ToonerGAN has yielded compelling observations. Kartik Thakral, Shashikant Prasad, Stuti Aswani, Mayank Vatsa, Richa Singh 0001 |
CVPR | 1 |
| 2024 | HyperSpaceX: Radial and Angular Exploration of HyperSpherical Dimensions
Chiranjeev Chiranjeev, Muskan Dosi, Kartik Thakral, Mayank Vatsa, Richa Singh 0001 |
ECCV (88) | 3 |
| 2023 | DF-Platter: Multi-Face Heterogeneous Deepfake DatasetabstractDeepfake detection is gaining significant importance in the research community. While most of the research efforts are focused towards high-quality images and videos with controlled appearance of individuals, deepfake generation algorithms now have the capability to generate deep-fakes with low-resolution, occlusion, and manipulation of multiple subjects. In this research, we emulate the real-world scenario of deepfake generation and propose the DF-Platter dataset, which contains (i) both low-resolution and high-resolution deepfakes generated using multiple generation techniques and (ii) single-subject and multiple-subject deepfakes, with face images of Indian ethnicity. Faces in the dataset are annotated for various attributes such as gender, age, skin tone, and occlusion. The dataset is prepared in 116 days with continuous usage of 32 GPUs accounting to 1,800 GB cumulative memory. With over 500 GBs in size, the dataset contains a total of 133,260 videos encompassing three sets. To the best of our knowledge, this is one of the largest datasets containing vast variability and multiple challenges. We also provide benchmark results under multiple evaluation settings using popular and state-of-the-art deepfake detection models, for c0 images and videos along with c23 and c40 compression variants. The results demonstrate a significant performance reduction in the deepfake detection task on low-resolution deep-fakes. Furthermore, existing techniques yield declined detection accuracy on multiple-subject deepfakes. It is our assertion that this database will improve the state-of-the-art by extending the capabilities of deepfake detection algorithms to real-world scenarios. The database is available at: http://iab-rubric.org/df-platter-database. Kartik Narayan, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, Richa Singh 0001 |
CVPR | 3 |
| 2023 | Are Face Detection Models Biased?abstractThe presence of bias in deep models leads to unfair outcomes for certain demographic subgroups. Research in bias focuses primarily on facial recognition and attribute prediction with scarce emphasis on face detection. Existing studies consider face detection as binary classification into ‘face’ and ‘non-face’ classes. In this work, we investigate possible bias in the domain of face detection through facial region localization which is currently unexplored. Since facial region localization is an essential task for all face recognition pipelines, it is imperative to analyze the presence of such bias in popular deep models. Most existing face detection datasets lack suitable annotation for such analysis. Therefore, we web-curate the Fair Face Localization with Attributes (F2LA) dataset and manually annotate more than 10 attributes per face, including facial localization information. Utilizing the extensive annotations from F2LA, an experimental setup is designed to study the performance of four pre-trained face detectors. We observe (i) a high disparity in detection accuracies across gender and skin-tone, and (ii) interplay of confounding factors beyond demography. The F2LA data and associated annotations can be accessed at http://iab-rubric.org/index.php/F2LA. Surbhi Mittal, Kartik Thakral, Puspita Majumdar, Mayank Vatsa, Richa Singh 0001 |
FG | 2 |
| 2023 | PhygitalNet: Unified Face Presentation Attack Detection via One-Class Isolation LearningabstractFace biometric systems are shown to be vulnerable to various kinds of presentation attacks including physical and digital attacks. Existing research generally focuses on individual attacks and very few focus on generalizability across digital and physical attacks. In this research, we propose PhygitalNet model that generalizes to both physical and digital presentation attacks on face biometric systems. The proposed model is based on novel one-class iSOLatiOn Learning (SOLO Learning) which is a two-step training process aimed at reducing of the covariate shift between the bonafide samples of the physical as well as digital attack dataset in the pre-training step. In the downstream step, the algorithm introduces a novel single-class iSOLatiOn loss (SOLO loss) function that isolates the samples belonging to the bonafide class away from the samples of the attacked class for both the attack methods. Experimental results show that PhygitalNet achieves a significant performance gain when compared with the baseline techniques, evaluated on a combination of MLFP, MSU-MFSD dataset (for physical attack) and FaceForensics++ (for digital attack) datasets. Kartik Thakral, Surbhi Mittal, Mayank Vatsa, Richa Singh 0001 |
FG | 1 |
| 2022 | DeePhy: On Deepfake PhylogenyabstractDeepfake refers to tailored and synthetically generated videos which are now prevalent and spreading on a large scale, threatening the trustworthiness of the information available online. While existing datasets contain different kinds of deepfakes which vary in their generation technique, they do not consider progression of deepfakes in a “phylogenetic” manner. It is possible that an existing deepfake face is swapped with another face. This process of face swapping can be performed multiple times and the resultant deepfake can be evolved to confuse the deepfake detection algorithms. Further, many databases do not provide the employed generative model as target labels. Model attribution helps in enhancing the explainability of the detection results by providing information on the generative model employed. In order to enable the research community to address these questions, this paper proposes DeePhy, a novel Deepfake Phylogeny dataset which consists of 5040 deep-fake videos generated using three different generation techniques. There are 840 videos of one-time swapped deep-fakes, 2520 videos of two-times swapped deepfakes and 1680 videos of three-times swapped deepfakes. With over 30 GBs in size, the database is prepared in over 1100 hours using 18 GPUs of 1,352 GB cumulative memory. We also present the benchmark on DeePhy dataset using six deep-fake detection algorithms. The results highlight the need to evolve the research of model attribution of deepfakes and generalize the process over a variety of deepfake generation techniques. The database is available at: http://iab-rubric.org/deephy-database Kartik Narayan, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, Richa Singh 0001 |
IJCB | 3 |
| 2022 | Multi-task driven explainable diagnosis of COVID-19 using chest X-ray images
Aakarsh Malhotra, Surbhi Mittal, Puspita Majumdar, Saheb Chhabra, Kartik Thakral, Mayank Vatsa, Richa Singh 0001, Santanu Chaudhury, Ashwin Pudrod, Anjali Agrawal |
Pattern Recognit. | 5 |
| 2021 | AECNet: Attentive EfficientNet For Crowd CountingabstractIn the COVID pandemic situation, crowd counting became one of the tools to monitor if the social-distancing norms are being followed or not. However, in designing crowd counting algorithm, there are several challenges such as background noise, camera-to-objects distance, occlusion, and variations due to illumination, scale, and viewpoint. In this research, we propose a novel pipeline for density estimation in crowd counting. The proposed pipeline makes use of an encoder-decoder-based architecture in which we explore the family of EfficientN ets for the encoder architecture. For the decoder, we propose a deeper attention network to assist the model in a better distinction between foreground and background pixels. We empirically show that for a crowd counting dataset, the use of average pooling operation for any backbone architecture of encoder gives a significant improvement in performance. In terms of Mean Absolute Error, the proposed pipeline outperforms existing state-of-the-art techniques by a large margin on large-scale and small-scale counting datasets, UCF-QNRF and UCF _CC_50 dataset. We also achieve state-of-the-art results on the ShanghaiTech and Mall datasets. We additionally propose a crowd counting dataset captured using drones. We perform benchmark experiments on this dataset with existing and the proposed methods. The proposed dataset can be found at http://www.iab-rubric.org/resources/CrowdUAV.html. Muskan Dosi, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, Richa Singh 0001 |
FG | 2 |