Aparna Bharati

dblp:141/2085 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-6404-9466ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Security and privacy · 6 · 4 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
abstract
The architectural blueprint of today’s leading text-to-image models contains a fundamental flaw: an inability to handle logical composition. This survey investigates this breakdown across three core primitives—negation, counting, and spatial relations. Our analysis reveals a dramatic performance collapse: models that are accurate on single primitives fail precipitously when these are combined, exposing severe interference. We trace this failure to three key factors. First, training data show a near-total absence of explicit negations. Second, continuous attention architectures are fundamentally unsuitable for discrete logic. Third, evaluation metrics reward visual plausibility over constraint satisfaction. By analyzing recent benchmarks and methods, we show that current solutions and simple scaling cannot bridge this gap. Achieving genuine compositionality, we conclude, will require fundamental advances in representation and reasoning rather than incremental adjustments to existing architectures.
Mayank Vatsa, Aparna Bharati, Richa Singh 0001
AAAI2
2026 Shades of Generalization: Diversity Aware Open Set Deepfake Detection
Swetcha Reddy Tukkani, Aparna Bharati
FG2
2026 Understanding Human-Like Biases in VLMs via Subjective Face Analytics
abstract
Vision-Language Models (VLMs) effectively integrate visual and textual information, often relying on shared embedding spaces to align modalities. However, the extent to which these spaces capture complex, subjective human judgments, such as perceived facial trustworthiness, and whether they replicate associated human social biases, remains underexplored. This paper investigates the representation of subjective face attributes within VLM embedding spaces, examining whether these representations encode human-like biases and assessing their interpretability. Using probing techniques on face datasets annotated with human judgments, we analyze the structure of VLM embeddings1. Our findings demonstrate that similarity scores between face image and textual description in the VLM embedding space align with human ratings of subjective attributes, and crucially, these representations exhibit correlations and demographic disparities mirroring known biases in human social perception. We also show that the use of variable context via face and attribute-specific captions can provide increased diagnostic value for revealing and mitigating bias, by increasing the alignment of the VLM embedding space with human impressions. Interpreting the embedded social biases highlights the need for critical evaluation and bias-aware development of VLMs to mitigate the risk of perpetuating harmful stereotypes in downstream applications that involve Human-AI interaction.
Chaitanya Roygaga, Aparna Bharati
WACV2
2025 Is Perturbation-Based Image Protection Disruptive to Image Editing?
abstract
The remarkable image generation capabilities of state-of-the-art diffusion models, such as Stable Diffusion, can also be misused to spread misinformation and plagiarize copyrighted materials. To mitigate the potential risks associated with image editing, current image protection methods rely on adding imperceptible perturbations to images to obstruct diffusion-based editing. A fully successful protection for an image implies that the output of editing attempts is an undesirable, noisy image which is completely unrelated to the reference image. In our experiments with various perturbation-based image protection methods across multiple domains (natural scene images and artworks) and editing tasks (image-to-image generation and style editing), we discover that such protection does not achieve this goal completely. In most scenarios, diffusion-based editing of protected images generates a desirable output image which adheres precisely to the guidance prompt. Our findings suggest that adding noise to images may paradoxically increase their association with given text prompts during the generation process, leading to unintended consequences such as better resultant edits. Hence, we argue that perturbation-based methods may not provide a sufficient solution for robust image protection against diffusion-based editing.1
Qiuyu Tang, Bonor Ayambem, Mooi Choo Chuah, Aparna Bharati
ICIP4
2025 Implications of Neural Compression to Scientific Images
abstract
While neural compression has the potential to revolutionize image compression, recent studies have emphasized its ability to introduce subtle artifacts that could alter the image content.Concerned about the impact of such modifications on scientific images, this work explores the potential effects of neural compression on these images, focusing on two critical aspects: semantic understanding and forensic integrity.We use scientific image datasets to assess the performance of neural compression techniques on Visual Question Answering (VQA) and copy-move forgery detection tasks.Our findings indicate that the subtle changes introduced by neural compression do not significantly degrade the performance of state-ofthe-art solutions.In the experiments, neurally compressed images sufficiently preserved the original semantics and forensic traces.Moreover, compared to lossy techniques, e.g., JPEG compression, at similar bit-per-pixel (bpp) rates, neural compression demonstrates a superior ability to preserve both semantic content and forensic traces, even at high compression levels.Our results suggest that neural compression may provide a viable alternative to lossy compression for scientific images.
João P. Cardenuto, Joshua Krinsky, Lucas Nogueira, Aparna Bharati, Daniel Moreira
IH&MMSec4
2025 Learning the Power of "No": Foundation Models with Negations
abstract
Negation is a fundamental aspect of natural language reasoning, yet foundational vision-language models (VLMs) like CLIP face significant challenges in accurately interpreting it. These models often process text prompts holistically, making it difficult to isolate and understand the role of negated terms. To overcome this limitation, we present CC-Neg: a novel dataset consisting of 228,246 images, each paired with both true captions and their corresponding negated versions. CC-Neg provides a critical benchmark to assess and improve foundational VLMs' ability to process negations, focusing specifically on how the presence of terms like ‘not’ alters the semantic relationship between images and their textual descriptions. To illustrate the effectiveness of the CC-Neg dataset in enhancing negation understanding, we introduce the CoN-CLIP framework, which incorporates targeted modifications to CLIP's contrastive loss function. When trained with CC-Neg, CoN-CLIP achieves a 3.85% average improvement in top-1 accuracy for zero-shot image classification across eight datasets, and a 4.4% performance boost on challenging compositionality benchmarks such as SugarCREPE. These results highlight CoN-CLIP's enhanced understanding of the nuanced semantic relationships involving negation. Our code and the CC-Neg benchmark are available at: https://github.com/jaisidhsingh/CoN-CLIP.
Jaisidh Singh, Ishaan Shrivastava, Mayank Vatsa, Richa Singh 0001, Aparna Bharati
WACV5
2024 A Relative Data Diversity Measure for Synthetic Face Images
abstract
Assessing the nature of synthetic images created by generative models is crucial for ensuring their usefulness for downstream visual recognition tasks. In addition to qualitative evaluation of realism, image generation processes are quantitatively assessed in terms of fidelity and diversity. Existing measures take into account proximity of generated data distribution with real training data, but lack understanding of key attributes characterizing diverse real datasets. In this work, we investigate the properties of generated synthetic face images with respect to the data used during training. We define a relative data diversity measure that captures both mutable and immutable aspects of face images and can represent gain/loss in diversity between two datasets. Through comprehensive experiments using GANs and DDPMs on two face image datasets — VGG-Face2 and BUPT-BalancedFace — we show that our data diversity measure captures variability among images in a more meaningful manner than existing metrics. Additionally, the proposed D2D measure can capture information leakage from training samples to generated images, highlighting privacy-related issues in the generation process. Since both privacy and diversity are desired properties for synthetic face image generation, the proposed measure, provides a more robust evaluation of generative methods than other existing measures.
Cancan Zhang, Chaitanya Roygaga, Aparna Bharati
IJCB3
2024 Exploring Saliency Bias in Manipulation Detection
abstract
The social media-fuelled explosion of fake news and misinformation supported by tampered images has led to growth in the development of models and datasets for image manipulation detection. However, existing detection methods mostly treat media objects in isolation, without considering the impact of specific manipulations on viewer perception. Forensic datasets are usually analyzed based on the manipulation operations and corresponding pixel-based masks, but not on the semantics of the manipulation, i.e., type of scene, objects, and viewers’ attention to scene content. The semantics of the manipulation play an important role in spreading misinformation through manipulated images. In an attempt to encourage further development of semantic-aware forensic approaches to understand visual misinformation, we propose a framework to analyze the trends of visual and semantic saliency in popular image manipulation datasets and their impact on detection. https://github.com/CV-Lehigh/Bias_IMD
Joshua Krinsky, Alan Bettis, Qiuyu Tang, Daniel Moreira, Aparna Bharati
ICIP5
2024 SynthProv: Interpretable Framework for Profiling Identity Leakage
abstract
Generative Adversarial Networks (GANs) can generate hyperrealistic face images of synthetic identities based on a latent understanding of real images from a large training set. Despite their proficiency, the term "synthetic identity" remains ambiguous, and the uniqueness of the faces GANs produce is rarely assessed. Recent studies have found that identities from the training data can unintentionally appear in the faces generated by StyleGAN2, but the cause of this phenomenon is unclear. In this work, we propose a novel framework, SynthProv, that utilizes the improved interpolation ability of StyleGAN2 latent space and employs image composition to analyze leakage. This is the first method that goes beyond detection and traces the source or provenance of constituent identity signals in the generated image. Experiments show that SynthProv succeeds in both detection and provenance tasks using multiple matching strategies. We identify identities from FFHQ and CelebA-HQ training datasets with the highest leakage into the latent space as "leaking reals". Analyzing latent space behavior to evaluate generative model privacy via leakage is an important research direction, as undetected leaking reals pose a significant threat to training data privacy. Our code is available at https://github.com/jaisidhsingh/SynthProv.
Jaisidh Singh, Harshil Bhatia, Mayank Vatsa, Richa Singh 0001, Aparna Bharati
WACV5
2023 IdProv: Identity-Based Provenance for Synthetic Image Generation (Student Abstract)
abstract
Recent advancements in Generative Adversarial Networks (GANs) have made it possible to obtain high-quality face images of synthetic identities. These networks see large amounts of real faces in order to learn to generate realistic looking synthetic images. However, the concept of a synthetic identity for these images is not very well-defined. In this work, we verify identity leakage from the training set containing real images into the latent space and propose a novel method, IdProv, that uses image composition to trace the source of identity signals in the generated image.
Harshil Bhatia, Jaisidh Singh, Gaurav Sangwan, Aparna Bharati, Richa Singh 0001, Mayank Vatsa
AAAI4
2022 In-group and Out-group Performance Bias in Facial Retouching Detection
abstract
Accuracy alone is not sufficient to establish the efficacy of an AI algorithm-issues of demographic bias are an important area of concern. Demographic bias in face recognition algorithms has attracted more attention from the re-search community to date, but bias can also be a problem for face image analysis algorithms, such as detection of manipulated face images. In this paper, we investigate performance of humans and algorithms at detecting retouched face images of subjects from different origin (America, India, China) and gender groups. To be representative of the state of retouching detection, we use eight different algorithms from the literature. In addition to overall human accuracy, differences across origin and gender of the human performing the task are analyzed. We observe different bias patterns, such as algorithms show higher in-group accuracy than out-group, while the extent of retouching and familiarity drives differences in detection accuracy for humans. This is the first work to analyze and compare bias exhibited by humans and algorithms in similar tasks of detecting retouched face images.
Aparna Bharati, Emma Connors, Mayank Vatsa, Richa Singh 0001, Kevin W. Bowyer
IJCB1
2021 Transformation-Aware Embeddings for Image Provenance
abstract
A dramatic rise in the flow of manipulated image content on the Internet has led to a prompt response from the media forensics research community. New mitigation efforts leverage cutting-edge data-driven strategies and increasingly incorporate usage of techniques from computer vision and machine learning to detect and profile the space of image manipulations. This paper addresses Image Provenance Analysis, which aims at discovering relationships among different manipulated image versions that share content. One important task in provenance analysis, like most visual understanding problems, is establishing a visual description and dissimilarity computation method that connects images that share full or partial content. But the existing handcrafted or learned descriptors - generally appropriate for tasks such as object recognition - may not sufficiently encode the subtle differences between near-duplicate image variants, which significantly characterize the provenance of any image. This paper introduces a novel data-driven learning-based approach that provides the context for ordering images that have been generated from a single image source through various transformations. Our approach learns transformation-aware embeddings using weak supervision via composited transformations and a rank-based Edit Sequence Loss. To establish the effectiveness of the proposed approach, comparisons are made with state-of-the-art handcrafted and deep-learning-based descriptors, as well as image matching approaches. Further experimentation validates the proposed approach in the context of image provenance analysis and improves upon existing approaches.
Aparna Bharati, Daniel Moreira, Patrick J. Flynn, Anderson Rocha 0001, Kevin W. Bowyer, Walter J. Scheirer
IEEE Trans. Inf. Forensics Secur.1
2021 Fast Local Spatial Verification for Feature-Agnostic Large-Scale Image Retrieval
abstract
Images from social media can reflect diverse viewpoints, heated arguments, and expressions of creativity, adding new complexity to retrieval tasks. Researchers working on Content-Based Image Retrieval (CBIR) have traditionally tuned their algorithms to match filtered results with user search intent. However, we are now bombarded with composite images of unknown origin, authenticity, and even meaning. With such uncertainty, users may not have an initial idea of what the search query results should look like. For instance, hidden people, spliced objects, and subtly altered scenes can be difficult for a user to detect initially in a meme image, but may contribute significantly to its composition. It is pertinent to design systems that retrieve images with these nuanced relationships in addition to providing more traditional results, such as duplicates and near-duplicates - and to do so with enough efficiency at large scale. We propose a new approach for spatial verification that aims at modeling object-level regions using image keypoints retrieved from an image index, which is then used to accurately weight small contributing objects within the results, without the need for costly object detection steps. We call this method the Objects in Scene to Objects in Scene (OS2OS) score, and it is optimized for fast matrix operations, which can run quickly on either CPUs or GPUs. It performs comparably to state-of-the-art methods on classic CBIR problems (Oxford 5K, Paris 6K, and Google-Landmarks), and outperforms them in emerging retrieval tasks such as image composite matching in the NIST MFC2018 dataset and meme-style imagery from Reddit.
Joel Brogan, Aparna Bharati, Daniel Moreira, Anderson Rocha 0001, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer
IEEE Trans. Image Process.2
2019 Beyond Pixels: Image Provenance Analysis Leveraging Metadata
abstract
Creative works, whether paintings or memes, follow unique journeys that result in their final form. Understanding these journeys, a process known as "provenance analysis," provides rich insights into the use, motivation, and authenticity underlying any given work. The application of this type of study to the expanse of unregulated content on the Internet is what we consider in this paper. Provenance analysis provides a snapshot of the chronology and validity of content as it is uploaded, re-uploaded, and modified over time. Although still in its infancy, automated provenance analysis for online multimedia is already being applied to different types of content. Most current works seek to build provenance graphs based on the shared content between images or videos. This can be a computationally expensive task, especially when considering the vast influx of content that the Internet sees every day. Utilizing non-content-based information, such as timestamps, geotags, and camera IDs can help provide important insights into the path a particular image or video has traveled during its time on the Internet without large computational overhead. This paper tests the scope and applicability of metadata-based inferences for provenance graph construction in two different scenarios: digital image forensics and cultural analytics.
Aparna Bharati, Daniel Moreira, Joel Brogan, Patricia Hale, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer
WACV1
2018 To Frontalize or Not to Frontalize: Do We Really Need Elaborate Pre-processing to Improve Face Recognition?
abstract
Face recognition performance has improved remarkably in the last decade. Much of this success can be attributed to the development of deep learning techniques such as convolutional neural networks (CNNs). While CNNs have pushed the state-of-the-art forward, their training process requires a large amount of clean and correctly labelled training data. If a CNN is intended to tolerate facial pose, then we face an important question: should this training data be diverse in its pose distribution, or should face images be normalized to a single pose in a pre-processing step? To address this question, we evaluate a number of facial landmarking algorithms and a popular frontalization method to understand their effect on facial recognition performance. Additionally, we introduce a new, automatic, single-image frontalization scheme that exceeds the performance of the reference frontalization algorithm for video-to-video face matching on the Point and Shoot Challenge (PaSC) dataset. Additionally, we investigate failure modes of each frontalization method on different facial yaw using the CMU Multi-PIE dataset. We assert that the subsequent recognition and verification performance serves to quantify the effectiveness of each pose correction scheme.
Sandipan Banerjee, Joel Brogan, Janez Krizaj, Aparna Bharati, Brandon RichardWebster, Vitomir Struc, Patrick J. Flynn, Walter J. Scheirer
WACV4
2018 Image Provenance Analysis at Scale
abstract
Prior art has shown it is possible to estimate, through image processing and computer vision techniques, the types and parameters of transformations that have been applied to the content of individual images to obtain new images. Given a large corpus of images and a query image, an interesting further step is to retrieve the set of original images whose content is present in the query image, as well as the detailed sequences of transformations that yield the query image given the original images. This is a problem that recently has received the name of image provenance analysis. In these times of public media manipulation (e.g., fake news and meme sharing), obtaining the history of image transformations is relevant for fact checking and authorship verification, among many other applications. This article presents an end-to-end processing pipeline for image provenance analysis, which works at real-world scale. It employs a cutting-edge image filtering solution that is custom-tailored for the problem at hand, as well as novel techniques for obtaining the provenance graph that expresses how the images, as nodes, are ancestrally connected. A comprehensive set of experiments for each stage of the pipeline is provided, comparing the proposed solution with state-of-the-art results, employing previously published datasets. In addition, this work introduces a new dataset of real-world provenance cases from the social media site Reddit, along with baseline results.
Aparna Bharati, Joel Brogan, Allan Pinto, Michael Parowski, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer
IEEE Trans. Image Process.1
2017 Demography-based facial retouching detection using subclass supervised sparse autoencoder
abstract
Digital retouching of face images is becoming more widespread due to the introduction of software packages that automate the task. Several researchers have introduced algorithms to detect whether a face image is original or retouched. However, previous work on this topic has not considered whether or how accuracy of retouching detection varies with the demography of face images. In this paper, we introduce a new Multi-Demographic Retouched Faces (MDRF) dataset, which contains images belonging to two genders, male and female, and three ethnicities, Indian, Chinese, and Caucasian. Further, retouched images are created using two different retouching software packages. The second major contribution of this research is a novel semi-supervised autoencoder incorporating “sub-class” information to improve classification. The proposed approach outperforms existing state-of-the-art detection algorithms for the task of generalized retouching detection. Experiments conducted with multiple combinations of ethnicities show that accuracy of retouching detection can vary greatly based on the demographics of the training and testing images.
Aparna Bharati, Mayank Vatsa, Richa Singh 0001, Kevin W. Bowyer
IJCB1
2017 U-Phylogeny: Undirected provenance graph construction in the wild
abstract
Deriving relationships between images and tracing back their history of modifications are at the core of Multimedia Phylogeny solutions, which aim to combat misinformation through doctored visual media. Nonetheless, most recent image phylogeny solutions cannot properly address cases of forged composite images with multiple donors, an area known as multiple parenting phylogeny (MPP). This paper presents a preliminary undirected graph construction solution for MPP, without any strict assumptions. The algorithm is underpinned by robust image representative keypoints and different geometric consistency checks among matching regions in both images to provide regions of interest for direct comparison. The paper introduces a novel technique to geometrically filter the most promising matches as well as to aid in the shared region localization task. The strength of the approach is corroborated by experiments with real-world cases, with and without image distractors (unrelated cases).
Aparna Bharati, Daniel Moreira, Allan Pinto, Joel Brogan, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer, Anderson Rocha 0001
ICIP1
2017 Spotting the difference: Context retrieval and analysis for improved forgery detection and localization
abstract
As image tampering becomes ever more sophisticated and commonplace, the need for image forensics algorithms that can accurately and quickly detect forgeries grows. In this paper, we revisit the ideas of image querying and retrieval to provide clues to better localize forgeries. We propose a method to perform large-scale image forensics on the order of one million images using the help of an image search algorithm and database to gather contextual clues as to where tampering may have taken place. In this vein, we introduce five new strongly invariant image comparison methods and test their effectiveness under heavy noise, rotation, and color space changes. Lastly, we show the effectiveness of these methods compared to passive image forensics using Nimble [1], a new, state-of-the-art dataset from the National Institute of Standards and Technology (NIST).
Joel Brogan, Paolo Bestagini, Aparna Bharati, Allan Pinto, Daniel Moreira, Kevin W. Bowyer, Patrick J. Flynn, Anderson Rocha 0001, Walter J. Scheirer
ICIP3
2017 Provenance filtering for multimedia phylogeny
abstract
Departing from traditional digital forensics modeling, which seeks to analyze single objects in isolation, multimedia phylogeny analyzes the evolutionary processes that influence digital objects and collections over time. One of its integral pieces is provenance filtering, which consists of searching a potentially large pool of objects for the most related ones with respect to a given query, in terms of possible ancestors (donors or contributors) and descendants. In this paper, we propose a two-tiered provenance filtering approach to find all the potential images that might have contributed to the creation process of a given query q. In our solution, the first (coarse) tier aims to find the most likely “host” images - the major donor or background - contributing to a composite/doctored image. The search is then refined in the second tier, in which we search for more specific (potentially small) parts of the query that might have been extracted from other images and spliced into the query image. Experimental results with a dataset containing more than a million images show that the two-tiered solution underpinned by the context of the query is highly useful for solving this difficult task.
Allan Pinto, Daniel Moreira, Aparna Bharati, Joel Brogan, Kevin W. Bowyer, Patrick J. Flynn, Walter J. Scheirer, Anderson Rocha 0001
ICIP3
2016 Detecting Facial Retouching Using Supervised Deep Learning
abstract
Digitally altering, or retouching, face images is a common practice for images on social media, photo sharing websites, and even identification cards when the standards are not strictly enforced. This research demonstrates the effect of digital alterations on the performance of automatic face recognition, and also introduces an algorithm to classify face images as original or retouched with high accuracy. We first introduce two face image databases with unaltered and retouched images. Face recognition experiments performed on these databases show that when a retouched image is matched with its original image or an unaltered gallery image, the identification performance is considerably degraded, with a drop in matching accuracy of up to 25%. However, when images are retouched with the same style, the matching accuracy can be misleadingly high in comparison with matching original images. To detect retouching in face images, a novel supervised deep Boltzmann machine algorithm is proposed. It uses facial parts to learn discriminative features to classify face images as original or retouched. The proposed approach for classifying images as original or retouched yields an accuracy of over 87% on the data sets introduced in this paper and over 99% on three other makeup data sets used by previous researchers. This is a substantial increase in accuracy over the previous state-of-the-art algorithm, which has shown <;50% accuracy in classifying original and retouched images from the ND-IIITD retouched faces database.
Aparna Bharati, Richa Singh 0001, Mayank Vatsa, Kevin W. Bowyer
IEEE Trans. Inf. Forensics Secur.1