EDBT 2026 Demo / reviewers in the wild / expert
Ashwin Swaminathan
dblp:61/754
· DBLP profile ↗
34ranked-venue papers
13as first author
11since 2021 · last 2025
0000-0002-4279-369XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 11 since 2021Security and privacy · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 3 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling up Image Segmentation across Data and TasksabstractTraditional segmentation models, while effective in isolated tasks, often fail to generalize to more complex and open-ended segmentation problems, such as free-form, open-vocabulary, and in-the-wild scenarios. To bridge this gap, we propose to scale up image segmentation across diverse datasets and tasks such that the knowledge across different tasks and datasets can be integrated while improving the generalization ability. Mixed-Query Transformer (MQ-Former), a novel segmentation framework, is introduced and designed to scale seamlessly across both data size and task diversity. It is built upon a dynamic object query mechanism called mixed query, which fuses different types of queries using cross-attention. This hybrid approach enables the model to balance between instance- and stuff-level segmentation, providing enhanced scalability for handling diverse object types. We further enhance scalability by leveraging synthetic data-generating segmentation masks and captions for pixel-level and open-vocabulary tasks-drastically reducing the need for costly human annotations. By training on multiple datasets and tasks at scale, MQ-Former continuously improves performance as the volume and diversity of data and tasks increase. It exhibits strong generalization capabilities, boosting performance in open-set segmentation tasks SeginW by 7 points. These advancements mark a key step toward universal, scalable segmentation models capable of addressing the demands of real-world applications. Zhaowei Cai, Hao Yang 0043, Ashwin Swaminathan, R. Manmatha, Stefano Soatto |
CVPR | 4 |
| 2025 | Enhancing Numerical Prediction of MLLMS With Soft Labeling
Zhaowei Cai, Hao Yang 0043, Davide Modolo, Ashwin Swaminathan |
ICCV | 5 |
| 2024 | Multi-Modal Hallucination Control by Visual Information GroundingabstractGenerative Vision-Language Models (VLMs) are prone to generate plausible-sounding textual answers that, however, are not always grounded in the input image. We investigate this phenomenon, usually referred to as “hallucination” and show that it stems from an excessive reliance on the language prior. In particular, we show that as more tokens are generated, the reliance on the visual prompt decreases, and this behavior strongly correlates with the emergence of hallucinations. To reduce hallucinations, we introduce Multi-Modal Mutual-Information Decoding (M3ID), a new sampling method for prompt amplification. M3ID amplifies the influence of the reference image over the language prior, hence favoring the generation of tokens with higher mutual information with the visual prompt. M3ID can be applied to any pre-trained autoregressive VLM at inference time without necessitating further training and with minimal computational overhead. If training is an option, we show that M3ID can be paired with Direct Preference Optimization (DPO) to improve the model's reliance on the prompt image without requiring any labels. Our empirical findings show that our algorithms maintain the fluency and linguistic capabilities of pre-trained VLMs while reducing hallucinations by mitigating visually ungrounded answers. Specifically, for the LLaVA 13B model, M3ID and M3ID+DPO reduce the percentage of hallucinated objects in captioning tasks by 25% and 28%, respectively, and improve the accuracy on VQA benchmarks such as POPE by 21% and 24%. Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, Stefano Soatto |
CVPR | 7 |
| 2024 | CPR: Retrieval Augmented Generation for Copyright ProtectionabstractRetrieval Augmented Generation (RAG) is emerging as a flexible and robust technique to adapt models to private users data without training, to handle credit attribution, and to allow efficient machine unlearning at scale. However, RAG techniques for image generation may lead to parts of the retrieved samples being copied in the model's output. To reduce risks of leaking private information contained in the retrieved set, we introduce Copy-Protected generation with Retrieval (CPR), a new method for RAG with strong copyright protection guarantees in a mixed-private setting for diffusion models. CPR allows to condition the output of diffusion models on a set of retrieved images, while also guaranteeing that unique identifiable information about those example is not exposed in the generated outputs. In particular, it does so by sampling from a mixture of public (safe) distribution and private (user) distribution by merging their diffusion scores at inference. We prove that CPR satisfies Near Access Freeness (NAF) which bounds the amount of information an attacker may be able to extract from the generated images. We provide two algorithms for copyright protection, CPR-KL and CPR-Choose. Unlike previously proposed rejection-sampling-based NAF methods, our methods enable efficient copyright-protected sampling with a single run of backward diffusion. We show that our method can be applied to any pre-trained conditional diffusion model, such as Stable Diffusion or unCLIP. In particular, we empirically show that applying CPR on top of unCLIP improves quality and text-to-image alignment of the generated results (81.4 to 83.17 on TIFA benchmark), while enabling credit attribution, copy-right protection, and deterministic, constant time, unlearning. Aditya Golatkar, Alessandro Achille, Luca Zancato, Yu-Xiang Wang 0003, Ashwin Swaminathan, Stefano Soatto |
CVPR | 5 |
| 2024 | THRONE: An Object-Based Hallucination Benchmark for the Free-Form Generations of Large Vision-Language ModelsabstractMitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which we term “Type I hallucinations”. Instead, they focus on hallucinations responding to very specific question formats-typically a multiple-choice response regarding a particular object or attribute-which we term “Type II hallucinations”. Additionally, such benchmarks often require external API calls to models which are subject to change. In practice, we observe that a reduction in Type II hallucinations does not lead to a reduction in Type I hallucinations but rather that the two forms of halluci-nations are often anti-correlated. To address this, we propose THRONE, a novel object-based automatic framework for quantitatively evaluating Type I hallucinations in LVLM free-form outputs. We use public language models (LMs) to identify hallucinations in LVLM responses and compute informative metrics. By evaluating a large selection of recent LVLMs using public datasets, we show that an improvement in existing metrics do not lead to a reduction in Type I hallucinations, and that established benchmarks for measuring Type I hallucinations are incomplete. Finally, we provide a simple and effective data augmentation method to reduce Type I and Type II hallucinations as a strong baseline. Prannay Kaul, Zhizhong Li 0001, Hao Yang 0043, Yonatan Dukler, Ashwin Swaminathan, C. J. Taylor, Stefano Soatto |
CVPR | 5 |
| 2024 | On the Scalability of Diffusion-based Text-to-Image GenerationabstractScaling up model and data size has been quite successful for the evolution of LLMs. However, the scaling law for the diffusion based text-to-image (T2I) models is not fully explored. It is also unclear how to efficiently scale the model for better performance at reduced cost. The different training settings and expensive training cost make a fair model comparison extremely difficult. In this work, we empirically study the scaling properties of diffusion based T2I models by performing extensive and rigours ablations on scaling both denoising backbones and training set, including training scaled UNet and Transformer variants ranging from 0.4B to 4B parameters on datasets upto 600M images. For model scaling, we find the location and amount of cross attention distinguishes the performance of existing UNet designs. And increasing the transformer blocks is more parameter-efficient for improving text-image alignment than increasing channel numbers. We then identify an efficient UNet variant, which is 45% smaller and 28% faster than SDXL's UNet. On the data scaling side, we show the quality and diversity of the training set matters more than simply dataset size. Increasing caption density and diversity improves text-image alignment performance and the learning efficiency. Finally, we provide scaling functions to predict the text-image alignment performance as functions of the scale of model size, compute and dataset size. Orchid Majumder, Yusheng Xie, R. Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, Stefano Soatto |
CVPR | 7 |
| 2024 | Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
Benjamin Biggs, Arjun Seshadri, Achin Jain, Aditya Golatkar, Yusheng Xie, Alessandro Achille, Ashwin Swaminathan, Stefano Soatto |
ECCV (63) | 8 |
| 2023 | Learning Expressive Prompting With Residuals for Vision TransformersabstractPrompt learning is an efficient approach to adapt transformers by inserting learnable set of parameters into the input and intermediate representations of a pre-trained model. In this work, we present Expressive Prompts with Residuals (EXPRES) which modifies the prompt learning paradigm specifically for effective adaptation of vision transformers (ViT). Our method constructs downstream representations via learnable “output” tokens (shal-low prompts), that are akin to the learned class tokens of the ViT. Further for better steering of the downstream representation processed by the frozen transformer, we introduce residual learnable tokens that are added to the output of various computations. We apply EXPRES for image classification and few-shot semantic segmentation, and show our method is capable of achieving state of the art prompt tuning on 3/3 categories of the VTAB benchmark. In addition to strong performance, we observe that our approach is an order of magnitude more prompt efficient than existing visual prompting baselines. We analytically show the computational benefits of our approach over weight space adaptation techniques like finetuning. Lastly we systematically corroborate the architectural design of our method via a series of ablation experiments. Rajshekhar Das, Yonatan Dukler, Avinash Ravichandran, Ashwin Swaminathan |
CVPR | 4 |
| 2023 | A Meta-Learning Approach to Predicting Performance and Data RequirementsabstractWe propose an approach to estimate the number of samples required for a model to reach a target performance. We find that the power law, the de facto principle to estimate model performance, leads to a large error when using a small dataset (e.g., 5 samples per class) for extrapolation. This is because the log-performance error against the log-dataset size follows a nonlinear progression in the few-shot regime followed by a linear progression in the high-shot regime. We introduce a novel piecewise power law (PPL) that handles the two data regimes differently. To estimate the parameters of the PPL, we introduce a random forest regressor trained via meta learning that generalizes across classification/detection tasks, ResNet/ViT based architectures, and random/pre-trained initializations. The PPL improves the performance estimation on average by 37% across 16 classification and 33% across 10 detection datasets, compared to the power law. We further extend the PPL to provide a confidence bound and use it to limit the prediction horizon that reduces over-estimation of data by 76% on classification and 91% on detection datasets. Achin Jain, Gurumurthy Swaminathan, Paolo Favaro, Hao Yang 0043, Avinash Ravichandran, Hrayr Harutyunyan, Alessandro Achille, Onkar Dabeer, Bernt Schiele, Ashwin Swaminathan, Stefano Soatto |
CVPR | 10 |
| 2023 | SAFE: Machine Unlearning With Shard GraphsabstractWe present Synergy Aware Forgetting Ensemble (SAFE), a method to adapt large models on a diverse collection of data while minimizing the expected cost to remove the influence of training samples from the trained model. This process, also known as selective forgetting or unlearning, is often conducted by partitioning a dataset into shards, training fully independent models on each, then ensembling the resulting models. Increasing the number of shards reduces the expected cost to forget but at the same time it increases inference cost and reduces the final accuracy of the model since synergistic information between samples is lost during the independent model training. Rather than treating each shard as independent, SAFE introduces the notion of a shard graph, which allows incorporating limited information from other shards during training, trading off a modest increase in expected forgetting cost with a significant increase in accuracy, all while still attaining complete removal of residual influence after forgetting. SAFE uses a lightweight system of adapters which can be trained while reusing most of the computations. This allows SAFE to be trained on shards an order-of-magnitude smaller than current state-of-the-art methods (thus reducing the forgetting costs) while also maintaining high accuracy, as we demonstrate empirically on fine-grained computer vision datasets. Yonatan Dukler, Benjamin Bowman, Alessandro Achille, Aditya Golatkar, Ashwin Swaminathan, Stefano Soatto |
ICCV | 5 |
| 2023 | Your representations are in the network: composable and parallel adaptation for large scale modelsabstractWe present a framework for transfer learning that efficiently adapts a large base-model by learning lightweight cross-attention modules attached to its intermediate activations.
We name our approach InCA (Introspective-Cross-Attention) and show that it can efficiently survey a network’s representations and identify strong performing adapter models for a downstream task.
During training, InCA enables training numerous adapters efficiently and in parallel, isolated from the frozen base model. On the ViT-L/16 architecture, our experiments show that a single adapter, 1.3% of the full model, is able to reach full fine-tuning accuracy on average across 11 challenging downstream classification tasks.
Compared with other forms of parameter-efficient adaptation, the isolated nature of the InCA adaptation is computationally desirable for large-scale models. For instance, we adapt ViT-G/14 (1.8B+ parameters) quickly with 20+ adapters in parallel on a single V100 GPU (76% GPU memory reduction) and exhaustively identify its most useful representations.
We further demonstrate how the adapters learned by InCA can be incrementally modified or combined for flexible learning scenarios and our approach achieves state of the art performance on the ImageNet-to-Sketch multi-task benchmark. Yonatan Dukler, Alessandro Achille, Hao Yang 0043, Varsha Vivek, Luca Zancato, Benjamin Bowman, Avinash Ravichandran, Charless C. Fowlkes, Ashwin Swaminathan, Stefano Soatto |
NeurIPS | 9 |
| 2011 | Information-theoretic database building and querying for mobile Augmented Reality applicationsabstractRecently, there has been tremendous interest in the area of mobile Augmented Reality (AR) with applications including navigation, social networking, gaming and education. Current generation mobile phones are equipped with camera, GPS and other sensors, e.g., magnetic compass, accelerometer, gyro in addition to having ever increasing computing/graphics capabilities and memory storage. Mobile AR applications process the output of one or more sensors to augment the real world view with useful information. This paper's focus is on the camera sensor output, and describes the building blocks for a vision-based AR system. We present information-theoretic techniques to build and maintain an image (feature) database based on reference images, and for querying the captured input images against this database. Performance results using standard image sets are provided demonstrating superior recognition performance even with dramatic reductions in feature database size. Pawan K. Baheti, Ashwin Swaminathan, Murali Chari, Serafin Diaz, Slawek Grzechnik |
ISMAR | 2 |
| 2010 | Relating Reputation and Money in Online MarketsabstractReputation in online economic systems is typically quantified using counters that specify positive and negative feedback from past transactions and/or some form of transaction network analysis that aims to quantify the likelihood that a network user will commit a fraudulent transaction. These approaches can be deceiving to honest users from numerous perspectives. We take a radically different approach with the goal of guaranteeing to a buyer that a fraudulent seller cannot disappear from the system with profit following a set of fabricated transactions that total a certain monetary limit. Even in the case of stolen identity, such an adversary cannot produce illegal profit unless a buyer decides to pay over the suggested limit. Ashwin Swaminathan, Renan G. Cattelan, Ydo Wexler, Cherian V. Mathew, Darko Kirovski |
ACM Trans. Web | 1 |
| 2009 | Tampering identification using Empirical Frequency ResponseabstractWith the widespread popularity of digital images and the presence of easy-to-use image editing software, content integrity can no longer be taken for granted, and there is a strong need for techniques that not only detect the presence of tampering but also identify its type. This paper focusses on tampering-type identification and introduces a new approach based on the empirical frequency response (EFR) to address this problem. We show that several types of tampering operations, both linear shift invariant (LSI) and non-LSI, can be characterized consistently and distinctly by their EFRs. We then extend the approach to estimate the EFR for scenarios where only the final image is available. Theoretical reasoning supported by experimental results verify the effectiveness of this method for identifying the type of a tampering operation. Wei-Hong Chuang, Ashwin Swaminathan, Min Wu 0001 |
ICASSP | 2 |
| 2009 | Secure image retrieval through feature protectionabstractThis paper addresses the problem of image retrieval from an encrypted database, where data confidentiality is preserved both in the storage and retrieval process. The paper focuses on image feature protection techniques which enable similarity comparison among protected features. By utilizing both signal processing and cryptographic techniques, three schemes are investigated and compared, including bit-plane randomization, random projection, and randomized unary encoding. Experimental results show that secure image retrieval can achieve comparable retrieval performance to conventional image retrieval techniques without revealing information about image content. This work enriches the area of secure information retrieval and can find applications in secure online services for images and videos. Wenjun Lu, Avinash L. Varna, Ashwin Swaminathan, Min Wu 0001 |
ICASSP | 3 |
| 2009 | Relating Reputation and Money in On-line MarketsabstractReputation in on-line economic systems is typically quantified using counters that specify positive and negative feedback from past transactions and/or some form of transaction network analysis that aims to quantify the likelihood that a network user will commit a fraudulent transaction. These approaches can be deceiving to honest users from numerous perspectives. We take a radically different approach with a goal to guarantee to a buyer that a seller cannot disappear from the system with profit following a set of transactions that total a certain monetary limit. Even in the case of stolen identity, an adversary cannot produce illegal profit unless a buyer decides to pay over the suggested sales limit. Ashwin Swaminathan, Renan G. Cattelan, Cherian V. Mathew, Ydo Wexler, Darko Kirovski |
Web Intelligence | 1 |
| 2009 | Essential PagesabstractResults to Web search queries are ranked using heuristics that typically analyze the global link topology, user behavior, and content relevance. We point to a particular inefficiency of such methods: information redundancy. In queries where learning about a subject is an objective, modern search engines return relatively unsatisfactory results as they consider the query coverage by each page individually, not a set of pages as a whole. We address this problem using essential pages. If we denote as $\mathbb{S}_Q$ the total knowledge that exists on the Web about a given query $Q$, we want to build a search engine that returns a set of essential pages $E_Q$ that maximizes the information covered over $\mathbb{S}_Q$. We present a preliminary prototype that optimizes the selection of essential pages; we draw some informal comparisons with respect to existing search engines; and finally, we evaluate our prototype using a blind-test user study. Ashwin Swaminathan, Cherian V. Mathew, Darko Kirovski |
Web Intelligence | 1 |
| 2009 | Intrinsic sensor noise features for forensic analysis on scanners and scanned imagesabstractA large portion of digital images available today are acquired using digital cameras or scanners. While cameras provide digital reproduction of natural scenes, scanners are often used to capture hard-copy art in a more controlled environment. In this paper, new techniques for nonintrusive scanner forensics that utilize intrinsic sensor noise features are proposed to verify the source and integrity of digital scanned images. Scanning noise is analyzed from several aspects using only scanned image samples, including through image denoising, wavelet analysis, and neighborhood prediction, and then obtain statistical features from each characterization. Based on the proposed statistical features of scanning noise, a robust scanner identifier is constructed to determine the model/brand of the scanner used to capture a scanned image. Utilizing these noise features, we extend the scope of acquisition forensics to differentiating scanned images from camera-taken photographs and computer-generated graphics. The proposed noise features also enable tampering forensics to detect postprocessing operations on scanned images. Experimental results are presented to demonstrate the effectiveness of employing the proposed noise features for performing various forensic analysis on scanners and scanned images. Hongmei Gou, Ashwin Swaminathan, Min Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2009 | Fingerprinting compressed multimedia signalsabstractDigital fingerprinting is a technique to deter unauthorized redistribution of multimedia content by embedding a unique identifying signal in each legally distributed copy. The embedded fingerprint can later be extracted and used to trace the originator of an unauthorized copy. A group of users may collude and attempt to create a version of the content that cannot be traced back to any of them. As multimedia data is commonly stored in compressed form, this paper addresses the problem of fingerprinting compressed signals. Analysis is carried out to show that due to the quantized nature of the host signal and the embedded fingerprint, directly extending traditional fingerprinting techniques for uncompressed signals to the compressed case leads to low collusion resistance. To overcome this problem and improve the collusion resistance, a new technique for fingerprinting compressed signals called Anti-Collusion Dither (ACD) is proposed, whereby a random dither signal is added to the compressed host before embedding so as to make the effective host signal appear more continuous. The proposed technique is shown to reduce the accuracy with which attackers can estimate the host signal, and from an information theoretic perspective, the proposed ACD technique increases the maximum number of users that can be supported by the fingerprinting system under a given attack. Both analytical and experimental studies confirm that the proposed technique increases the probability of identifying a guilty user and can approximately quadruple the collusion resistance compared to conventional Gaussian fingerprinting. Avinash L. Varna, Shan He 0002, Ashwin Swaminathan, Min Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2008 | A decision theoretic framework for analyzing binary hash-based content identification systemsabstractContent identification has many applications, ranging from preventing illegal sharing of copyrighted content on video sharing websites, to automatic identification and tagging of content. Several content identification techniques based on watermarking or robust hashes have been proposed in the literature, but they have mostly been evaluated through experiments. This paper analyzes binary hash-based content identification schemes under a decision theoretic framework and presents a lower bound on the length of the hash required to correctly identify multimedia content that may have undergone modifications. A practical scheme for content identification is evaluated under the proposed framework. The results obtained through experiments agree very well with the performance suggested by the theoretical analysis. Avinash L. Varna, Ashwin Swaminathan, Min Wu 0001 |
Digital Rights Management Workshop | 2 |
| 2008 | Image acquisition forensics: Forensic analysis to identify imaging sourceabstractWith widespread availability of digital images and easy-to-use image editing softwares, the origin and integrity of digital images has become a serious concern. This paper introduces the problem of image acquisition forensics and proposes a fusion of a set of signal processing features to identify the source of digital images. Our results show that the devices' color interpolation coefficients and noise statistics can jointly serve as good forensic features to help accurately trace the origin of the input image to its production process and to differentiate between images produced by cameras, cell phone cameras, scanners, and computer graphics. Further, the proposed features can also be extended to determining the brand and model of the device. Thus, the techniques introduced in this work provide a unified framework for image acquisition forensics. Christine McKay, Ashwin Swaminathan, Hongmei Gou, Min Wu 0001 |
ICASSP | 2 |
| 2008 | A pattern classification framework for theoretical analysis of component forensicsabstractComponent forensics is an emerging methodology for forensic analysis that aims at estimating the algorithms and parameters in each component of a digital device. This paper proposes a theoretical foundation to examine the performance limits of component forensics. Using ideas from pattern classification theory, we define formal notions of identifiability of components in the information processing chain. We show that the parameters of certain device components can be accurately identified only in controlled settings through semi non-intrusive forensics, while the parameters of some others can be computed directly from the available sample data via complete non-intrusive analysis. We then extend the proposed theoretical framework to quantify and improve the accuracies and confidence in component parameter identification for several forensic applications. Ashwin Swaminathan, Min Wu 0001, K. J. Ray Liu |
ICASSP | 1 |
| 2008 | Digital Image Forensics via Intrinsic FingerprintsabstractDigital imaging has experienced tremendous growth in recent decades, and digital camera images have been used in a growing number of applications. With such increasing popularity and the availability of low-cost image editing software, the integrity of digital image content can no longer be taken for granted. This paper introduces a new methodology for the forensic analysis of digital camera images. The proposed method is based on the observation that many processing operations, both inside and outside acquisition devices, leave distinct intrinsic traces on digital images, and these intrinsic fingerprints can be identified and employed to verify the integrity of digital data. The intrinsic fingerprints of the various in-camera processing operations can be estimated through a detailed imaging model and its component analysis. Further processing applied to the camera captured image is modelled as a manipulation filter, for which a blind deconvolution technique is applied to obtain a linear time-invariant approximation and to estimate the intrinsic fingerprints associated with these postcamera operations. The absence of camera-imposed fingerprints from a test image indicates that the test image is not a camera output and is possibly generated by other image production processes. Any change or inconsistencies among the estimated camera-imposed fingerprints, or the presence of new types of fingerprints suggest that the image has undergone some kind of processing after the initial capture, such as tampering or steganographic embedding. Through analysis and extensive experimental studies, this paper demonstrates the effectiveness of the proposed framework for nonintrusive digital image forensics. Ashwin Swaminathan, Min Wu 0001, K. J. Ray Liu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2007 | Optimization of Input Pattern for Semi Non-Intrusive Component Forensics of Digital CamerasabstractThis paper considers the problem of semi non-intrusive component forensics and proposes a methodology to identify the algorithms and parameters employed by various processing modules inside a digital camera. The proposed analysis techniques assume the availability of the camera; and introduce a forensic methodology to estimate the parameters of the color interpolation and white balancing algorithms employed in cameras. We devise testing conditions, and design good input patterns to improve the overall accuracy in parameter estimation. As demonstrated by the results in the paper, the proposed techniques provide a much lower estimation bias and variance compared to non-intrusive analysis. The features obtained from component forensic analysis provide useful evidence for such applications as analyzing technology evolution trend, detecting technology infringement/licensing, protecting intellectual property rights, and determining camera source. Ashwin Swaminathan, Min Wu 0001, K. J. Ray Liu |
ICASSP (2) | 1 |
| 2007 | Collusion-Resistant Fingerprinting for Compressed Multimedia SignalsabstractMost existing collusion-resistant fingerprinting techniques are for fingerprinting uncompressed signals. In this paper, we first study the performance of the traditional Gaussian based spread spectrum sequences for fingerprinting compressed signals and show that the system can be easily defeated by averaging or taking the median of a few copies. To overcome the collusion problem for compressed multimedia host signals, we propose a technique called anti-collusion dithering to mimic an uncompressed signal. Results show higher probability of catching a colluder using the proposed scheme compared to using Gaussian based fingerprints. Avinash L. Varna, Shan He 0002, Ashwin Swaminathan, Min Wu 0001, Haiming Lu, Zengxiang Lu |
ICASSP (2) | 3 |
| 2007 | Noise Features for Image Tampering Detection and SteganalysisabstractWith increasing availability of low-cost image editing softwares, the authenticity of digital images can no longer be taken for granted. Digital images have also been used as cover data for transmitting secret information in the field of steganography. In this paper, we introduce a new set of features for multimedia forensics to determine if a digital image is an authentic camera output or if it has been tampered or embedded with hidden data. We perform such image forensic analysis employing three sets of statistical noise features, including those from denoising operations, wavelet analysis, and neighborhood prediction. Our experimental results demonstrate that the proposed method can effectively distinguish digital images from their tampered or stego versions. Hongmei Gou, Ashwin Swaminathan, Min Wu 0001 |
ICIP (6) | 2 |
| 2007 | Analysis of Nonlinear Collusion Attacks on Fingerprinting Systems for Compressed MultimediaabstractIn this paper, we analyze the effect of various collusion attacks on fingerprinting systems for compressed multimedia. We evaluate the effectiveness of the collusion attacks in terms of the probability of detection and accuracy in estimating the host signal. Our analysis shows that applying averaging collusion on copies of moderately compressed content gives a highly accurate estimation of the host, and can effectively remove the embedded fingerprints. Averaging is thus the best choice for an attacker as the probability of detection and the distortion introduced are the lowest. Avinash L. Varna, Shan He 0002, Ashwin Swaminathan, Min Wu 0001 |
ICIP (2) | 3 |
| 2007 | A Component Estimation Framework for Information ForensicsabstractWith a rapid growth of imaging technologies and an increasingly widespread usage of digital images and videos for a large number of high security and forensic applications, there is a strong need for techniques to verify the source and integrity of digital data. Component forensics is new approach for forensic analysis that aims to estimate the algorithms and parameters in each component of the digital device. In this paper, we develop a novel theoretical foundation to understand the fundamental performance limits of component forensics. We define formal notions of identifiability of components in the information processing chain, and present methods to quantify the accuracies at which the component parameters can be estimated. Building upon the proposed theoretical framework, we devise methods to improve the accuracies of component parameter estimation for a wide range of forensic applications. Ashwin Swaminathan, Min Wu 0001, K. J. Ray Liu |
MMSP | 1 |
| 2007 | Nonintrusive Component Forensics of Visual Sensors Using Output ImagesabstractRapid technology development and the widespread use of visual sensors have led to a number of new problems related to protecting intellectual property rights, handling patent infringements, authenticating acquisition sources, and identifying content manipulations. This paper introduces nonintrusive component forensics as a new methodology for the forensic analysis of visual sensing information, aiming to identify the algorithms and parameters employed inside various processing modules of a digital device by only using the device output data without breaking the device apart. We propose techniques to estimate the algorithms and parameters employed by important camera components, such as color filter array and color interpolation modules. The estimated interpolation coefficients provide useful features to construct an efficient camera identifier to determine the brand and model from which an image was captured. The results obtained from such component analysis are also useful to examine the similarities between the technologies employed by different camera models to identify potential infringement/licensing and to facilitate studies on technology evolution Ashwin Swaminathan, Min Wu 0001, K. J. Ray Liu |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2006 | Non-Intrusive Forensic Analysis of Visual Sensors Using Output ImagesabstractThis paper considers the problem of non-intrusive forensic analysis of the individual components in visual sensors and its implementation. As a new addition to the emerging area of forensic engineering, we present a framework for analyzing technologies employed inside digital cameras based on output images, and develop a set of forensic signal processing algorithms for visual sensors based on color array sensor and interpolation methods. We show through simulations that the proposed method is robust against compression and noise, and can help identify various processing components inside the camera. Such a non-intrusive forensic framework would provide useful evidence for analyzing technology infringement and evolution for visual sensors. Ashwin Swaminathan, Min Wu 0001, K. J. Ray Liu |
ICASSP (5) | 1 |
| 2006 | Image Tampering Identification using Blind DeconvolutionabstractDigital images have been used in growing number of applications from law enforcement and surveillance, to medical diagnosis and consumer photography. With such widespread popularity and the presence of low-cost image editing softwares, the integrity of image content can no longer be taken for granted. In this paper, we propose a novel technique based on blind deconvolution to verify image authenticity. We consider the direct output images of a camera as authentic, and introduce algorithms to detect further processing such as tampering applied to the image. Our proposed method is based on the observation that many tampering operations can be approximated as a combination of linear and non-linear components. We model the linear part of the tampering process as a filter, and obtain its coefficients using blind deconvolution. These estimated coefficients are then used to identify possible manipulations. We demonstrate the effectiveness of the proposed image authentication technique and compare our results with existing works. Ashwin Swaminathan, Min Wu 0001, K. J. Ray Liu |
ICIP | 1 |
| 2006 | Robust and secure image hashingabstractImage hash functions find extensive applications in content authentication, database search, and watermarking. This paper develops a novel algorithm for generating an image hash based on Fourier transform features and controlled randomization. We formulate the robustness of image hashing as a hypothesis testing problem and evaluate the performance under various image processing operations. We show that the proposed hash function is resilient to content-preserving modifications, such as moderate geometric and filtering distortions. We introduce a general framework to study and evaluate the security of image hashing systems. Under this new framework, we model the hash values as random variables and quantify its uncertainty in terms of differential entropy. Using this security framework, we analyze the security of the proposed schemes and several existing representative methods for image hashing. We then examine the security versus robustness tradeoff and show that the proposed hashing methods can provide excellent security and robustness. Ashwin Swaminathan, Yinian Mao, Min Wu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2005 | Security of feature extraction in image hashingabstractSecurity and robustness are two important requirements for image hash functions. We introduce "differential entropy" as a metric to quantify the amount of randomness in image hash functions and to study their security. We present a mathematical framework and derive expressions for the proposed security metric for various common image hashing schemes. Using the proposed security metric, we discuss the trade-offs between security and robustness in image hashing. Ashwin Swaminathan, Yinian Mao, Min Wu 0001 |
ICASSP (2) | 1 |
| 2004 | Image hashing resilient to geometric and filtering operationsabstractImage hash functions provide compact representations of images, which is useful for search and authentication applications. In this work, we have identified a general three step framework and proposed a new image hashing scheme that achieves a better overall performance than the existing approaches under various kinds of image processing distortions. By exploiting the properties of discrete polar Fourier transform and incorporating cryptographic keys, the proposed image hash is resilient to geometric and filtering operations, and is secure against guessing and forgery attacks. Ashwin Swaminathan, Yinian Mao, Min Wu 0001 |
MMSP | 1 |