EDBT 2026 Demo / reviewers in the wild / expert
Vinu Sankar Sadasivan
dblp:244/8052
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
4 papers |
Security and privacy of machine learning · 66% Digital forensics and information hiding · 34% | |
| Artificial intelligence
7 papers |
Trustworthy machine learning · 40% Language models and text generation · 23% Deep learning architectures and training · 12% |
Topics — the 27 heaviest of 29, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Security and privacy of machine learning
adversarial attack |
2.4 | 3 | 2025 | Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text · NeurIPS 2025 Fast Adversarial Attacks on Language Models In One GPU Minute · ICML 2024 Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks · ICLR 2024 |
Security and privacy of machine learning › adversarial attack
evasion attack |
0.9 | 1 | 2025 | Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text · NeurIPS 2025 |
Digital forensics and information hiding › synthetic media detection
machine-generated text detection |
0.9 | 1 | 2025 | Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text · NeurIPS 2025 |
Security and privacy of machine learning › adversarial attack › textual adversarial attack
paraphrase attack |
0.9 | 1 | 2025 | Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
hallucination |
0.8 | 1 | 2024 | Fast Adversarial Attacks on Language Models In One GPU Minute · ICML 2024 |
Natural language and speech › Language models and text generation
hallucination detection |
0.8 | 1 | 2024 | LLM-Check: Investigating Detection of Hallucinations in Large Language Models · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › representation analysis
hidden state analysis |
0.8 | 1 | 2024 | LLM-Check: Investigating Detection of Hallucinations in Large Language Models · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | LLM-Check: Investigating Detection of Hallucinations in Large Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | Fast Adversarial Attacks on Language Models In One GPU Minute · ICML 2024 |
Digital forensics and information hiding › digital forensics › multimedia forensics › image forensics
AI-generated image detection |
0.8 | 1 | 2024 | Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks · ICLR 2024 |
Security and privacy of machine learning › adversarial attack
jailbreak attack |
0.8 | 1 | 2024 | Fast Adversarial Attacks on Language Models In One GPU Minute · ICML 2024 |
Digital forensics and information hiding
watermarking |
0.8 | 1 | 2024 | Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks · ICLR 2024 |
Digital forensics and information hiding › watermarking › watermarking security › watermark attack
watermark removal |
0.8 | 1 | 2024 | Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks · ICLR 2024 |
Machine learning › Trustworthy machine learning › robustness
adversarial examples |
0.7 | 1 | 2023 | Exploring Geometry of Blind Spots in Vision models · NeurIPS 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.7 | 1 | 2023 | Exploring Geometry of Blind Spots in Vision models · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › robustness › data poisoning
unlearnable examples |
0.7 | 1 | 2023 | CUDA: Convolution-Based Unlearnable Datasets · CVPR 2023 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.7 | 1 | 2023 | Exploring Geometry of Blind Spots in Vision models · NeurIPS 2023 |
Security and privacy of machine learning
poisoning attack |
0.7 | 1 | 2023 | CUDA: Convolution-Based Unlearnable Datasets · CVPR 2023 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.4 | 1 | 2019 | Shallow RNN: Accurate Time-series Classification on Resource Constrained Devices · NeurIPS 2019 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.4 | 1 | 2019 | Shallow RNN: Accurate Time-series Classification on Resource Constrained Devices · NeurIPS 2019 |
Machine learning › Efficient and distributed learning › inference efficiency
resource-constrained inference |
0.4 | 1 | 2019 | Shallow RNN: Accurate Time-series Classification on Resource Constrained Devices · NeurIPS 2019 |
Natural language and speech › Language models and text generation › text generation
paraphrase generation |
0.3 | 1 | 2025 | Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
diffusion-based purification |
0.2 | 1 | 2024 | Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks · ICLR 2024 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks · ICLR 2024 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.2 | 1 | 2024 | LLM-Check: Investigating Detection of Hallucinations in Large Language Models · NeurIPS 2024 |
Security and privacy of machine learning
membership inference |
0.2 | 1 | 2024 | Fast Adversarial Attacks on Language Models In One GPU Minute · ICML 2024 |
Security and privacy of machine learning
privacy attack |
0.2 | 1 | 2024 | Fast Adversarial Attacks on Language Models In One GPU Minute · ICML 2024 |
Methods — techniques the papers use, named apart from their topics
large language model · 1.7adversarial paraphrasing · 1.7spoofing attack · 1.5model substitution attack · 1.5gradient-free attack · 1.5diffusion purification attack · 1.5beam search · 1.5adversarial training · 1.3output probability analysis · 0.8attention map analysis · 0.8convolution-based perturbation · 0.7recurrent neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated TextabstractThe increasing capabilities of Large Language Models (LLMs) have raised concerns about their misuse in AI-generated plagiarism and social engineering. While various AI-generated text detectors have been proposed to mitigate these risks, many remain vulnerable to simple evasion techniques such as paraphrasing. However, recent detectors have shown greater robustness against such basic attacks. In this work, we introduce \textbf{Adversarial Paraphrasing}, a training-free attack framework that universally humanizes any AI-generated text to evade detection more effectively. Our approach leverages an off-the-shelf instruction-following LLM to paraphrase AI-generated content under the guidance of an AI text detector, producing adversarial examples that are specifically optimized to bypass detection. Extensive experiments show that our attack is both broadly effective and highly transferable across several detection systems. For instance, compared to simple paraphrasing attack—which, ironically, increases the true positive at 1\% false positive (T@1\%F) by 8.57\% on RADAR and 15.03\% on Fast-DetectGPT—adversarial paraphrasing, guided by OpenAI-RoBERTa-Large, reduces T@1\%F by 64.49\% on RADAR and a striking 98.96\% on Fast-DetectGPT. Across a diverse set of detectors—including neural network-based, watermark-based, and zero-shot approaches—our attack achieves an average T@1\%F reduction of 87.88\% under the guidance of OpenAI-RoBERTa-Large.
We also analyze the tradeoff between text quality and our attack success to find that our method can significantly reduce detection rates, with mostly a slight degradation in text quality.
Our novel adversarial setup highlights the need for more robust and resilient detection strategies in the light of increasingly sophisticated evasion techniques. Yize Cheng, Vinu Sankar Sadasivan, Mehrdad Saberi, Shoumik Saha, Soheil Feizi |
NeurIPS | 2 |
| 2024 | Robustness of AI-Image Detectors: Fundamental Limits and Practical AttacksabstractIn light of recent advancements in generative AI models, it has become essential to distinguish genuine content from AI-generated one to prevent the malicious usage of fake materials as authentic ones and vice versa. Various techniques have been introduced for identifying AI-generated images, with watermarking emerging as a promising approach. In this paper, we analyze the robustness of various AI-image detectors including watermarking and classifier-based deepfake detectors. For watermarking methods that introduce subtle image perturbations (i.e., low perturbation budget methods), we reveal a fundamental trade-off between the evasion error rate (i.e., the fraction of watermarked images detected as non-watermarked ones) and the spoofing error rate (i.e., the fraction of non-watermarked images detected as watermarked ones) upon an application of a diffusion purification attack. In this regime, we also empirically show that diffusion purification effectively removes watermarks with minimal changes to images. For high perturbation watermarking methods where notable changes are applied to images, the diffusion purification attack is not effective. In this case, we develop a model substitution adversarial attack that can successfully remove watermarks. Moreover, we show that watermarking methods are vulnerable to spoofing attacks where the attacker aims to have real images (potentially obscene) identified as watermarked ones, damaging the reputation of the developers. In particular, by just having black-box access to the watermarking method, we show that one can generate a watermarked noise image which can be added to the real images to have them falsely flagged as watermarked ones. Finally, we extend our theory to characterize a fundamental trade-off between the robustness and reliability of classifier-based deep fake detectors and demonstrate it through experiments. Code is available at https://github.com/mehrdadsaberi/watermark_robustness. Mehrdad Saberi, Vinu Sankar Sadasivan, Keivan Rezaei, Aounon Kumar, Atoosa Malemir Chegini, Wenxiao Wang 0002, Soheil Feizi |
ICLR | 2 |
| 2024 | Fast Adversarial Attacks on Language Models In One GPU MinuteabstractIn this paper, we introduce a novel class of fast, beam search-based adversarial attack (BEAST) for Language Models (LMs). BEAST employs interpretable parameters, enabling attackers to balance between attack speed, success rate, and the readability of adversarial prompts. The computational efficiency of BEAST facilitates us to investigate its applications on LMs for jailbreaking, eliciting hallucinations, and privacy attacks. Our gradient-free targeted attack can jailbreak aligned LMs with high attack success rates within one minute. For instance, BEAST can jailbreak Vicuna-7B-v1.5 under one minute with a success rate of 89% when compared to a gradient-based baseline that takes over an hour to achieve 70% success rate using a single Nvidia RTX A6000 48GB GPU. BEAST can also generate adversarial suffixes for successful jailbreaks that can transfer to unseen prompts and unseen models such as GPT-4-Turbo. Additionally, we discover a unique outcome wherein our untargeted attack induces hallucinations in LM chatbots. Through human evaluations, we find that our untargeted attack causes Vicuna-7B-v1.5 to produce $\sim$15% more incorrect outputs when compared to LM outputs in the absence of our attack. We also learn that 22% of the time, BEAST causes Vicuna to generate outputs that are not relevant to the original prompt. Further, we use BEAST to generate adversarial prompts in a few seconds that can boost the performance of existing membership inference attacks for LMs. We believe that our fast attack, BEAST, has the potential to accelerate research in LM security and privacy. Vinu Sankar Sadasivan, Shoumik Saha, Gaurang Sriramanan, Priyatham Kattakinda, Atoosa Malemir Chegini, Soheil Feizi |
ICML | 1 |
| 2024 | LLM-Check: Investigating Detection of Hallucinations in Large Language ModelsabstractWhile Large Language Models (LLMs) have become immensely popular due to their outstanding performance on a broad range of tasks, these models are prone to producing hallucinations— outputs that are fallacious or fabricated yet often appear plausible or tenable at a glance. In this paper, we conduct a comprehensive investigation into the nature of hallucinations within LLMs and furthermore explore effective techniques for detecting such inaccuracies in various real-world settings. Prior approaches to detect hallucinations in LLM outputs, such as consistency checks or retrieval-based methods, typically assume access to multiple model responses or large databases. These techniques, however, tend to be computationally expensive in practice, thereby limiting their applicability to real-time analysis. In contrast, in this work, we seek to identify hallucinations within a single response in both white-box and black-box settings by analyzing the internal hidden states, attention maps, and output prediction probabilities of an auxiliary LLM. In addition, we also study hallucination detection in scenarios where ground-truth references are also available, such as in the setting of Retrieval-Augmented Generation (RAG). We demonstrate that the proposed detection methods are extremely compute-efficient, with speedups of up to 45x and 450x over other baselines, while achieving significant improvements in detection performance over diverse datasets. Gaurang Sriramanan, Siddhant Bharti, Vinu Sankar Sadasivan, Shoumik Saha, Priyatham Kattakinda, Soheil Feizi |
NeurIPS | 3 |
| 2023 | CUDA: Convolution-Based Unlearnable DatasetsabstractLarge-scale training of modern deep learning models heavily relies on publicly available data on the web. This potentially unauthorized usage of online data leads to concerns regarding data privacy. Recent works aim to make unlearnable data for deep learning models by adding small, specially designed noises to tackle this issue. However, these methods are vulnerable to adversarial training (AT) and/or are computationally heavy. In this work, we propose a novel, model-free, Convolution-based Unlearnable DAtaset (CUDA) generation technique. CUDA is generated using controlled class-wise convolutions with filters that are randomly generated via a private key. CUDA encourages the network to learn the relation between filters and labels rather than informative features for classifying the clean data. We develop some theoretical analysis demonstrating that CUDA can successfully poison Gaussian mixture data by reducing the clean data performance of the optimal Bayes classifier. We also empirically demonstrate the effectiveness of CUDA with various datasets (CIFAR-10, CIFAR-100, ImageNet-100, and Tiny-ImageNet), and architectures (ResNet-18, VGG-16, Wide ResNet-34-10, DenseNet-121, DeIT, EfficientNetV2-S, and MobileNetV2). Our experiments show that CUDA is robust to various data augmentations and training approaches such as smoothing, AT with different budgets, transfer learning, and fine-tuning. For instance, training a ResNet-18 on ImageNet-100 CUDA achieves only 8.96%, 40.08%, and 20.58% clean test accuracies with empirical risk minimization (ERM), L∞AT, and L2 AT, respectively. Here, ERM on the clean training data achieves a clean test accuracy of 80.66%. CUDA exhibits unlearnability effect with ERM even when only a fraction of the training dataset is perturbed. Furthermore, we also show that CUDA is robust to adaptive defenses designed specifically to break it. Vinu Sankar Sadasivan, Mahdi Soltanolkotabi, Soheil Feizi |
CVPR | 1 |
| 2023 | Exploring Geometry of Blind Spots in Vision modelsabstractDespite the remarkable success of deep neural networks in a myriad of settings, several works have demonstrated their overwhelming sensitivity to near-imperceptible perturbations, known as adversarial attacks. On the other hand, prior works have also observed that deep networks can be under-sensitive, wherein large-magnitude perturbations in input space do not induce appreciable changes to network activations. In this work, we study in detail the phenomenon of under-sensitivity in vision models such as CNNs and Transformers, and present techniques to study the geometry and extent of “equi-confidence” level sets of such networks. We propose a Level Set Traversal algorithm that iteratively explores regions of high confidence with respect to the input space using orthogonal components of the local gradients. Given a source image, we use this algorithm to identify inputs that lie in the same equi-confidence level set as the source image despite being perceptually similar to arbitrary images from other classes. We further observe that the source image is linearly connected by a high-confidence path to these inputs, uncovering a star-like structure for level sets of deep networks. Furthermore, we attempt to identify and estimate the extent of these connected higher-dimensional regions over which the model maintains a high degree of confidence. Sriram Balasubramanian, Gaurang Sriramanan, Vinu Sankar Sadasivan, Soheil Feizi |
NeurIPS | 3 |
| 2019 | Shallow RNN: Accurate Time-series Classification on Resource Constrained DevicesabstractRecurrent Neural Networks (RNNs) capture long dependencies and context, and 2 hence are the key component of typical sequential data based tasks. However, the sequential nature of RNNs dictates a large inference cost for long sequences even if the hardware supports parallelization. To induce long-term dependencies, and yet admit parallelization, we introduce novel shallow RNNs. In this architecture, the first layer splits the input sequence and runs several independent RNNs. The second layer consumes the output of the first layer using a second RNN thus capturing long dependencies. We provide theoretical justification for our architecture under weak assumptions that we verify on real-world benchmarks. Furthermore, we show that for time-series classification, our technique leads to substantially improved inference time over standard RNNs without compromising accuracy. For example, we can deploy audio-keyword classification on tiny Cortex M4 devices (100MHz processor, 256KB RAM, no DSP available) which was not possible using standard RNN models. Similarly, using SRNN in the popular Listen-Attend-Spell (LAS) architecture for phoneme classification [4], we can reduce the lag inphoneme classification by 10-12x while maintaining state-of-the-art accuracy. Don Kurian Dennis, Durmus Alp Emre Acar, Vikram Mandikal, Vinu Sankar Sadasivan, Venkatesh Saligrama, Harsha Vardhan Simhadri, Prateek Jain 0002 |
NeurIPS | 4 |