EDBT 2026 Demo / reviewers in the wild / expert
Jiameng Pu
dblp:198/9527
· DBLP profile ↗
7ranked-venue papers
4as first author
3since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 5 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Deepfake Text Detection: Limitations and OpportunitiesabstractRecent advances in generative models for language have enabled the creation of convincing synthetic text or deepfake text. Prior work has demonstrated the potential for misuse of deepfake text to mislead content consumers. Therefore, deepfake text detection, the task of discriminating between human and machine-generated text, is becoming increasingly critical. Several defenses have been proposed for deepfake text detection. However, we lack a thorough understanding of their real-world applicability. In this paper, we collect deepfake text from 4 online services powered by Transformer-based tools to evaluate the generalization ability of the defenses on content in the wild. We develop several low-cost adversarial attacks, and investigate the robustness of existing defenses against an adaptive attacker. We find that many defenses show significant degradation in performance under our evaluation scenarios compared to their original claimed performance. Our evaluation shows that tapping into the semantic information in the text content is a promising approach for improving the robustness and generalization performance of deepfake text detection schemes. Jiameng Pu, Zain Sarwar, Sifat Muhammad Abdullah, Abdullah Rehman, Yoonjin Kim, Parantapa Bhattacharya, Mobin Javed, Bimal Viswanath |
SP | 1 |
| 2021 | T-Miner: A Generative Approach to Defend Against Trojan Attacks on DNN-based Text Classification
Ahmadreza Azizi, Ibrahim Asadullah Tahmid, Asim Waheed, Neal Mangaokar, Jiameng Pu, Mobin Javed, Chandan K. Reddy, Bimal Viswanath |
USENIX Security Symposium | 5 |
| 2021 | Deepfake Videos in the Wild: Analysis and DetectionabstractAI-manipulated videos, commonly known as deepfakes, are an emerging problem. Recently, researchers in academia and industry have contributed several (self-created) benchmark deepfake datasets, and deepfake detection algorithms. However, little effort has gone towards understanding deepfake videos in the wild, leading to a limited understanding of the real-world applicability of research contributions in this space. Even if detection schemes are shown to perform well on existing datasets, it is unclear how well the methods generalize to real-world deepfakes. To bridge this gap in knowledge, we make the following contributions: First, we collect and present the largest dataset of deepfake videos in the wild, containing 1,869 videos from YouTube and Bilibili, and extract over 4.8M frames of content. Second, we present a comprehensive analysis of the growth patterns, popularity, creators, manipulation strategies, and production methods of deepfake content in the real-world. Third, we systematically evaluate existing defenses using our new dataset, and observe that they are not ready for deployment in the real-world. Fourth, we explore the potential for transfer learning schemes and competition-winning techniques to improve defenses. Jiameng Pu, Neal Mangaokar, Lauren Kelly, Parantapa Bhattacharya, Kavya Sundaram, Mobin Javed, Bolun Wang, Bimal Viswanath |
WWW | 1 |
| 2020 | NoiseScope: Detecting Deepfake Images in a Blind SettingabstractRecent advances in Generative Adversarial Networks (GANs) have significantly improved the quality of synthetic images or deepfakes. Photorealistic images generated by GANs start to challenge the boundary of human perception of reality, and brings new threats to many critical domains, e.g., journalism, and online media. Detecting whether an image is generated by GAN or a real camera has become an important yet under-investigated area. In this work, we propose a blind detection approach called NoiseScope for discovering GAN images among other real images. A blind approach requires no a priori access to GAN images for training, and demonstrably generalizes better than supervised detection schemes. Our key insight is that, similar to images from cameras, GAN images also carry unique patterns in the noise space. We extract such patterns in an unsupervised manner to identify GAN images. We evaluate NoiseScope on 11 diverse datasets containing GAN images, and achieve up to 99.68% F1 score in detecting GAN images. We test the limitations of NoiseScope against a variety of countermeasures, observing that NoiseScope holds robust or is easily adaptable. Jiameng Pu, Neal Mangaokar, Bolun Wang, Chandan K. Reddy, Bimal Viswanath |
ACSAC | 1 |
| 2020 | Jekyll: Attacking Medical Image Diagnostics using Deep Generative ModelsabstractAdvances in deep neural networks (DNNs) have shown tremendous promise in the medical domain. However, the deep learning tools that are helping the domain, can also be used against it. Given the prevalence of fraud in the healthcare domain, it is important to consider the adversarial use of DNNs in manipulating sensitive data that is crucial to patient healthcare. In this work, we present the design and implementation of a DNN-based image translation attack on biomedical imagery. More specifically, we propose Jekyll, a neural style transfer framework that takes as input a biomedical image of a patient and translates it to a new image that indicates an attacker-chosen disease condition. The potential for fraudulent claims based on such generated ‘fake’ medical images is significant, and we demonstrate successful attacks on both X-rays and retinal fundus image modalities. We show that these attacks manage to mislead both medical professionals and algorithmic detection schemes. Lastly, we also investigate defensive measures based on machine learning to detect images generated by Jekyll. Neal Mangaokar, Jiameng Pu, Parantapa Bhattacharya, Chandan K. Reddy, Bimal Viswanath |
EuroS&P | 2 |
| 2020 | Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data AugmentationabstractMachine learning has been widely applied to building security applications. However, many machine learning models require the continuous supply of representative labeled data for training, which limits the models' usefulness in practice. In this paper, we use bot detection as an example to explore the use of data synthesis to address this problem. We collected the network traffic from 3 online services in three different months within a year (23 million network requests). We develop a stream-based feature encoding scheme to support machine learning models for detecting advanced bots. The key novelty is that our model detects bots with extremely limited labeled data. We propose a data synthesis method to synthesize unseen (or future) bot behavior distributions. The synthesis method is distribution-aware, using two different generators in a Generative Adversarial Network to synthesize data for the clustered regions and the outlier regions in the feature space. We evaluate this idea and show our method can train a model that outperforms existing methods with only 1% of the labeled data. We show that data synthesis also improves the model's sustainability over time and speeds up the retraining. Finally, we compare data synthesis and adversarial retraining and show they can work complementary with each other to improve the model generalizability. Steve T. K. Jan, Qingying Hao, Tianrui Hu, Jiameng Pu, Sonal Oswal, Gang Wang 0011, Bimal Viswanath |
SP | 4 |
| 2016 | Multiview clustering based on Robust and Regularized Matrix ApproximationabstractPattern recognition tasks such as the data classification and clustering usually can be represented by the perspective of multiple views or feature spaces. Obviously, the accuracy of the classification and clustering should be greatly improved if we carefully consider the discriminabilities from multiple views and explore the complementary information among them. However, multiple features also bring new challenges to handle them. In the literature, many existed multiview feature learning methods dealt with different views equally, thus they couldn't optimally utilize the complementary property of them. On the other hand, the matrix factorization based clustering algorithms usually adopt the conventional ℓ2-norm based squared residue minimization to measure the loss, which is easily influenced by the outliers and noises from the multiple sources of input. In this paper, we propose a novel multiview data clustering algorithm based on the matrix factorization to relieve the above issues. The basic idea of the proposed Robust and Regularized Matrix Approximation (RRMA) is that the observed data matrix could be low-rank approximated by a cluster centroid matrix and a cluster indicator matrix, respectively, and the major contributions of our work lie in the introduction of the robust ℓ2,1-norm and ensemble manifold regularization to regularize the matrix factorization and make the model more discriminative for multiview data clustering. We properly adjust the importance of different views by assigning a set of trainable weights on the views. Moreover, we propose an efficient solution featured with impactful updating rules to seek the local optimal parameters. Encouraging experimental results on numerous public multiview datasets demonstrate the superiority of our model compared to some state-of-the-art methods. Jiameng Pu, Qian Zhang 0009, Lefei Zhang, Bo Du 0001, Jane You |
ICPR | 1 |