Qi Pang

dblp:44/8421 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
18since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 FABLE: Batched Evaluation on Confidential Lookup Tables in 2PC
Zhengyuan Su, Qi Pang, Simon Beyzerov, Wenting Zheng
USENIX Security Symposium2
2025 Iterative Gradient Corrected Semisupervised Seismic Impedance Inversion via Swin Transformer
abstract
Seismic impedance inversion is essential for sub-surface exploration, facilitating precise lithological interpretation by reconstructing subsurface impedance. Although recent deep learning-based methods have advanced this field, many rely on direct mapping from observation to model space, which increases solution uncertainty due to the presence of a large null space, impacting inversion accuracy. To address this issue, we propose an iterative method that operates within the model space, applying progressive gradient correction to incrementally refine the current model towards a physically plausible solution, effectively reducing non-uniqueness and improving inversion robustness compared to single-step updates. The effectiveness of this iterative framework is further strengthened by a semi-supervised learning approach, which critically depends on both the network architecture and the design of the loss function. While most DL methods use convolutional architectures, their localized nature limits the capture of long-range dependencies critical for seismic inversion. To overcome this, we introduce USTNet, a hybrid UNet-Swin Transformer architecture that captures multi-scale features, improving inversion precision. To further ensure consistency with subsurface structure, structural priors are incorporated into the loss function, reinforcing spatial coherence. Experiments on synthetic and field data confirm that the proposed method significantly outperforms conventional and several state-of-the-art deep learning approaches in accuracy.
Qi Pang, Hongling Chen, Jinghuai Gao
IEEE Trans. Geosci. Remote. Sens.1
2024 MPCDiff: Testing and Repairing MPC-Hardened Deep Learning Models
Qi Pang, Yuanyuan Yuan 0001, Shuai Wang 0011
NDSS1
2024 Communication Bounds for the Distributed Experts Problem
abstract
In this work, we study the experts problem in the distributed setting where an expert's cost needs to be aggregated across multiple servers. Our study considers various communication models such as the message-passing model and the broadcast model, along with multiple aggregation functions, such as summing and taking the $\ell_p$ norm of an expert's cost across servers. We propose the first communication-efficient protocols that achieve near-optimal regret in these settings, even against a strong adversary who can choose the inputs adaptively. Additionally, we give a conditional lower bound showing that the communication of our protocols is nearly optimal. Finally, we implement our protocols and demonstrate empirical savings on the HPO-B benchmarks.
Qi Pang, Trung Tran, David P. Woodruff, Zhihao Zhang 0001, Wenting Zheng
NeurIPS2
2024 No Free Lunch in LLM Watermarking: Trade-offs in Watermarking Design Choices
abstract
Advances in generative models have made it possible for AI-generated text, code, and images to mirror human-generated content in many applications. Watermarking, a technique that aims to embed information in the output of a model to verify its source, is useful for mitigating the misuse of such AI-generated content. However, we show that common design choices in LLM watermarking schemes make the resulting systems surprisingly susceptible to attack---leading to fundamental trade-offs in robustness, utility, and usability. To navigate these trade-offs, we rigorously study a set of simple yet effective attacks on common watermarking systems, and propose guidelines and defenses for LLM watermarking in practice.
Qi Pang, Shengyuan Hu 0001, Wenting Zheng, Virginia Smith
NeurIPS1
2024 BOLT: Privacy-Preserving, Accurate and Efficient Inference for Transformers
abstract
The advent of transformers has brought about significant advancements in traditional machine learning tasks. However, their pervasive deployment has raised concerns about the potential leakage of sensitive information during inference. Existing approaches using secure multiparty computation (MPC) face limitations when applied to transformers due to the extensive model size and resource-intensive matrix-matrix multiplications. In this paper, we present BOLT, a privacy-preserving inference framework for transformer models that supports efficient matrix multiplications and nonlinear computations. Combined with our novel machine learning optimizations, BOLT reduces the communication cost by 10.91×. Our evaluation on diverse datasets demonstrates that BOLT maintains comparable accuracy to floating-point models and achieves 4.8-9.5× faster inference across various network settings compared to the state-of-the-art system.
Qi Pang, Jinhao Zhu, Helen Möllering, Wenting Zheng, Thomas Schneider 0003
SP1
2024 Provably Valid and Diverse Mutations of Real-World Media Data for DNN Testing
abstract
Deep neural networks (DNNs) often accept high-dimensional media data (e.g., photos, text, and audio) and understand their perceptual content (e.g., a cat). To test DNNs,diverseinputs are needed to trigger mis-predictions. Some preliminary works use byte-level mutations or domain-specific filters (e.g., foggy), whose enabled mutations may be limited and likely error-prone. State-of-the-art (SOTA) works employ deep generative models to generate (infinite) inputs. Also, to keep the mutated inputs perceptuallyvalid(e.g., a cat remains a “cat” after mutation), existing efforts rely on imprecise and less generalizable heuristics. This study revisits two key objectives in media input mutation—perception diversity (Div) and validity (Val) — in a rigorous manner based on manifold, a well-developed theory capturing perceptions of high-dimensional media data in a low-dimensional space. We show important results thatDivandValinextricably bound each other, and prove that SOTA generative model-based methods fundamentally fail to mutatereal-world media data(either sacrificingDivorVal). In contrast, we discuss the feasibility of mutating real-world media data with provably highDivandValbased on manifold.Following, we concretize the technical solution of mutating media data of various formats (images, audios, text) via aunifiedmanner based on manifold. Specifically, when media data are projected into a low-dimensional manifold, the data can be mutated by walking on the manifold with certain directions and step sizes. When contrasted with the input data, the mutated data exhibit encouragingDivin the perceptual traits (e.g., lying vs. standing dog) while retaining reasonably highVal(i.e., a dog remains a dog).We implement our techniques inDeepwalkfor testing DNNs.Deepwalkconstructs manifolds for media data offline. In online testing,Deepwalkwalks on manifolds to generate mutated media data with provably highDivandVal. Our evaluation tests DNNs executing various tasks (e.g., classification, self-driving, machine translation) and media data of different types (image, audio, text).Deepwalkoutperforms prior methods in terms of the testing comprehensiveness and can find more error-triggering inputs with higher quality. The tested DNNs, after repaired usingDeepwalk’sfindings, exhibit better accuracy.
Yuanyuan Yuan 0001, Qi Pang, Shuai Wang 0011
IEEE Trans. Software Eng.2
2023 Detecting and Repairing Deviated Outputs of Compressed Models
Yichen Li 0001, Qi Pang, Dongwei Xiao, Zhibo Liu 0001, Shuai Wang 0011
ACML2
2023 Byzantine-Robust Federated Learning with Optimal Statistical Rates
abstract
We propose Byzantine-robust federated learning protocols with nearly optimal statistical rates based on recent progress in high dimensional robust statistics. In contrast to prior work, our proposed protocols improve the dimension dependence and achieve a near-optimal statistical rate for strongly convex losses. We also provide statistical lower bound for the problem. For experiments, we benchmark against competing protocols and show the empirical superiority of the proposed protocols.
Banghua Zhu, Lun Wang 0001, Qi Pang, Shuai Wang 0011, Jiantao Jiao, Dawn Song, Michael I. Jordan
AISTATS3
2023 Secure Federated Correlation Test and Entropy Estimation
abstract
We propose the first federated correlation test framework compatible with secure aggregation, namely FED-$\chi^2$. In our protocol, the statistical computations are recast as frequency moment estimation problems, where the clients collaboratively generate a shared projection matrix and then use stable projection to encode the local information in a compact vector. As such encodings can be linearly aggregated, secure aggregation can be applied to conceal the individual updates. We formally establish the security guarantee of FED-$\chi^2$ by proving that only the minimum necessary information (i.e., the correlation statistics) is revealed to the server. We show that our protocol can be naturally extended to estimate other statistics that can be recast as frequency moment estimations. By accommodating Shannon’e Entropy in FED-$\chi^2$, we further propose the first secure federated entropy estimation protocol, FED-$H$. The evaluation results demonstrate that FED-$\chi^2$ and FED-$H$ achieve good performance with small client-side computation overhead in several real-world case studies.
Qi Pang, Lun Wang 0001, Shuai Wang 0011, Wenting Zheng, Dawn Song
ICML1
2023 Revisiting Neuron Coverage for DNN Testing: A Layer-Wise and Distribution-Aware Criterion
abstract
Various deep neural network (DNN) coverage criteria have been proposed to assess DNN test inputs and steer input mutations. The coverage is characterized via neurons having certain outputs, or the discrepancy between neuron outputs. Nevertheless, recent research indicates that neuron coverage criteria show little correlation with test suite quality. In general, DNNs approximate distributions, by incorporating hierarchical layers, to make predictions for inputs. Thus, we champion to deduce DNN behaviors based on its approximated distributions from a layer perspective. A test suite should be assessed using its induced layer output distributions. Accordingly, to fully examine DNN behaviors, input mutation should be directed toward diversifying the approximated distributions. This paper summarizes eight design requirements for DNN coverage criteria, taking into account distribution properties and practical concerns. We then propose a new criterion, Neural Coverage (nlc),that satisfies all design requirements. NLC treats a single DNN layer as the basic computational unit (rather than a single neuron) and captures four critical properties of neuron output distributions. Thus, NL C accurately describes how DNNs comprehend inputs via approximated distributions. We demonstrate that NLC is significantly correlated with the diversity of a test suite across a number of tasks (classification and generation) and data formats (image and text). Its capacity to discover DNN prediction errors is promising. Test input mutation guided by NLC results in a greater quality and diversity of exposed erroneous behaviors.
Yuanyuan Yuan 0001, Qi Pang, Shuai Wang 0011
ICSE2
2023 ADI: Adversarial Dominating Inputs in Vertical Federated Learning Systems
abstract
Vertical federated learning (VFL) system has recently become prominent as a concept to process data distributed across many individual sources without the need to centralize it. Multiple participants collaboratively train models based on their local data in a privacy-aware manner. To date, VFL has become a de facto solution to securely learn a model among organizations, allowing knowledge to be shared without compromising privacy of any individuals. Despite the prosperous development of VFL systems, we find that certain inputs of a participant, named adversarial dominating inputs (ADIs), can dominate the joint inference towards the direction of the adversary's will and force other (victim) participants to make negligible contributions, losing rewards that are usually offered regarding the importance of their contributions in federated learning scenarios. We conduct a systematic study on ADIs by first proving their existence in typical VFL systems. We then propose gradient-based methods to synthesize ADIs of various formats and exploit common VFL systems. We further launch greybox fuzz testing, guided by the saliency score of "victim" participants, to perturb adversary-controlled inputs and systematically explore the VFL attack surface in a privacy-preserving manner. We conduct an in-depth study on the influence of critical parameters and settings in synthesizing ADIs. Our study reveals new VFL attack opportunities, promoting the identification of unknown threats before breaches and building more secure VFL systems.
Qi Pang, Yuanyuan Yuan 0001, Shuai Wang 0011, Wenting Zheng
SP1
2022 MDPFuzz: testing models solving Markov decision processes
abstract
The Markov decision process (MDP) provides a mathematical frame- work for modeling sequential decision-making problems, many of which are crucial to security and safety, such as autonomous driving and robot control. The rapid development of artificial intelligence research has created efficient methods for solving MDPs, such as deep neural networks (DNNs), reinforcement learning (RL), and imitation learning (IL). However, these popular models solving MDPs are neither thoroughly tested nor rigorously reliable.
Qi Pang, Yuanyuan Yuan 0001, Shuai Wang 0011
ISSTA1
2022 Unveiling Hidden DNN Defects with Decision-Based Metamorphic Testing
abstract
Contemporary DNN testing works are frequently conducted using metamorphic testing (MT). In general, de facto MT frameworks mutate DNN input images using semantics-preserving mutations and determine if DNNs can yield consistent predictions. Nevertheless, we find that DNNs may rely on erroneous decisions (certain components on the DNN inputs) to make predictions, which may still retain the outputs by chance. Such DNN defects would be neglected by existing MT frameworks. Erroneous decisions, however, would likely result in successive mis-predictions over diverse images that may exist in real-life scenarios.
Yuanyuan Yuan 0001, Qi Pang, Shuai Wang 0011
ASE2
2022 PrivGuard: Privacy Regulation Compliance Made Easier
Lun Wang 0001, Usmann Khan, Joseph P. Near, Qi Pang, Jithendaraa Subramanian, Neel Somani, Peng Gao 0008, Andrew Low, Dawn Song
USENIX Security Symposium4
2022 Automated Side Channel Analysis of Media Software with Manifold Learning
Yuanyuan Yuan 0001, Qi Pang, Shuai Wang 0011
USENIX Security Symposium2
2022 NoLeaks: Differentially Private Causal Discovery Under Functional Causal Model
abstract
Causal inference is widely used in clinical research, economic analysis, and other fields. As is the case with many statistical data, the findings of causal discovery (i.e., causal graph) might leak demographic information of participants. For example, a causal link between one genome and a rare disease can reveal the participation of a minority patient in genome-ide association studies. To date, differential privacy has served as the de facto foundation for guaranteeing the privacy of causal discovery algorithms. However, existing approaches to protecting causal discovery from privacy leakage rely heavily on private conditional independence tests, which generate a considerable amount of noise and are thus prone to inaccuracy. As a result of their limited accuracy and scalability, they are insufficient for non-trivial datasets (e.g., those with more than ten variables). In this paper, we advocate a novel focus on enforcing privacy for causal discovery algorithms based on functional causal models. First, we propose NOLEAKS, a differentially private causal discovery algorithm, which manifests both high accuracy and efficiency compared with prior works. Second, we design a quasi-Newton numerical optimization algorithm for solving NOLEAKS in a highly efficient way. Third, we evaluate NOLEAKS using both public benchmarks and synthetic data. We observe that NOLEAKS achieves comparable performance or even surpasses the state-of-the-art (non-private) approaches. We also find encouraging results that NOLEAKS can smoothly scale to large datasets, on which existing works would fail. Through a case study and a downstream application, we observe encouraging results on the versatile usages of NOLEAKS.
Pingchuan Ma 0004, Zhenlan Ji, Qi Pang, Shuai Wang 0011
IEEE Trans. Inf. Forensics Secur.3
2021 mID: Tracing Screen Photos via Moiré Patterns
Yushi Cheng, Xiaoyu Ji 0001, Lixu Wang, Qi Pang, Yi-Chao Chen 0001, Wenyuan Xu 0001
USENIX Security Symposium4
2020 Towards practical differentially private causal graph discovery
abstract
Causal graph discovery refers to the process of discovering causal relation graphs from purely observational data. Like other statistical data, a causal graph might leak sensitive information about participants in the dataset. In this paper, we present a differentially private causal graph discovery algorithm, Priv-PC, which improves both utility and running time compared to the state-of-the-art. The design of Priv-PC follows a novel paradigm called sieve-and-examine which uses a small amount of privacy budget to filter out “insignificant” queries, and leverages the remaining budget to obtain highly accurate answers for the “significant” queries. We also conducted the first sensitivity analysis for conditional independence tests including conditional Kendall’s τ and conditional Spearman’s ρ. We evaluated Priv-PC on 7 public datasets and compared with the state-of-the-art. The results show that Priv-PC achieves 10.61 to 293.87 times speedup and better utility. The implementation of Priv-PC, including the code used in our evaluation, is available at https://github.com/sunblaze-ucb/ Priv-PC-Differentially-Private-Causal-Graph-Discovery.
Lun Wang 0001, Qi Pang, Dawn Song
NeurIPS2
2016 Probabilistic linguistic term sets in multi-attribute group decision making
Qi Pang, Hai Wang 0005, Zeshui Xu
Inf. Sci.1