VLDB 2026 Research / reviewers in the wild / expert
Xingbin Wang
dblp:119/6492
· DBLP profile ↗
22ranked-venue papers
10as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoLoRA: A Collaborative Scheduling Framework for Multi-Tenant LoRA LLM InferenceabstractLarge Language Models (LLM) incur substantial resource costs during inference, driving widespread interest in Parameter-Efficient Fine-Tuning (PEFT) techniques. Among these, LoRA dramatically reduces overhead by updating only a few low-rank adapters. However, Multi-tenant LoRA LLM inference faces challenges from heterogeneous requests and latency-throughput trade-offs. Moreover, inefficient adapter reuse, poor cache management, and non-adaptive batching strategies severely restrict inference efficiency, service quality, resource utilization, and fairness. To address these challenges, we propose CoLoRA—a collaborative scheduling framework for multi-tenant LoRA LLM inference, comprising four core modules: (1) Adaptive Priority Scheduling (APS), which dynamically integrates queue waiting time, adapter residency status, and SLA urgency to compute task priorities; (2) Adapter-Aware Scheduling (AAS), which enhances cache management by prioritizing SLA-critical, frequently used, and fairly shared adapters, thus reducing cold-start latency and fragmentation; (3) Load-Aware Batch Scheduling (LBS), which combines real-time GPU utilization and queue depth to adaptively form batches and coalesce tasks targeting the same adapter, thereby improving parallelism while controlling latency; and (4) Unified Scheduler (US), which periodically gathers system metadata to orchestrate the submodules collaboratively and employs a feedback loop to optimize global strategies online. Evaluation on realistic multi-tenant workloads and popular open-source LLM shows that CoLoRA, compared to conventional baselines, increases overall system throughput by 56.5%, reduces P95 latency of online requests by 34%, and significantly enhances GPU utilization and tenant-level fairness, demonstrating its promise for large-scale inference services. Zechao Lin, Xingbin Wang, Dan Meng 0002, Rui Hou 0001 |
ASP-DAC | 2 |
| 2026 | AegisX: An Acceleration Framework for Moving Target Defenses to Boost Adversarial Robustness and Computational EfficiencyabstractMoving target defenses are a key proactive technique for defending against adversarial attacks by increasing the dynamics, randomness and diversification of the system, and constructs an ensemble of multiple DNNs to obtain stronger defense. However, their models run noticeably slower on existing DNN accelerators than single-model inference. Moreover, they lack the hardware assistance for supporting scheduling mechanism of moving target defenses.Accordingly, our work is the first to propose a novel acceleration framework for moving target defenses against adversarial attacks, called AegisX, which provides architecture support for accelerating the ensemble of moving target defenses with dynamic grouping and early stop. Our hardware architecture integrates three key innovations: 1) a novel scheduler enabling concurrent operator executions and supporting moving target defense scheduling patterns; 2) a utilization-aware resource allocation strategy that fully exploits temporal and spatial sharing; 3) a hierarchical scheduling mechanism with early stop and an idleness-aware resource borrowing scheme to utilize idle computing cores effectively. Benchmark evaluations reveal substantial gains in throughput and energy efficiency. Xingbin Wang, Chaochao Zhang, Rui Hou 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2025 | ROBIN: A Novel Framework for Accelerating Robust Multi-Variant TrainingabstractRobust variants represent a promising method to enhance model robustness against adversarial attacks through exploring diverse neural network architectures. However, the significant computational demand of training multiple variants often restricts adversarial defense techniques to a narrow range of model architectures, thus failing to fully exploit the robustness benefits of architectural variations. In this paper, we first reveal that function-preserving knowledge transfer can significantly speed up adversarial training of different architecture variants. Then, we propose ROBIN, a framework for accelerating adversarially robust multi-variant training. By utilizing the architectural similarities among variants, ROBIN facilitates efficient weight transformation across models via two tensor-level atomic operations, hastening the convergence of multiple variants. Our experiments indicate that ROBIN can accelerate the adversarial training process of various architecture variants by 2.56 × to 4.27 ×, enabling efficient exploration of robust network architectures. Yan Wang 0122, Xingbin Wang, Yulan Su, Sisi Zhang, Zechao Lin, Dan Meng 0002, Rui Hou 0001 |
ASP-DAC | 2 |
| 2025 | LoRATEE: A Secure and Efficient Inference Framework for Multi-Tenant LoRA LLMs Based on TEEabstractLow-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning approach that adaptes pre-trained Large Language Models (LLMs) to multi-tenant tasks by generating a variety of LoRA adapters. However, this approach faces significant security challenges and is particularly susceptible to malicious servers stealing model parameters and sensitive data. Existing research on addressing security risks in multi-tenant environments remains constrained and insufficient. This paper explores the security challenges and proposes the LoRATEE framework, which embeds LoRA adapters within a server-side Trusted Execution Environment (TEE) and employs a lightweight One-Time Pad (OTP) encryption mechanism to ensure secure data transmission. Additionally, we design a dynamic LoRA adapter prefetching mechanism to reduce I/O latency. Moreover, a LoRA adapter module equivalence-sharing strategy based on Parameter-Efficient Fine-Tuning (PEFT) and minimalist design principles was introduced to optimize adapters loading. Experimental results show that LoRATEE maintains inference efficiency while securing multi-tenant LoRA LLMs systems. Zechao Lin, Sisi Zhang, Xingbin Wang, Yulan Su, Yan Wang 0122, Rui Hou 0001, Dan Meng 0002 |
ICASSP | 3 |
| 2025 | Jack of All Trades, Master of None: PMP-Guided Adaptive Multi-Teacher Distillation with Meta-LearningabstractTo enhance the robustness and accuracy of the small model, existing approaches combine adversarial training with knowledge distillation, introducing a comprehensive single-teacher model to improve the performance of the student model (small model). However, due to the limited knowledge of a teacher model, it appears "knowledge gain saturation" phenomenon. Therefore, we propose a PMP-Guided Adaptive Multi-Teacher Distillation with Meta-Learning. Pontryagin’s Maximum Principle is employed to solve the issue of inconsistent teaching objectives among teachers causing distinct optimization directions. Meanwhile, Meta-learning-network is designed to tackle the problem of a student struggling to balance the learned knowledge. A series of experiments conducted on public datasets demonstrate that our approach outperforms the state-of-the-art methods against various adversarial attacks. Sisi Zhang, Zechao Lin, Xingbin Wang, Yulan Su, Yan Wang 0122, Rui Hou 0001, Dan Meng 0002 |
ICASSP | 3 |
| 2025 | Poseidon: A NAS-Based Ensemble Defense Method Against Multiple Perturbations
Yulan Su, Sisi Zhang, Zechao Lin, Xingbin Wang, Lutan Zhao, Dan Meng 0002, Rui Hou 0001 |
MMM (3) | 4 |
| 2025 | RobSparse: Automatic Search for GPU-Friendly Robust and Sparse Vision Transformers
Yulan Su, Sisi Zhang, Yan Wang 0122, Xingbin Wang, Lutan Zhao, Dan Meng 0002, Rui Hou 0001 |
MMM (3) | 4 |
| 2024 | SecPaging: Secure Enclave Paging with Hardware-Enforced Protection against Controlled-Channel AttacksabstractAs a prevalent privacy-preserving technology, Trusted Execution Environment has become widely adopted in numerous commercial processors. Nonetheless, they remain susceptible to various controlled-channel attacks. Untrusted operating systems can deduce enclave secrets by manipulating page tables or observing allocation- or swap-based page faults. In this paper, we propose SecPaging, a novel secure enclave paging mechanism based on hardware-enforced and microcode-supported protection to prevent these attacks. First, enclave PTEs are protected through hardware isolation, preventing privileged attackers from malicious tampering or observations. Second, an Eager-Allocation mechanism is employed to prevent allocation-based controlled-channel attacks. Besides, a Record-Reload mechanism is proposed to prevent swap-based controlled-channel attacks. We simulate SecPaging on real SGX. Experiments demonstrate that controlled channel attacks can be defended with minimal performance overhead. Yunkai Bai, Peinan Li, Yubiao Huang, Shiwen Wang 0002, Xingbin Wang, Dan Meng 0002, Rui Hou 0001 |
DAC | 5 |
| 2024 | Garrison: A High-Performance GPU-Accelerated Inference System for Adversarial Ensemble DefenseabstractIn the face of huge threats from adversarial attacks, developing an efficient defense mechanism is crucial for deep learning systems. Adversarial ensemble defense method is one of the most effective techniques for defending against adversarial attacks, which constructs ensembles of multiple DNNs to improve model's robustness. However, deploying ensemble defense methods on existing DNN inference systems is inefficient and impractical due to their dynamics and randomness. To this end, we propose an inference system for adversarial ensemble defense called Garrison, which can deliver robust and low-latency predictions using Multi-Instance GPUs. Garrison employs a multi-granularity GPU partitioning strategy, optimizing hardware utilization by capitalizing on the intrinsic heterogeneity of GPUs. It also integrates a reinforcement learning-based scheduling mechanism, enabling random ensemble of diverse defense models to enhance robustness while maintaining bounded latency. Our evaluations show that Garrison can improve adversarial robustness by up to 24.5%, while accelerating ensemble inference by up to 6.6X compared to the state-of-the-art inference framework. Yan Wang 0122, Xingbin Wang, Zechao Lin, Yulan Su, Sisi Zhang, Rui Hou 0001, Dan Meng 0002 |
DAC | 2 |
| 2024 | EnTurbo: Accelerate Confidential Serverless Computing via Parallelizing Enclave Startup ProcedureabstractServerless computing has gained widespread attention, and Trusted Execution Environments (TEEs) are well-suited for safeguarding user privacy. However, the additional startup procedure introduced by TEEs imposes considerable performance overhead on confidential serverless workloads. This paper introduces a novel parallelized enclave startup design, EnTurbo, which eliminates the integrity dependence of the enclave startup procedure, accelerating it while ensuring its security. Additionally, EnTurbo parallelizes the measurement procedure, enabling multi-thread measurement for acceleration with provable security. We evaluate EnTurbo by running confidential serverless workloads on SGX simulation mode. Results show that EnTurbo effectively speeds up enclave serverless by 1.42x-6.48x (SGXv1) and 1.33x-3.76x (SGXv2). Yifan Zhu 0008, Peinan Li, Yunkai Bai, Yubiao Huang, Shiwen Wang 0002, Xingbin Wang, Dan Meng 0002, Rui Hou 0001 |
DAC | 6 |
| 2024 | FakeGuard: Novel Architecture Support for Deepfake Detection Networks
Xingbin Wang, Dan Meng 0002, Rui Hou 0001, Yan Wang 0122 |
Euro-Par (2) | 1 |
| 2024 | EnsGuard: A Novel Acceleration Framework for Adversarial Ensemble LearningabstractTo defend against various adversarial attacks, it is essential to develop a robust and high computing efficiency defence framework. Adversarial ensemble learning is the most effective technique for defending against adversarial example attacks, which constructs ensembles of multiple DNNs with adversarial training to obtain stronger defense. However, ensemble models run noticeably slower on existing DNN accelerators than single-model inference. Deploying ensemble models on the existing DNN accelerators leads to many critical issues such as the underutilization of hardware resources. To tackle emerging challenges, we propose EnsGuard, a dynamic asymmetric multi-core systolic array architecture for adversarial ensemble learning inference to fully exploit both static and dynamic parallelism of ensemble models. Specifically, on the hardware level, we propose a novel instruction set extension and develop efficient architecture components to fully exploit the new hardware abstraction of scattered idle computing cores, and use them to dynamically create on-the-fly Neural Processing Units (fNPUs). Moreover, we propose a computing power recycle mechanism to run on-the-fly models (small models) on fNPUs by carefully orchestrating execution order of ensemble models for maximizing hardware resources and bandwidth utilization. On the software level, EnsGuard adopts an integrated hardware/randomized ensemble co-design optimizer aiming at winning both faster inference and higher adversarial robustness. On top of that, a multi-model mapping method based on decision tree is proposed to enable the interleaving of different DNN executions both spatially and temporally, and mitigate straggler problems. Evaluation with a diverse set of workloads shows significant gains in throughput (4.4×) and energy reduction (3.2×). Xingbin Wang, Yan Wang 0122, Yulan Su, Sisi Zhang, Dan Meng 0002, Rui Hou 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | A Hybrid Sparse-dense Defensive DNN Accelerator Architecture against Adversarial Example AttacksabstractUnderstanding how to defend against adversarial attacks is crucial for ensuring the safety and reliability of these systems in real-world applications. Various adversarial defense methods are proposed, which aim at improving the robustness of neural networks against adversarial attacks by changing the model structure, adding detection networks, and adversarial purification network. However, deploying adversarial defense methods in existing DNN accelerators or defensive accelerators leads to many key issues. To address these challenges, this article proposessDNNGuard, an elastic heterogeneous DNN accelerator architecture that can efficiently orchestrate the simultaneous execution of original (target) DNN networks and thedetectalgorithm or network. It not only supports for dense DNN detect algorithms, but also allows for sparse DNN defense methods and other mixed dense-sparse (e.g., dense-dense and sparse-dense) workloads to fully exploit the benefits of sparsity. sDNNGuard with a CPU core also supports the non-DNN computing and allows the special layer of the neural network, and used for the conversion for sparse storage format for weights and activation values. To reduce off-chip traffic and improve resources utilization, a new hardware abstraction with elastic on-chip buffer/computing resource management is proposed to achieve dynamical resource scheduling mechanism. We propose anextended AI instruction setfor neural networks synchronization, task scheduling and efficient data interaction. Experiment results show that sDNNGuard can effectively validate the legitimacy of the input samples in parallel with the target DNN model, achieving an average 1.42× speedup compared with the state-of-the-art accelerators. Xingbin Wang, Boyan Zhao, Yulan Su, Sisi Zhang, Fengkai Yuan, Dan Meng 0002, Rui Hou 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2023 | A Efficient Adaptive Data Rate Algorithm in LoRaWAN Networks: K-ADR
Xingbin Wang |
APNOMS | 3 |
| 2022 | Mimic Octopus Attack: Dynamic Camouflage Adversarial Examples Using Mimetic Feature for 3D Humans
Jing Li 0114, Sisi Zhang, Xingbin Wang, Rui Hou 0001 |
Inscrypt | 3 |
| 2022 | Multi-modal Face Anti-spoofing Using Channel Cross Fusion Network and Global Depth-Wise Convolution
Shidong Chen, Mengfan Tang, Xingbin Wang |
KSEM (2) | 5 |
| 2021 | NASGuard: A Novel Accelerator Architecture for Robust Neural Architecture Search (NAS) NetworksabstractDue to the wide deployment of deep learning applications in safety-critical systems, robust and secure execution of deep learning workloads is imperative. Adversarial examples, where the inputs are carefully designed to mislead the machine learning model is among the most challenging attacks to detect and defeat. The most dominant approach for defending against adversarial examples is to systematically create a network architecture that is sufficiently robust. Neural Architecture Search (NAS) has been heavily used as the de facto approach to design robust neural network models, by using the accuracy of detecting adversarial examples as a key metric of the neural network’s robustness. While NAS has been proven effective in improving the robustness (and accuracy in general), the NAS-generated network models run noticeably slower on typical DNN accelerators than the hand-crafted networks, mainly because DNN accelerators are not optimized for robust NAS-generated models. In particular, the inherent multi-branch nature of NAS-generated networks causes unacceptable performance and energy overheads.To bridge the gap between the robustness and performance efficiency of deep learning applications, we need to rethink the design of AI accelerators to enable efficient execution of robust (auto-generated) neural networks. In this paper, we propose a novel hardware architecture, NASGuard, which enables efficient inference of robust NAS networks. NASGuard leverages a heuristic multi-branch mapping model to improve the efficiency of the underlying computing resources. Moreover, NASGuard addresses the load imbalance problem between the computation and memory-access tasks from multi-branch parallel computing. Finally, we propose a topology-aware performance prediction model for data prefetching, to fully exploit the temporal and spatial localities of robust NAS-generated architectures. We have implemented NASGuard with Verilog RTL. The evaluation results show that NASGuard achieves an average speedup of 1.74× over the baseline DNN accelerator. Xingbin Wang, Boyan Zhao, Rui Hou 0001, Amro Awad, Zhihong Tian 0001, Dan Meng 0002 |
ISCA | 1 |
| 2020 | DNNGuard: An Elastic Heterogeneous DNN Accelerator Architecture against Adversarial AttacksabstractRecent studies show that Deep Neural Networks (DNN) are vulnerable to adversarial samples that are generated by perturbing correctly classified inputs to cause the misclassification of DNN models. This can potentially lead to disastrous consequences, especially in security-sensitive applications such as unmanned vehicles, finance and healthcare. Existing adversarial defense methods require a variety of computing units to effectively detect the adversarial samples. However, deploying adversary sample defense methods in existing DNN accelerators leads to many key issues in terms of cost, computational efficiency and information security. Moreover, existing DNN accelerators cannot provide effective support for special computation required in the defense methods. Xingbin Wang, Rui Hou 0001, Boyan Zhao, Fengkai Yuan, Dan Meng 0002, Xuehai Qian |
ASPLOS | 1 |
| 2020 | SNA: A Siamese Network Accelerator to Exploit the Model-Level Parallelism of Hybrid Network StructureabstractSiamese network is compute-intensive learning model with growing applicability in a wide range of domains. However, state-of-art deep neural network (DNN) accelerators would not work efficiently for Siamese network, as their designs do not account for the algorithm properties of Siamese network. In this paper, we propose a Siamese network accelerator called SNA, the first Simultaneous Multi-Threading (SMT) hardware architecture to perform Siamese network inference with high performance and energy efficiency. We devise an adaptive inter-model computing resource partition and flexible on-chip buffer management mechanism based on the model parallelism and SMT design philosophy. Our architecture is implemented in Verilog and synthesized in a 65nm technology using Synopsys design tools. We also evaluate it with several typical Siamese networks. Compared to the state-of-art accelerator, on average, the SNA architecture offers 2.1x speedup and 1.48x energy reduction. Xingbin Wang, Boyan Zhao, Rui Hou 0001, Dan Meng 0002 |
DATE | 1 |
| 2019 | NPUFort: a secure architecture of DNN accelerator against model inversion attackabstractDeep neural network (DNN) models are widely used for inference in many application scenarios. DNN accelerators are not designed with security in mind, but for higher performance and lower energy consumption. Hence, they are suffering from the security risk of being attacked. The insecure design flaws of existing DNN accelerators can be exploited to recover the structure of DNN model from the plain instructions, thus the runtime environment can be controlled to obtain the weights of DNN model. Furthermore, the structure of DNN model running on the accelerator is acquired by the side channel information and interrupt status register. To protect general DNN accelerator from being attacked by model inversion attack, this paper proposes a secure and general architecture called NPUFort, which guarantees the confidentiality of the parameters of DNN model and mitigates side-channel information leakage. The experimental results demonstrate the feasibility and effectiveness of the secure architecture of DNN accelerators with negligible performance overhead. Xingbin Wang, Rui Hou 0001, Yifan Zhu 0008, Dan Meng 0002 |
CF | 1 |
| 2012 | An R package suite for microarray meta-analysis in quality control, differentially expressed gene analysis and pathway enrichment detectionabstractSUMMARY: With the rapid advances and prevalence of high-throughput genomic technologies, integrating information of multiple relevant genomic studies has brought new challenges. Microarray meta-analysis has become a frequently used tool in biomedical research. Little effort, however, has been made to develop a systematic pipeline and user-friendly software. In this article, we present MetaOmics, a suite of three R packages MetaQC, MetaDE and MetaPath, for quality control, differentially expressed gene identification and enriched pathway detection for microarray meta-analysis. MetaQC provides a quantitative and objective tool to assist study inclusion/exclusion criteria for meta-analysis. MetaDE and MetaPath were developed for candidate marker and pathway detection, which provide choices of marker detection, meta-analysis and pathway analysis methods. The system allows flexible input of experimental data, clinical outcome (case-control, multi-class, continuous or survival) and pathway databases. It allows missing values in experimental data and utilizes multi-core parallel computing for fast implementation. It generates informative summary output and visualization plots, operates on different operation systems and can be expanded to include new algorithms or combine different types of genomic data. This software suite provides a comprehensive tool to conveniently implement and compare various genomic meta-analysis pipelines. AVAILABILITY: http://www.biostat.pitt.edu/bioinfo/software.htm CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xingbin Wang, Dongwan D. Kang, Kui Shen, Chi Song, Shuya Lu, Lun-Ching Chang, Serena G. Liao, Zhiguang Huo, Shaowu Tang, Naftali Kaminski, Etienne Sibille, George C. Tseng |
Bioinform. | 1 |
| 2012 | Detecting disease-associated genes with confounding variable adjustment and the impact on genomic meta-analysis: With application to major depressive disorderabstractBACKGROUND: Detecting candidate markers in transcriptomic studies often encounters difficulties in complex diseases, particularly when overall signals are weak and sample size is small. Covariates including demographic, clinical and technical variables are often confounded with the underlying disease effects, which further hampers accurate biomarker detection. Our motivating example came from an analysis of five microarray studies in major depressive disorder (MDD), a heterogeneous psychiatric illness with mostly uncharacterized genetic mechanisms. RESULTS: We applied a random intercept model to account for confounding variables and case-control paired design. A variable selection scheme was developed to determine the effective confounders in each gene. Meta-analysis methods were used to integrate information from five studies and post hoc analyses enhanced biological interpretations. Simulations and application results showed that the adjustment for confounding variables and meta-analysis improved detection of biomarkers and associated pathways. CONCLUSIONS: The proposed framework simultaneously considers correction for confounding variables, selection of effective confounders, random effects from paired design and integration by meta-analysis. The approach improved disease-related biomarker and pathway detection, which greatly enhanced understanding of MDD neurobiology. The statistical framework can be applied to similar experimental design encountered in other complex and heterogeneous diseases. Xingbin Wang, Chi Song, Etienne Sibille, George C. Tseng |
BMC Bioinform. | 1 |