Yan Wang 0122

dblp:59/2227-122 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0005-2485-0899ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond similarity: Mutual information-guided retrieval for in-context learning in VQA
Zezhong Lv, Jian Zhao 0006, Yan Wang 0122, Yuchen Yuan, Yuchu Jiang, Wenqi Ren, Xuelong Li 0001
Pattern Recognit.4
2025 ROBIN: A Novel Framework for Accelerating Robust Multi-Variant Training
abstract
Robust variants represent a promising method to enhance model robustness against adversarial attacks through exploring diverse neural network architectures. However, the significant computational demand of training multiple variants often restricts adversarial defense techniques to a narrow range of model architectures, thus failing to fully exploit the robustness benefits of architectural variations. In this paper, we first reveal that function-preserving knowledge transfer can significantly speed up adversarial training of different architecture variants. Then, we propose ROBIN, a framework for accelerating adversarially robust multi-variant training. By utilizing the architectural similarities among variants, ROBIN facilitates efficient weight transformation across models via two tensor-level atomic operations, hastening the convergence of multiple variants. Our experiments indicate that ROBIN can accelerate the adversarial training process of various architecture variants by 2.56 × to 4.27 ×, enabling efficient exploration of robust network architectures.
Yan Wang 0122, Xingbin Wang, Yulan Su, Sisi Zhang, Zechao Lin, Dan Meng 0002, Rui Hou 0001
ASP-DAC1
2025 LoRATEE: A Secure and Efficient Inference Framework for Multi-Tenant LoRA LLMs Based on TEE
abstract
Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning approach that adaptes pre-trained Large Language Models (LLMs) to multi-tenant tasks by generating a variety of LoRA adapters. However, this approach faces significant security challenges and is particularly susceptible to malicious servers stealing model parameters and sensitive data. Existing research on addressing security risks in multi-tenant environments remains constrained and insufficient. This paper explores the security challenges and proposes the LoRATEE framework, which embeds LoRA adapters within a server-side Trusted Execution Environment (TEE) and employs a lightweight One-Time Pad (OTP) encryption mechanism to ensure secure data transmission. Additionally, we design a dynamic LoRA adapter prefetching mechanism to reduce I/O latency. Moreover, a LoRA adapter module equivalence-sharing strategy based on Parameter-Efficient Fine-Tuning (PEFT) and minimalist design principles was introduced to optimize adapters loading. Experimental results show that LoRATEE maintains inference efficiency while securing multi-tenant LoRA LLMs systems.
Zechao Lin, Sisi Zhang, Xingbin Wang, Yulan Su, Yan Wang 0122, Rui Hou 0001, Dan Meng 0002
ICASSP5
2025 Jack of All Trades, Master of None: PMP-Guided Adaptive Multi-Teacher Distillation with Meta-Learning
abstract
To enhance the robustness and accuracy of the small model, existing approaches combine adversarial training with knowledge distillation, introducing a comprehensive single-teacher model to improve the performance of the student model (small model). However, due to the limited knowledge of a teacher model, it appears "knowledge gain saturation" phenomenon. Therefore, we propose a PMP-Guided Adaptive Multi-Teacher Distillation with Meta-Learning. Pontryagin’s Maximum Principle is employed to solve the issue of inconsistent teaching objectives among teachers causing distinct optimization directions. Meanwhile, Meta-learning-network is designed to tackle the problem of a student struggling to balance the learned knowledge. A series of experiments conducted on public datasets demonstrate that our approach outperforms the state-of-the-art methods against various adversarial attacks.
Sisi Zhang, Zechao Lin, Xingbin Wang, Yulan Su, Yan Wang 0122, Rui Hou 0001, Dan Meng 0002
ICASSP5
2025 RobSparse: Automatic Search for GPU-Friendly Robust and Sparse Vision Transformers
Yulan Su, Sisi Zhang, Yan Wang 0122, Xingbin Wang, Lutan Zhao, Dan Meng 0002, Rui Hou 0001
MMM (3)3
2024 Garrison: A High-Performance GPU-Accelerated Inference System for Adversarial Ensemble Defense
abstract
In the face of huge threats from adversarial attacks, developing an efficient defense mechanism is crucial for deep learning systems. Adversarial ensemble defense method is one of the most effective techniques for defending against adversarial attacks, which constructs ensembles of multiple DNNs to improve model's robustness. However, deploying ensemble defense methods on existing DNN inference systems is inefficient and impractical due to their dynamics and randomness. To this end, we propose an inference system for adversarial ensemble defense called Garrison, which can deliver robust and low-latency predictions using Multi-Instance GPUs. Garrison employs a multi-granularity GPU partitioning strategy, optimizing hardware utilization by capitalizing on the intrinsic heterogeneity of GPUs. It also integrates a reinforcement learning-based scheduling mechanism, enabling random ensemble of diverse defense models to enhance robustness while maintaining bounded latency. Our evaluations show that Garrison can improve adversarial robustness by up to 24.5%, while accelerating ensemble inference by up to 6.6X compared to the state-of-the-art inference framework.
Yan Wang 0122, Xingbin Wang, Zechao Lin, Yulan Su, Sisi Zhang, Rui Hou 0001, Dan Meng 0002
DAC1
2024 FakeGuard: Novel Architecture Support for Deepfake Detection Networks
Xingbin Wang, Dan Meng 0002, Rui Hou 0001, Yan Wang 0122
Euro-Par (2)4
2024 EnsGuard: A Novel Acceleration Framework for Adversarial Ensemble Learning
abstract
To defend against various adversarial attacks, it is essential to develop a robust and high computing efficiency defence framework. Adversarial ensemble learning is the most effective technique for defending against adversarial example attacks, which constructs ensembles of multiple DNNs with adversarial training to obtain stronger defense. However, ensemble models run noticeably slower on existing DNN accelerators than single-model inference. Deploying ensemble models on the existing DNN accelerators leads to many critical issues such as the underutilization of hardware resources. To tackle emerging challenges, we propose EnsGuard, a dynamic asymmetric multi-core systolic array architecture for adversarial ensemble learning inference to fully exploit both static and dynamic parallelism of ensemble models. Specifically, on the hardware level, we propose a novel instruction set extension and develop efficient architecture components to fully exploit the new hardware abstraction of scattered idle computing cores, and use them to dynamically create on-the-fly Neural Processing Units (fNPUs). Moreover, we propose a computing power recycle mechanism to run on-the-fly models (small models) on fNPUs by carefully orchestrating execution order of ensemble models for maximizing hardware resources and bandwidth utilization. On the software level, EnsGuard adopts an integrated hardware/randomized ensemble co-design optimizer aiming at winning both faster inference and higher adversarial robustness. On top of that, a multi-model mapping method based on decision tree is proposed to enable the interleaving of different DNN executions both spatially and temporally, and mitigate straggler problems. Evaluation with a diverse set of workloads shows significant gains in throughput (4.4×) and energy reduction (3.2×).
Xingbin Wang, Yan Wang 0122, Yulan Su, Sisi Zhang, Dan Meng 0002, Rui Hou 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2