Wendong Xu

dblp:84/10613 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture
abstract
Low-rank adaptation (LoRA) is a predominant parameter-efficient finetuning method for adapting large language models (LLMs) to downstream tasks. Meanwhile, Compute-in-Memory (CIM) architectures demonstrate superior energy efficiency due to their array-level parallel in-memory computing designs. In this article, we propose deploying the LoRA-finetuned LLMs on the hybrid CIM architecture (i.e., pretrained weights onto energy-efficient Resistive Random-Access Memory (RRAM) and LoRA branches onto noise-free Static Random-Access Memory (SRAM)), reducing the energy cost to about 3% compared with the Nvidia A100 GPU. However, the inherent noise of RRAM on the saved weights leads to performance degradation, simultaneously. To address this issue, we design a novel Hardware-aware Low-rank Adaptation (HaLoRA) method. The key insight is to train a LoRA branch that is robust toward such noise and then deploy it on noise-free SRAM, while the extra cost is negligible since the parameters of LoRAs are much fewer than pretrained weights (e.g., 0.15% for LLaMA-3.2 1B model). To improve the robustness towards the noise, we theoretically analyze the gap between the optimization trajectories of the LoRA branch under both ideal and noisy conditions and further design an extra loss to minimize the upper bound of this gap. Therefore, we can enjoy both energy efficiency and accuracy during inference. Experiments finetuning the Qwen and LLaMA series demonstrate the effectiveness of HaLoRA across multiple reasoning tasks, achieving up to 22.7 improvement in average score while maintaining robustness at various noise types and noise levels.
Taiqiang Wu, Chenchen Ding, Wenyong Zhou, Yuxin Cheng, Xincheng Feng, Wendong Xu, Chufan Shi, Zhengwu Liu, Ngai Wong 0001
ACM Trans. Design Autom. Electr. Syst.7
2026 PPD: A Portable and Highly Parallel Dispatching System for Deep Learning
Wendong Xu, Yuhao Ji, Yueting Li 0001, Yuxuan Zhao 0001, Zhengwu Liu, Bei Yu 0001, Ngai Wong 0001
ACM Trans. Design Autom. Electr. Syst.1
2025 A Custom RISC-V ISA with Scalable Processing Units for Efficient Neural Network Inference
abstract
A customized RISC-V ISA with integrated digital accelerators offers a promising solution to improve energy efficiency in neural network inference.However, it often requires multiple instructions per accelerator operation, which limits computational efficiency during deep neural network inference.To overcome the instruction overhead, this design introduces a dedicated instruction set that enables scalable and fine-grained accelerator control.By incorporating the pattern-driven instruction mode, this design exploits the neural layer regularity to support efficient instruction iteration.Furthermore, this digital accelerator leverages hardware reuse for logic operations, forming a fusion-style architecture that integrates reconfigurable components.Experimental results demonstrate that the custom RISC-V ISA achieves an average runtime speedup of 8.26× and reduces the instruction count by 14.71×.This design also yields an average 8.73× reduction in cycles per instruction across MobileNetV2, ResNet50, VGG19, EfficientNet, and DenseNet-BC, validating its effectiveness across representative benchmarks.Additionally, it improves average energy efficiency by 1.74×, outperforming state-of-the-art designs.
Yueting Li 0001, Wanshuang Lin, Wendong Xu, Ngai Wong 0001, Weisheng Zhao 0001
CF3
2025 TransGER: Transformer-Based CNN-BiGRU Architecture for sEMG Gesture Recognition in Time-Frequency Domain
Yuhan Yuan, Anming Dong, Wendong Xu, Yubing Han, Jiguo Yu, You Zhou 0006
WASA (3)3
2025 A Learned Performance Model With Transfer Learning Across GPUs on Tensorized Instructions
abstract
The training and inference efficiency of ever-larger deep neural networks highly rely on the performance of tensor operators on specific hardware accelerators. Therefore, a performance tuning framework with tensorized instruction compilation for automatic tensor generation is necessary for efficient deployment. These novel tensorized instruction, along with the emerging machine learning models, bring tremendous engineering challenges in compilation-based methods. They suffer from a large design space exploration with rough measurement accuracy and poor transferability among specialized instructions with certain hardware constraints. This paper presents a novel performance model for automatic code optimization with tensorized instruction. Central to the performance model is the assignment feature that not only clearly specifies the behaviour of instruction with computation and data movement abstraction, but also formally defines the matching problem from algorithm to tensorized instructions. Meanwhile, a simple yet efficient design with attention-inspired modules to accurately predict the performance of optimized tensor program by capturing global and long-range dependencies within a complete scheduling space. Compared with state-of-the-arts, our performance model can predict the optimal implementation of code configurations with tensorized instruction to reduce inference latency and search time by up to 1.21× and 3.41× on modern DNN benchmarks. Furthermore, with pre-trained parameters, our performance can quickly adapt to different workloads and platforms on tensorized instruction via transfer learning.
Wendong Xu, Bei Yu 0001
IEEE Trans. Parallel Distributed Syst.3