EDBT 2026 Demo / reviewers in the wild / expert
Wendong Xu
dblp:84/10613
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hardware-aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory ArchitectureabstractLow-rank adaptation (LoRA) is a predominant parameter-efficient finetuning method for adapting large language models (LLMs) to downstream tasks. Meanwhile, Compute-in-Memory (CIM) architectures demonstrate superior energy efficiency due to their array-level parallel in-memory computing designs. In this article, we propose deploying the LoRA-finetuned LLMs on the hybrid CIM architecture (i.e., pretrained weights onto energy-efficient Resistive Random-Access Memory (RRAM) and LoRA branches onto noise-free Static Random-Access Memory (SRAM)), reducing the energy cost to about 3% compared with the Nvidia A100 GPU. However, the inherent noise of RRAM on the saved weights leads to performance degradation, simultaneously. To address this issue, we design a novel Hardware-aware Low-rank Adaptation (HaLoRA) method. The key insight is to train a LoRA branch that is robust toward such noise and then deploy it on noise-free SRAM, while the extra cost is negligible since the parameters of LoRAs are much fewer than pretrained weights (e.g., 0.15% for LLaMA-3.2 1B model). To improve the robustness towards the noise, we theoretically analyze the gap between the optimization trajectories of the LoRA branch under both ideal and noisy conditions and further design an extra loss to minimize the upper bound of this gap. Therefore, we can enjoy both energy efficiency and accuracy during inference. Experiments finetuning the Qwen and LLaMA series demonstrate the effectiveness of HaLoRA across multiple reasoning tasks, achieving up to 22.7 improvement in average score while maintaining robustness at various noise types and noise levels. Taiqiang Wu, Chenchen Ding, Wenyong Zhou, Yuxin Cheng, Xincheng Feng, Wendong Xu, Chufan Shi, Zhengwu Liu, Ngai Wong 0001 |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2026 | PPD: A Portable and Highly Parallel Dispatching System for Deep Learning
Wendong Xu, Yuhao Ji, Yueting Li 0001, Yuxuan Zhao 0001, Zhengwu Liu, Bei Yu 0001, Ngai Wong 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2025 | A Custom RISC-V ISA with Scalable Processing Units for Efficient Neural Network InferenceabstractA customized RISC-V ISA with integrated digital accelerators offers a promising solution to improve energy efficiency in neural network inference.However, it often requires multiple instructions per accelerator operation, which limits computational efficiency during deep neural network inference.To overcome the instruction overhead, this design introduces a dedicated instruction set that enables scalable and fine-grained accelerator control.By incorporating the pattern-driven instruction mode, this design exploits the neural layer regularity to support efficient instruction iteration.Furthermore, this digital accelerator leverages hardware reuse for logic operations, forming a fusion-style architecture that integrates reconfigurable components.Experimental results demonstrate that the custom RISC-V ISA achieves an average runtime speedup of 8.26× and reduces the instruction count by 14.71×.This design also yields an average 8.73× reduction in cycles per instruction across MobileNetV2, ResNet50, VGG19, EfficientNet, and DenseNet-BC, validating its effectiveness across representative benchmarks.Additionally, it improves average energy efficiency by 1.74×, outperforming state-of-the-art designs. Yueting Li 0001, Wanshuang Lin, Wendong Xu, Ngai Wong 0001, Weisheng Zhao 0001 |
CF | 3 |
| 2025 | TransGER: Transformer-Based CNN-BiGRU Architecture for sEMG Gesture Recognition in Time-Frequency Domain
Yuhan Yuan, Anming Dong, Wendong Xu, Yubing Han, Jiguo Yu, You Zhou 0006 |
WASA (3) | 3 |
| 2025 | A Learned Performance Model With Transfer Learning Across GPUs on Tensorized InstructionsabstractThe training and inference efficiency of ever-larger deep neural networks highly rely on the performance of tensor operators on specific hardware accelerators. Therefore, a performance tuning framework with tensorized instruction compilation for automatic tensor generation is necessary for efficient deployment. These novel tensorized instruction, along with the emerging machine learning models, bring tremendous engineering challenges in compilation-based methods. They suffer from a large design space exploration with rough measurement accuracy and poor transferability among specialized instructions with certain hardware constraints. This paper presents a novel performance model for automatic code optimization with tensorized instruction. Central to the performance model is the assignment feature that not only clearly specifies the behaviour of instruction with computation and data movement abstraction, but also formally defines the matching problem from algorithm to tensorized instructions. Meanwhile, a simple yet efficient design with attention-inspired modules to accurately predict the performance of optimized tensor program by capturing global and long-range dependencies within a complete scheduling space. Compared with state-of-the-arts, our performance model can predict the optimal implementation of code configurations with tensorized instruction to reduce inference latency and search time by up to 1.21× and 3.41× on modern DNN benchmarks. Furthermore, with pre-trained parameters, our performance can quickly adapt to different workloads and platforms on tensorized instruction via transfer learning. Wendong Xu, Bei Yu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |