Yuhang Du

dblp:49/9819 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 ReCA: Integrated Acceleration for Real-Time and Efficient Cooperative Embodied Autonomous Agents
Zishen Wan, Yuhang Du, Mohamed Ibrahim 0002, Jiayi Qian, Jason Jabbour, Yang Zhao 0013, Tushar Krishna, Arijit Raychowdhury, Vijay Janapa Reddi
ASPLOS (2)2
2025 ToMamba: Towards Token-Efficient Mamba Architecture on FPGA
abstract
The State Space Model (SSM), particularly the Mamba implementation, has demonstrated impressive capabilities across various domains. It offers a significant reduction in computational complexity compared to Transformers while achieving higher algorithm accuracy. However, the ineffectiveness of spatially unfolding the SSM layer leads to increased latency as sentence length grows, especially when being deployed on FPGA. Previous token reduction methods introduced in Transformers fail to maintain high performance in Mamba. Moreover, the dispersed outliers, complex model structure and variety of non-linear operators obstruct its efficient implementation on FPGA. To address these challenges, we propose ToMamba, the first algorithm-architecture co-design to optimize Mamba implementation. At the algorithmic level, ToMamba incorporates a novel progressive token merging algorithm with minimal hardware consumption and a hardware-aware fine-grained quantization strategy. On the hardware side, a dualflow systolic array is designed to unify convolution and matrix multiplication, supporting both weight stationary and output stationary dataflow. A fine-grained pipeline design is adopted for SSM computation to maximize hardware efficiency and enhance throughput. Furthermore, efficient hardware architecture and approximation method for nonlinear function units are proposed. To enable merging after the Mamba layer, ToMamba also adopts a dedicated data mapping scheme. Comprehensive evaluations across multiple benchmarks demonstrate that the token reduction method of ToMamba achieves 10% sparsity with only 0.25% accuracy loss, improving up to 16.89% in accuracy compared to previous methods. ToMamba hardware implementation on U280 FPGA achieves up to 636.00×/11.01×/1.39× speedup compared to Intel Xeon Platinum 8369B CPU, NVIDIA Tesla A100 GPU and ASIC platforms and 1280×/44.32× energy efficiency improvement compared to CPU and GPU platforms.
Kejia Shi, Yuhang Du, Jianli Chen, Jun Yu 0010, Kun Wang 0005
ICCAD3
2025 Generative AI in Embodied Systems: System-Level Analysis of Performance, Efficiency and Scalability
abstract
Embodied systems, where generative autonomous agents engage with the physical world through integrated perception, cognition, action, and advanced reasoning powered by large language models (LLMs), hold immense potential for addressing complex, long-horizon, multi-objective tasks in realworld environments. However, deploying these systems remains challenging due to prolonged runtime latency, limited scalability, and heightened sensitivity, leading to significant system inefficiencies. In this paper, we aim to understand the workload characteristics of embodied agent systems and explore optimization solutions. We systematically categorize these systems into four paradigms and conduct benchmarking studies to evaluate their task performance and system efficiency across various modules, agent scales, and embodied tasks. Our benchmarking studies uncover critical challenges, such as prolonged planning and communication latency, redundant agent interactions, complex low-level control mechanisms, memory inconsistencies, exploding prompt lengths, sensitivity to self-correction and execution, sharp declines in success rates, and reduced collaboration efficiency as agent numbers increase. Leveraging these profiling insights, we suggest system optimization strategies to improve the performance, efficiency, and scalability of embodied agents across different paradigms. This paper presents the first system-level analysis of embodied AI agents, and explores opportunities for advancing future embodied system design.
Zishen Wan, Jiayi Qian, Yuhang Du, Jason Jabbour, Yilun Du, Yang Zhao 0013, Arijit Raychowdhury, Tushar Krishna, Vijay Janapa Reddi
ISPASS3
2024 Thinking and Moving: An Efficient Computing Approach for Integrated Task and Motion Planning in Cooperative Embodied AI Systems
abstract
Cooperative embodied AI systems, where multiple agents collaborate to accomplish complex, long-horizon tasks, show significant promise for real-world applications. These systems integrate perception, cognition, and action through integrated task and motion planning (TAMP), leveraging the advanced reasoning and communication capabilities of large language models (LLMs). However, their efficiency is often hindered by challenges such as high computational latency and redundant communication, largely due to the reliance on LLMs for sequential planning decisions.
Zishen Wan, Yuhang Du, Mohamed Ibrahim 0002, Yang Zhao 0013, Tushar Krishna, Arijit Raychowdhury
ICCAD2
2023 Probabilistic-Attention Fusion-Based Lithium-ion Battery Pack Multivariate Prediction Method
abstract
Lithium-ion battery packs are widely employed in various applications, such as electric vehicles, energy storage system solutions, and other practical uses. Inconsistency between battery cells is one of the main factors affecting the performance of the battery pack. However, existing studies have paid insufficient attention to the prediction of inconsistency evolution trends within a battery pack. They neglect the impact of local variables on the accuracy of the results in the prediction process, while the results lack the corresponding ability to express uncertainty. To address the aforementioned issues, this paper proposes a probabilistic-attention fusion-based method for multivariate prediction of lithium-ion battery pack performance. Firstly, it utilizes LSTM to identify correlations between parameters for multiparameter prediction. Then, an attention mechanism is introduced to enhance the model's focus on key information by assigning attention weights to the input data. Finally, probabilistic modeling is incorporated to provide the model with the capability to express uncertainties. The effectiveness of the proposed method is corroborated through rigorous testing using real-world battery data obtained from laboratory experiments.
Yuhang Du, Datong Liu, Yu Peng 0002
IECON1
2021 Dual Batch Size Training: An efficient MGD adaptive batch size method
abstract
Mini-batch Gradient Descent (MGD) has become a standard for deep learning model training. For a long period of time, the size of the mini-batch (also known as batch size) is set empirically as a fixed value, while recent works have demonstrated it could have crucial effects on the training. Although there already exist several adaptive batch size methods, they either add significant overhead to the training, or lack of robustness on various training scenarios. In this work, an adaptive batch size method to accelerate MGD training is proposed, whose basic idea is to concurrently run training with two batch sizes, and choose batch size based on the comparison of the evaluated history performance. It can be easily implemented on various deep learning platforms, and the experiment results suggest that the proposed method achieves high performance and strong robustness with acceptable and controllable overhead, which outperforms existing adaptive batch size methods.
Yuhang Du, Wenfeng Shen, Baohua Liu, Weijia Lu
ICTAI1
2019 DeepRec: A deep neural network approach to recommendation with item embedding and weighted loss function
Wen Zhang 0001, Yuhang Du, Taketoshi Yoshida
Inf. Sci.2
2018 DRI-RCNN: An approach to deceptive review identification using recurrent convolutional neural network
Wen Zhang 0001, Yuhang Du, Taketoshi Yoshida, Qing Wang 0001
Inf. Process. Manag.2