Huadong Dai

dblp:61/3869 · DBLP profile ↗
← Back
26ranked-venue papers
0as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 9 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Self-Attention Proximal Policy Optimization for Beam Tracking in mmWave Communications
abstract
Millimeter-wave (mmWave) communication systems are highly sensitive to user mobility and environmental changes, often suffering from rapid beam direction shifts and frequent link blockages. These dynamics pose significant challenges to maintaining reliable beam tracking. To enable accurate and robust beam tracking in dynamic environments, we propose a Self-Attention Proximal Policy Optimization (SAPPO) algorithm. The beam tracking task is modeled as a Partially Observable Markov Decision Process (POMDP), where the agent aims to maximize Received Signal Strength (RSS) using only current and previous signal observations, without relying on channel models or complex signal processing. By integrating a multi-head self-attention (MHSA) mechanism with a Gated Recurrent Unit (GRU), the policy network effectively captures temporal dependencies, enhancing the agent’s perception of user motion and channel state variations. To validate the approach, we construct three test scenarios with varying mobility and blockage complexity using the DeepMIMO dataset. Experimental results demonstrate that SAPPO consistently outperforms baseline methods such as DQN and PPO in terms of tracking accuracy, stability, and robustness. Notably, in challenging environments with frequent Line-Of-Sight (LOS)/Non-Line-Of-Sight (NLOS) transitions, SAPPO achieves a beam tracking success rate exceeding 95%, highlighting its strong adaptability and reliable performance under dynamic conditions.
Jie Tan 0002, Xiaoguang Ren, Huadong Dai
IEEE J. Sel. Areas Commun.4
2026 Semantic-Oriented Image Transmission and Resource Allocation for UAV Networks
Jianchao Zheng, Weilu Wang, Xiancai Yao, Huadong Dai, Jinshu Su
IEEE Trans. Commun.7
2025 AIRES: A General Framework for Efficient Intrinsic Rewards Based on Attention Mechanisms
abstract
Efficient exploration in high-dimensional observation spaces remains a critical challenge in deep reinforcement learning, particularly in scenarios with sparse extrinsic rewards. A promising approach is to encourage exploration by estimating intrinsic rewards based on the novelty of observations. However, there is a gap between the observed novelty and the actual effectiveness of exploration, as both environmental stochasticity and the agent’s actions may influence observations. To accurately evaluate the novelty contributed by agent exploration in intrinsic rewards, we propose the AIRES (Attention-driven Intrinsic Reward for Exploration Strategy) framework. AIRES leverages the attention mechanisms to analyze the relationship within trajectory sequences generated by agent-environment interactions, employing attention weights to quantify the relevance of observations to actions. By applying attention weights to intrinsic rewards, the novelty brought by agent exploration is enhanced and the impact of environmental stochasticity is reduced. Extensive experiments demonstrate that AIRES significantly enhances the performance of prominent intrinsic reward methods, establishing it as a robust and scalable solution for efficient exploration.
Guoli Wu, Xiaoguang Ren, Huadong Dai
ECAI7
2025 LintLLM: An Open-Source Verilog Linting Framework Based on Large Language Models
Zhigang Fang 0002, Renzhi Chen, Yang Guo 0003, Huadong Dai, Lei Wang 0011
ACM Great Lakes Symposium on VLSI5
2025 NearFetch: Saving Inter-Module Bandwidth in Many-Chip-Module GPUs
abstract
As Graphics Processing Units (GPUs) face increasing computing demands that surpass single-module capabilities due to transistor scaling and lithography constraints, the necessity for expanding the module count within GPUs grows. This escalation faces a significant challenge: the total inter-module bandwidth in many-chip-module GPUs is limited by manufacturing constraints in organic substrates or silicon interposers. Unlike Central Processing Units (CPUs), which are latency-sensitive, GPUs leverage their high thread-level parallelism to effectively hide memory access latency through simultaneous multithreading. This attribute makes GPUs inherently sensitive to bandwidth constraints, making the efficient exploitation of available inter-module bandwidth important. In this paper, we identify that fetching data from faraway memory in many-chip-module GPUs can easily cause bandwidth contention which degrades the real achieved data bandwidth per GPU module compared to fetching data from nearby memory. To further analyze this problem, we introduce the Inter-Module Bandwidth per Access (IBPA) metric for quantifying bandwidth usage and finding the network hop count directly impacts the IBPA and network contention. Next, we propose NearFetch, a routing-based solution to reduce IBPA. NearFetch works due to the fact that GPU modules along the routing path are typically much closer to the source GPU module while these GPU modules can supply $29.1 \%$ of data for high-sharing applications. NearFetch consists of two primary components: a data forwarding scheme, enabling data forwarding when the data resides in a remote GPU module, and a topology-aware Miss Status Handling Register (MSHR) coalescing scheme, responsible for recording the memory address information for future use in case of a data miss. By leveraging the data locality among various GPU modules, NearFetch substantially minimizes inter-module bandwidth usage, eliminating the need to fetch data from distant memory partitions. Our evaluation of NearFetch within the context of many-chip-module GPUs, across applications exhibiting diverse degrees of data locality, reveals that it reduces IBPA by $4 2. 6 \%$ and enhances performance by an average of $52.2 \%$ (with up to $9 8. 1 \%$ improvement) for high-sharing workloads.
Guangda Zhang, Shiqing Zhang, Huadong Dai
HPCA5
2025 RTLBench: A Multi-Dimensional Benchmark Suite for Evaluating LLM-Generated RTL Code
abstract
The rapid advancement of large language models (LLMs) has enabled automated Register Transfer Level (RTL) code generation, accelerating chip design workflows. However, existing benchmarks focus mainly on syntax and functionality, overlooking critical engineering aspects such as lint compliance, readability, and coding style. To address this gap, we propose RTLBench, a benchmark suite of 160 copyright-free RTL cases sourced from textbooks and open-source projects. RTLBench features a multi-dimensional evaluation framework covering syntax, functionality, lint compliance, readability, and style consistency. To assess subjective code quality metrics, it also incorporates an LLM-as-a-judge mechanism. We evaluated 24 state-of-the-art LLMs using RTLBench, finding that while several models perform well in syntax and functionality, most fall short on engineering quality. To address this, we propose Log2BetterRTL, a log-driven feedback system that transforms EDA tool diagnostics into iterative improvement prompts. It improves syntax correctness by up to 18.13 %, boosts functional correctness by 14.38 %, reduces lint violations by up to 229, and raises clarity scores by 0.51. These results demonstrate RTLBench's effectiveness in evaluating and enhancing LLMgenerated RTL, bridging the gap between generative AI and industrial-grade hardware design. The suite and scripts are available at: https://fangzhigang32.github.io/RTLBench.
Zhigang Fang 0002, Renzhi Chen, Yang Guo 0003, Huadong Dai, Lei Wang 0011
ICCD4
2025 UGPU: Dynamically Constructing Unbalanced GPUs for Enhanced Resource Efficiency
abstract
Different GPU generations have various numbers of SMs but still keep the balanced idea during the manufacture, i.e., the proportion of compute and memory resources within a single physical GPU is similar.Although GPU applications have different characteristics, it is still uncommon and uneconomic to build unbalanced physical GPUs for customers.With their powerful computational capabilities, GPUs are widely used in the cloud to accelerate diverse workloads from multiple users, creating opportunities to explore the unbalanced GPU concept in multitasking environments.In this paper, we take the first step in exploring the feasibility and performance benefits of building unbalanced GPUs.Specifically, these unbalanced GPUs, referred to as GPU slices, are dynamically constructed with dedicated compute and memory resources from a single physical GPU to effectively address the diverse demands of co-executing applications, achieving high performance during the execution.However, there are two challenges that must to be solved.First, determining the size of unbalanced GPU slices during execution is challenging, as predicting GPU performance under varying resource allocations is inherently difficult.Second, reallocating memory resources after partitioning requires extensive data migration, with traditional methods leading to unacceptable performance degradation.To address the first challenge, UGPU employs a demand-aware resource partitioning algorithm that partitions resources dynamically without relying on a complex or inaccurate performance model.For the second challenge, UGPU introduces PageMove, a novel mechanism for efficient page migration between different memory dies within an HBM stack.Our key insight is that all memory channels already have physical connections to all through-silicon via (TSV) within a DRAM stack, while different bank groups can transfer data at the same time.PageMove slightly modifies DRAM architecture, uses a customized memory address mapping, designs a new parallel page migration mode (PPMM) and updates the virtual memory management scheme.By doing this,
Xia Zhao 0004, Guangda Zhang, Lu Wang 0019, Huadong Dai
ISCA4
2024 LLM-Based Processor Verification: A Case Study for Neuronnorphic Processor
abstract
With the increasing complexity of the hardware design, conducting verification before the tapeout is of utmost importance. Simulation-based verification remains the primary method owing to its scalability and flexibility. A comprehensive verification of modern processors usually requires numerous effective tests to cover all possible conditions and use cases, leading to significant time, resource, and manual effort even with the EDA. Moreover, novel domain specific architecture (DSA), such as neuromorphic processors, will exacerbate the challenge of verification. Fortunately, emerging large language models (LLMs) have been demonstrating a powerful ability to complete specific tasks assigned by human instructions. In this paper, we explore the challenges and opportunities encountered when using the LLMs to accelerate the DSA verification using the proposed LLM-based workflow consisting of test generation, compilation&simulation, and result collection&processing. By verifying a RISC-V core and a neuromorphic processor, we examine the capabilities and limitations of the LLMs when using them for the function verification of traditional processors and emerging DSA. In the experiment, 36$C$programs and 128 assembly snippets for the RISC-V core and the neuromorphic processor are generated using an advanced LLM to demonstrate our claim. The experimental results show that the code coverage based on the LLM test generation can reach 89% and 91% for the above two architectures respectively, showing a promising research direction for the future processor verification in the new golden age for computer architecture.
Yifei Deng, Renzhi Chen, Jingyue Zhao, Huadong Dai, Yuhua Tang
DATE7
2024 LLM - TG: Towards Automated Test Case Generation for Processors Using Large Language Models
abstract
Design verification (DV) has existed for decades and is crucial for identifying potential bugs before chip tape- out. Hand-crafting test cases is time-consuming and error-prone, even for experienced verification engineers. Prior work has attempted to lighten this burden by rule-guided random test case generation. However, this approach does not eliminate the manual effort required to write rules that describe detailed hardware behavior. Motivated by advances in large language models (LLMs), we explore their potential to capture register transfer level (RTL) behavior and construct prompts for test case generation based on RTL behavior. First, we introduce a prompt framework, LLM - Driven Test Generation (LLM - TG), to generate test cases, thereby enhancing LLMs' test generation capabilities. Additionally, we provide an open-source prompt library that offers a set of standardized prompts for processor verification, aiming to improve test generation efficiency. Lastly, we use an LLM to verify a 12-stage, multi-issue, out-of-order RV64GC processor, achieving at least an 8.34 % increase in block coverage and at least a 5.8 % increase in expression coverage compared to the state-of-the-art (SOTA) methods, LLM4DV and RISCV- DV. The prompt library is available at https://github.com/LLM-TGIPrompt_Library.
Yifei Deng, Renzhi Chen, Yuanfeng Luo, Jingyue Zhao, Zhong Wan, Yongbao Ai, Huadong Dai
ICCD10
2024 MOTPE/D: Hardware and Algorithm Co-design for Reconfigurable Neuromorphic Processor
abstract
Recent advances in hardware/algorithm co-design for spiking neural networks have demonstrated its potential for jointly optimizing algorithmic performance while minimizing hardware overhead. However, the gigantic mixed-variable hard-ware/algorithm co-design space and time-consuming hardware verification still pose an intractable challenge for solutions exploration. To tackle these problems, 1) we propose a generic three-phase hardware/algorithm co-design framework. In this framework, 2) we target a reconfigurable neuromorphic processor, and parameterize the hardware and network architecture in a unified design space. 3) We propose a generic analytical model to estimate the parameter size and power consumption, which can support fast candidate evaluation during the exploration. 4) We extend vanilla TPE (a single-objective optimization algorithm) to MOTPE/D, a generic Multi-objective optimization (MOO) algorithm, by introducing a decomposition strategy.
Renzhi Chen, Xun Xiao, Jingyue Zhao, Zhenhua Zhu 0002, Huadong Dai, Yuhua Tang
ICCD7
2024 PLRUT: Pseudo Label and Re-detection Boosted Unsupervised Tracking of Unmanned Aerial Vehicle Objects
Jun Wang 0041, Huadong Dai, Bo Zhang 0007, Shan Qin, Jian Zhao 0006
PRCV (12)2
2024 A Fast and Safe Neuromorphic Approach for Obstacle Avoidance of Unmanned Aerial Vehicle
abstract
Obstacle avoidance is a crucial task in unmanned aerial vehicles (UAV) motion planning. The accuracy and consistency of real-time visual information affect the gener-ation of obstacle avoidance commands, raising higher safety demands for obstacle avoidance. The neuromorphic computing-based obstacle avoidance solution can address these challenges. Dynamic vision sensors (DVS) exhibit low latency, low power consumption, and high dynamic range as novel neuromorphic sensors. Spiking neural networks (SNN) also leverage the same mechanism to efficiently process asynchronous and sparse event data generated by DVS, offering latency and energy efficiency advantages. Additionally, the optimal estimation method effectively mitigates the impact of noise and interference within the system, reducing the influence of errors on the algorithm and enhancing safety. Based on these considerations, this paper proposes a fast and safe obstacle avoidance framework. DVS is used to acquire event data from the environment, and a hardware-compatible lightweight SNN is employed to extract dynamic obstacle position information from the data. Compared to baseline methods, this approach reduces latency by 85%. Furthermore, two estimation methods are used to predict the movement of obstacles, ensuring flight safety by generating different UAV obstacle avoidance actions based on confidence intervals, even in the presence of obstacle information errors and omissions.
Zhong Wan, Xun Xiao, Jingyue Zhao, Junbo Tie, Renzhi Chen, Guangda Zhang, Huadong Dai
SMC10
2023 KURL: A Knowledge-Guided Reinforcement Learning Model for Active Object Tracking
Jie Tan 0002, Xiaoguang Ren, Weiya Ren, Huadong Dai
ACML5
2023 Air-to-Ground Active Object Tracking via Reinforcement Learning
Weiya Ren, Jie Tan 0002, Xiaochuan Zhang, Xiaoguang Ren, Huadong Dai
ICANN (6)6
2023 Joint Optimization of Trajectory and Image Transmission in Multi-UAV Semantic Communication Networks
abstract
Semantic communication is considered the key promoter and basic paradigm of future 6G networks and applications. In this paper, we investigate a multi-unmanned aerial vehicle (UAV) semantic communication framework, where the trajectories and communication services are jointly optimized for image transmission of ground users. We aim to minimize the transmission delay while considering constraints such as the UAV’s energy threshold, collision avoidance, and the bandwidth of the multi-UAV system. We propose a value decomposition based multi-agent deep reinforcement learning (VD-MADRL) algorithm to solve this problem, which explores the joint optimization scheme of UAV trajectory, channel allocation, and semantic information selection. Simulations demonstrate that the proposed algorithm greatly reduces system energy consumption and transmission delay compared to other traditional UAVs’ trajectory planning algorithms.
Xiancai Yao, Jianchao Zheng, Huadong Dai
ICPADS4
2022 Efficient maintenance for maximal bicliques in bipartite graph streams
Ziyi Ma, Yikun Hu 0001, Jianye Yang 0001, Chubo Liu, Huadong Dai
World Wide Web6
2021 Automated Software Vulnerability Detection via Pre-trained Context Encoder and Self Attention
Guang Kou, Huadong Dai
ICDF2C5
2021 micROS.BT: An Event-Driven Behavior Tree Framework for Swarm Robots
abstract
In this paper, we propose micROS.BT, an event-driven behavior tree (BT) framework aiming at supporting swarm-robot coordination. Compared with other BT frame-works, micROS.BT implements the event-driven way under the multi-thread mode, which can effectively save computing resources. Moreover, in order to ensure swarm-robot coordination, we optimize the implementation of the traditional blackboard and propose the multi-mode blackboard, which supports inner-tree, inter-tree, and inter-robot data sharing. Furthermore, considering the limited modularity of a single tree, micROS.BT realizes a mechanism called hierarchical tree management which involves inter-tree notifying and waiting functionalities, while ensuring that each tree is independent and self-scheduled. The effectiveness of micROS.BT is verified by simulation and real-robot experiments for different system settings, showing that a substantial improvement is achieved in comparison with the traditional BT implementations.
Yunlong Wu 0002, Huadong Dai, Xiaodong Yi 0002, Xuejun Yang
IROS3
2020 Attentional Fused Temporal Transformation Network for Video Action Recognition
abstract
Effective spatiotemporal feature representation is crucial to the video-based action recognition task. Focusing on discriminate spatiotemporal feature learning, we propose Attentional Fused Temporal Transformation Network (AttnTTN) for action recognition on top of popular Temporal Segment Network (TSN) framework. In the network, Attentional Fusion Module (AttnFM) is designed to fuse the appearance and motion features at multiple ConvNet levels for each video snippet, forming a short-term video descriptor. With fused features as inputs, Temporal Transformation Networks (TTN) are employed to model middle-term temporal transformation between the neighboring temporal snippets following a sequential order. AttnTTN achieves the state-of-the-art results on two most popular action recognition datasets: UCF101 and HMDB51.
Ke Yang 0004, Huadong Dai, Tianlong Shen, Peng Qiao, Xin Niu 0002, Dongsheng Li 0001, Yong Dou
ICASSP3
2020 An Actor-based Programming Framework for Swarm Robotic Systems
abstract
Programming cooperative tasks for autonomous swarm robotic systems has always been challenging. In this paper, we introduce a concept `Actor', as a virtualization for robot platforms. Every robot platform in the swarm robotic system carries out the task and interacts with others as an Actor. We designed an Actor-based framework for the management of autonomous swarm robotic systems including modules and interfaces for the Actor, the collective Actor, and task management. The Actor-based framework enables task developers to explicitly model cooperative tasks without intricacies about the detailed robotic algorithms or the specific robot brands, and eases the burden on robotic algorithm developers by providing common functionalities. The proposed framework is implemented in C++ and validated quantitatively and qualitatively with a swarm of thirty drones by simulations and a swarm of ten drones by in-field tests.
Bin Di, Ruihao Li 0001, Huadong Dai, Xiaodong Yi 0002, Xuejun Yang
IROS4
2018 Enhanced Visual Loop Closing for Laser-Based SLAM
abstract
Three-dimensional (3D) laser-based simultaneous localization and mapping (SLAM) can provide real-time pose information and construct accurate 3D map. However, detecting loop closures is a challenging task in the 3D laser-based SLAM because of the heavy computational overheads. In this paper, we propose a visual method to simultaneously detect and correct loop closures in the 3D laser-based SLAM based on prior work. In particular, we improve the experiments and evaluate our method by analyzing computational errors. The experimental results on the KITTI dataset prove that our method can efficiently reduce motion accumulation errors and successfully ensure the consistency performance of loop closure correction.
Zulun Zhu, Shaowu Yang, Huadong Dai
ASAP3
2018 Loop Detection and Correction of 3D Laser-Based SLAM with Visual Information
abstract
Three-dimensional (3D) laser-based simultaneous localization and mapping (SLAM) can provide real-time pose information and construct accurate 3D map. However, detecting loop closures is a challenge in the 3D laser-based SLAM for expensive computation of algorithms. In this paper, we propose a visual method to detect and correct loop closures. We introduce visual bags-of-words techniques for loop closure detection in the 3D laser-based SLAM. Time of computing similarities between points clouds can be saved. Our method maintains visual keyframes, each of which associates with its pose and segmentation of laser point clouds. Our experiments on KITTI dataset prove that our method can efficiently reduce motion accumulation errors and successfully ensure the real-time performance of loop closure correction.
Zulun Zhu, Shaowu Yang, Huadong Dai
CASA3
2015 Complementary Synthesis for Encoder with Flow Control Mechanism
abstract
Complementary synthesis automatically generates an encoder's decoder with the assumption that the encoder's all input variables can always be uniquely determined by its output symbol sequence. However, to prevent the faster encoder from overwhelming the slower decoder, many encoders employ flow control mechanism that fails this assumption. Such encoders, when their output symbol sequences are too fast to be processed by the decoders, will stop transmitting data symbols, but instead transmitting idle symbols that can only uniquely determine a subset of the encoder's input variables. And the decoder should recognize and discard these idle symbols. This mechanism fails the assumption of all complementary synthesis algorithms, because some input variables can't be uniquely determined by the idle symbol. A novel algorithm is proposed to handle such encoders. First, it identifies all input variables that can be uniquely determined, and takes them as flow control variables. Second, it infers a predicate over these flow control variables that enables all other input variables to be uniquely determined. Third, it characterizes the decoder's Boolean function with Craig interpolant. Experimental results on several complex encoders indicate that this algorithm can always correctly identify the flow control variables, infer the predicates and generate the decoder's Boolean functions.
ShengYu Shen, Qingbo Wu 0003, Huadong Dai, Yan Jia 0001
ACM Trans. Design Autom. Electr. Syst.4
2013 Residency-Aware Virtual Machine Communication Optimization: Design Choices and Techniques
abstract
Network I/O workloads are dominating in many data centers and cloud computing environments today. One way to improve inter Virtual Machine (VM) communication efficiency is to support co-resident VM communication by using shared memory based approaches and to resort to the traditional TCP/IP for inter-VM communications between VMs that are located on different physical hosts. Although a number of independent efforts are dedicated to improving communication efficiency between co-resident VMs, they differ from one another in terms of how the inter-VM communication optimization is carried out and where in the software stack the shared memory channel is established. In this paper, we provide an in-depth overview of the design choices and techniques for optimizing the performance of the co-resident inter-VM communication, with dual objectives. First, we describe the core design guidelines and key issues for optimizing inter-VM communication by using shared memory based mechanisms. Typical issues include choices of implementation layer in the software stack, seamless agility for VM live migration and VM dynamic deployment support, multilevel transparency. Second, we conduct a comprehensive analysis of representative state-of-the-art research efforts and implementation techniques based on the core design guidelines. We also give an analysis of future requirements in advanced features such as reliability, security and stability. The research reported in this paper not only provides the reference for developing the next generation of inter-VM communication optimization mechanisms, but also offers opportunities for both cloud infrastructure providers and cloud service consumers to improve inter-VM communication efficiency in virtualized platforms.
Yi Ren 0008, Ling Liu 0001, Qi Zhang 0009, Qingbo Wu 0003, Jinzhu Kong, Jianbo Guan, Huadong Dai
IEEE CLOUD8
2013 Scratchpad memory allocation for arrays in permutation graphs
Li Wang 0027, Xuejun Yang, Huadong Dai
Sci. China Inf. Sci.3
2012 A fast and transparent communication protocol for co-resident virtual machines
abstract
Network I/O workloads are dominating in most of the Cloud datacenters today. One way to improve inter-VM communication efficiency is to support co-resident VM communication using a faster communication protocol than the traditional TCP/IP commonly used regardless whether VMs are located on the same
Yi Ren 0008, Ling Liu 0001, Xiaojian Liu 0005, Jinzhu Kong, Huadong Dai, Qingbo Wu 0003, Yuan Li 0011
CollaborateCom5