EDBT 2026 Demo / reviewers in the wild / expert
Changwu Zhang
dblp:246/7476
· DBLP profile ↗
4ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-2918-8647ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FAMS: A FrAmework of Memory-Centric Mapping for DNNs on Systolic Array AcceleratorsabstractIn recent years, deep neural networks (DNNs) have experienced rapid development. These DNNs demonstrate significant variations in architecture and scale, creating a substantial demand for domain-specific accelerators that are optimized for both high performance and low energy consumption. Systolic array accelerators, due to their efficient dataflow and parallel processing capabilities, offer significant advantages when performing computations for DNNs. Existing studies frequently overlook various hardware constraints in systolic array accelerators when representing mapping strategies. This oversight includes ignoring the differences in delays between communication and computation operations, as well as overlooking the capacities of multilevel memory hierarchies. Such omissions can lead to inaccuracies in predicting accelerator performance and inefficiencies in system design. We propose the FAMS framework, which introduces a memory-centric notation capable of fully representing the mapping of DNN operations on systolic array accelerators. Memory-centric notation moves away from the idealized assumptions of previous notations and considers various hardware constraints, thereby expanding the effective design and mapping spaces. The FAMS framework also includes a cycle-accurate simulator, which takes the hardware configurations, task descriptions, and mapping strategy represented by memory-centric notation as inputs, providing various metrics such as latency and energy consumption. The experimental results demonstrate that our proposed FAMS framework reduces latency by up to 29.7% and increases throughput by 42.4% compared to the state-of-the-art TENET framework. Additionally, under hardware configurations with a MAC delay of 2 and 3 clock cycles, the FAMS framework enhances performance by 12.0% and 25.4%, respectively. Hao Sun 0023, Junzhong Shen, Zhongyi Tang, Changwu Zhang, Yang Shi 0008, Hengzhu Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | Rocket Landing Control with Random Annealing Jump Start Reinforcement LearningabstractRocket recycling is a crucial pursuit in aerospace technology, aimed at reducing costs and environmental impact in space exploration. The primary focus centers on rocket landing control, involving the guidance of a nonlinear under-actuated rocket with limited fuel in real-time. This challenging task prompts the application of reinforcement learning (RL), yet goal-oriented nature of the problem poses difficulties for standard RL algorithms due to the absence of intermediate reward signals. This paper, for the first time, significantly elevates the success rate of rocket landing control from 8% with a baseline controller to 97% on a high-fidelity rocket model using RL. Our approach, called Random Annealing Jump Start (RAJS), is tailored for real-world goal-oriented problems by leveraging prior feedback controllers as guide policy to facilitate environmental exploration and policy learning in RL. In each episode, the guide policy navigates the environment for the guide horizon, followed by the exploration policy taking charge to complete remaining steps. This jump-start strategy prunes exploration space, rendering the problem more tractable to RL algorithms. The guide horizon is sampled from a uniform distribution, with its upper bound annealing to zero based on performance metrics, mitigating distribution shift and mismatch issues in existing methods. Additional enhancements, including cascading jump start, refined reward and terminal condition, and action smoothness regulation, further improve policy performance and practical applicability. The proposed method is validated through extensive evaluation and Hardware-in-the-Loop testing, affirming the effectiveness, real-time feasibility, and smoothness of the proposed controller. Yuxuan Jiang 0011, Zhiqian Lan, Guojian Zhan, Shengbo Eben Li, Qi Sun 0004, Tianwen Yu, Changwu Zhang |
IROS | 9 |
| 2023 | A Survey of Memory-Centric Energy Efficient Computer ArchitectureabstractEnergy efficient architecture is essential to improve both the performance and power consumption of a computer system. However, modern computers suffer from the severe “memory wall” problem due to the significant performance gap between the processor technology and the memory technology. Thus, the computer architecture community is evolving from compute-centric to memory-centric designs to reduce the data movement overhead. This paper presents a comprehensive survey of the main challenges and recent advances in memory-centric energy efficient computer architecture. We summarize two research directions: improving the memory technology and processing closer to memory. The former focuses on optimizing the conventional memory technology and exploiting emerging non-volatile memory (NVM) technology. The latter talks about currently popular processing in memory (PIM) technology, including near-memory processing (NMP) and in-memory processing (IMP). Moreover, some other topics like hardware for machine learning (ML), ML for hardware, security, privacy, and reliability are gaining increasing attention and should be considered seriously in the design phase of a computer system. The community is facing various challenges and opportunities simultaneously, requiring researchers to have a more comprehensive understanding of this field which is also the goal of this paper. Changwu Zhang, Hao Sun 0023, Shuman Li, Hengzhu Liu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | Closer to Optimal Angle-Constrained Path PlanningabstractPlanning on grids and planning via sampling are the two classical mainstreams of path planning for intelligent agents, whose respective representatives are A* and RRT, including their variants, Theta* and RRT*. However, in the nonholonomic path planning, such us being under angle constraints, Theta* and Lazy Theta* may fail to generate a feasible path because the line-of-sight check (LoS-Check) will modify the original orientation of a state, which makes the planning process incomplete (cannot visit all possible states). Then, we propose a more delayed evaluation algorithm called Late LoS-Check A* (LLA*) to relax the angle constraints. Due to the nature of random sampling, RRT* is asymptotically optimal but still not optimal, then we propose LoS-Check RRT* (LoS-RRT*). In order to solve the problems caused by improper settings of the planning resolution, we propose the LoS-Slider (LoSS) smoothing method. Through experimental comparison, it can be found that angle-constrained versions of LLA* and LoS-RRT* can both generate the near-optimal paths. Meanwhile, the experiment result shows that LLA* performs better than Theta* and Lazy Theta* under angle constraints. The planned path will be even closer to the optimal (shortest) solution after the smoothing of LoSS algorithm. Changwu Zhang, Hengzhu Liu |
IJCNN | 1 |