Hailiang Hu

dblp:314/9029 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2026 LegoMap: Optimization for High-Throughput Transformer Computing on AI Engine-Based FPGAs
Hailiang Hu, Haodong Chang, Donghao Fang, Zhenrui Wang, Wuxi Li, Rongjian Liang, Bo Yuan 0001, Jiang Hu 0001
FCCM1
2026 A New Approach to Performance-Driven Analog IC Placement
abstract
A major obstacle in analog design automation is that circuit performance is sensitive to layout, yet accurately capturing this impact within layout tools is very expensive. To address this challenge, we propose a performance-driven analog IC placement approach, called VPlace, guided by machine learning. Our approach leverages a novel application of the VQ-VAE technique to improve robustness during the placement stage, in conjunction with a recent machine learning-based macromodeling method. We further demonstrate that data preparation strategies, which directly affect the efficiency of investigating the solution space, play an important role in determining both the accuracy of machine learning models and the resulting circuit performance. Experimental results show that VPlace achieves 22%-26% and 10%-16% performance improvements over an open-source analog layout tool and a prior machine learning–based performance-driven analog placement technique, respectively.
Donghao Fang, Hailiang Hu, Wuxi Li, Jiang Hu 0001
ISPD2
2025 Global Placement Exploiting Soft 2D Regularity
abstract
Cell placement is a step of paramount importance in chip physical design and requests relentless effort for continuous improvement. Recently, designs with two-dimensional (2D) processing element arrays have become popular primarily due to their deep neural network hardware applications. The 2D array regularity is similar to but different from the regularity of conventional datapath designs. To exploit the 2D array regularity, this work develops a new global placement technique, Placement of Arrays with SOft Regularity (PASOR), built upon RePlAce, the state-of-the-art placement framework. Experimental results from various designs show that the proposed approach can reduce global routing wirelength by 11% and 6% compared to RePlAce and a previous work on datapath driven placement, respectively.
Donghao Fang, Boyang Zhang 0007, Hailiang Hu, Wuxi Li, Bo Yuan 0001, Jiang Hu 0001
ACM Trans. Design Autom. Electr. Syst.3
2024 SysMix: Mixed-Size Placement for Systolic-Array-Based Hierarchical Designs
abstract
Systolic array designs are gaining popularity due to their applications in hardware acceleration for ML computing, such as CNNs and transformers. Increasingly large ML models necessitate very high circuit energy-efficiency, which is highly correlated with minimizing placement wirelength in chip physical design. However, existing placement techniques are mostly general purpose and overlook unique properties of systolic array designs. We propose a mixed-size placement approach, called SysMix, which is tailored for systolic arrays and leverage their partial regularity in hierarchical design methodologies. Experimental results from multiple CNN designs show that SysMix achieves 53% wirelength reduction and 15X speedup compared to a commercial placer and a state-of-the-art academic placer.
Donghao Fang, Hailiang Hu, Wuxi Li, Bo Yuan 0001, Jiang Hu 0001
ICCAD2
2023 Systolic Array Placement on FPGAs
abstract
Systolic array designs have regained popularity in recent years, particularly for their applications in accelerating CNN (Convolutional Neural Network) computing in hardware, including on FPGAs. However, existing FPGA layout techniques are primarily designed for general-purpose applications and have not fully leveraged the regularity of systolic arrays to enhance solution quality. This paper presents a new algorithmic approach for systolic array placement on FPGAs. Our approach enables 23% – 25% wirelength reduction for CNN circuits compared to an industrial tool and state-of-the-art academic methods. Moreover, it usually leads to significantly reduced routing resource utilization, accelerated placement runtime and improved timing performance.
Hailiang Hu, Donghao Fang, Wuxi Li, Bo Yuan 0001, Jiang Hu 0001
ICCAD1
2022 Mapping Large Scale Finite Element Computing on to Wafer-Scale Engines
abstract
The finite element method has wide applications and often presents a computing challenge due to huge problem sizes and slow convergence rate. A leading-edge computing acceleration approach is to leverage wafer-scale engine, which contains more than 800K processing elements. The effectiveness of this approach heavily depends on how to map a finite element computing task onto such enormous hardware space. A mapping method is introduced to partition an object space into computing kernels, which are further placed onto processing elements. This method achieves the best overall result in terms of computing accuracy and communication cost among all the ISPD 2021 contest participants.
Yishuang Lin, Rongjian Liang, Hailiang Hu, Jiang Hu 0001
ASP-DAC4
2022 Global Placement Exploiting Soft 2D Regularity
abstract
Cell placement is such a critical step for chip physical design that it needs many kinds of efforts for improvement. Recently, designs with 2D processing element arrays have become popular primarily due to their deep neural network computing applications. The 2D array regularity is similar to but different from the regularity of conventional datapath designs. To exploit the 2D array regularity, this work develops a new global placement technique built upon RePlAce, the latest state-of-the-art placement framework. Experimental results from various designs show that the proposed technique can reduce half-perimeter wirelength and Steiner tree wirelength by about $6%$ and $12%$, respectively.
Donghao Fang, Boyang Zhang 0007, Hailiang Hu, Wuxi Li, Bo Yuan 0001, Jiang Hu 0001
ISPD3