Baohui Xie

dblp:401/7482 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0006-6420-5517ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 50% Hardware accelerators and domain-specific architectures · 50%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
FPGA-based CNN accelerator
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Electronic design automation › physical design › placement › circuit placement
FPGA placement
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Electronic design automation
physical design
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025
Electronic design automation › physical design
placement
0.912025
DSPlacer: DSP Placement for FPGA-based CNN Accelerator · DAC 2025

Methods — techniques the papers use, named apart from their topics

min-cost flow · 0.9integer linear programming · 0.9graph convolutional network · 0.9
YearPublicationVenuePosition
2025 DSPlacer: DSP Placement for FPGA-based CNN Accelerator
abstract
Deploying convolutional neural networks (CNNs) on hardware platforms like Field Programmable Gate Arrays (FPGAs) has garnered significant attention due to their inherent flexibility and parallelism. Achieving optimal timing closure remains a critical challenge, as placement directly impacts clock frequency and throughput. Existing approaches often face scalability issues with large designs or fail to formalize placement rules into automated algorithms. In this paper, we propose DSPlacer, a novel DSP placement framework designed for diverse CNN accelerator architectures in the context of FPGA design. The proposed approach iteratively optimizes the placement of datapath DSPs to enhance timing performance. To achieve this, DSPlacer integrates several advanced techniques, including graph convolutional network-based datapath DSP identification, DSP graph construction, min-cost-flow DSP assignment, and integer linear programming (ILP)-based cascade constraint legalization. These techniques collectively address two key requirements for datapath DSP placement: (1) cascading datapath DSPs to achieve a compact layout, and (2) preserving direct datapath information between the processing system and programmable logic. The framework has been evaluated on multiple academic benchmarks and compared against AMD Xilinx Vivado 2020.2 and AMF-Placer 2.0. Experimental results demonstrate that DSPlacer improves Worst Negative Slack (WNS) by 32% and 65%, respectively, highlighting its efficacy and superiority.
Baohui Xie, Xinrui Zhu, Yuan Pu 0001, Tongkai Wu, Xiaofeng Zou, Bei Yu 0001, Tinghuan Chen
DAC1
2024 RISCSparse: Point Cloud Inference Engine on RISC-V Processor
abstract
Machine learning on point clouds is increasingly accessible at the edge, notably in applications such as autonomous driving. However, the sparse and irregular nature of point clouds presents significant latency challenges on general-purpose hardware. RISC-V, with its evolving ecosystem, offers a promising platform for embedding intelligence at the edge due to its full-stack scalability. This paper focuses on the advanced point cloud operation known as submanifold convolution (SC), deploying submanifold sparse convolutional networks (SSCNs) on a RISC-V System-on-Chip (SoC) designed within the Chipyard framework. We address three critical bottlenecks of SSCNs- Rule Map Construction (Mapping), Gather-MatMul-Scatter (GMS), and uncombined operation - to meet the real-time inference requirement for the on-chip implementation. By leveraging the RISC-V Vector extension and Gemmini, an open-source full-stack DNN accelerator generator, we vectorize the Mapping process, offload GEMM-related operations to the Gemmini Systolic Array, and cooperatively use the Systolic Array and vector processing units to reduce the memory footprint. Our evaluations show that the RISC-V-based SSCNs implementation achieves an average of 11.73× and 13.1× overall speedups with a small workload compared to TorchSparse on Edge-CPU, a state-of-the-art point cloud inference engine, for 3D segmentation and detection tasks, respectively. When contrasting with TorchSparse on an Edge-GPU, our implementation still delivers a notable improvement, with average speedups of 1.63× for 3D segmentation and 1.07× for detection tasks.
Shangran Lin, Xinrui Zhu, Baohui Xie, Tinghuan Chen, Cheng Zhuo, Qi Sun 0002, Bei Yu 0001
ICCAD3