Hongzheng Tian

dblp:347/5820 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0004-6253-5059ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Reconfigurable computing and FPGAs · 68% Hardware accelerators and domain-specific architectures · 32%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Artificial intelligence
2 papers
Motion planning and robot control · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs
FPGA accelerator
1.622025
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization · ACM Trans. Archit. Code Optim. 2025
Accelerating Autonomous Path Planning on FPGAs with Sparsity-Aware HW/SW Co-Optimizations · FPGA 2024
Mathematical optimization › continuous optimization › nonlinear optimization
quadratic programming
0.912025
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization · ACM Trans. Archit. Code Optim. 2025
Hardware accelerators and domain-specific architectures
quadratic programming solver
0.812024
Accelerating Autonomous Path Planning on FPGAs with Sparsity-Aware HW/SW Co-Optimizations · FPGA 2024
Robotics › Motion planning and robot control
path planning
0.522025
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization · ACM Trans. Archit. Code Optim. 2025
Accelerating Autonomous Path Planning on FPGAs with Sparsity-Aware HW/SW Co-Optimizations · FPGA 2024
Robotics › Motion planning and robot control › path planning › path generation
collision-free path generation
0.312025
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization · ACM Trans. Archit. Code Optim. 2025

Methods — techniques the papers use, named apart from their topics

preconditioned conjugate gradient · 4.1multi-level dataflow optimization · 2.6alternating direction method of multipliers · 2.6HW/SW co-design · 2.6task-level parallelism · 1.5operator splitting · 1.5hardware pipelining · 1.5
YearPublicationVenuePosition
2025 HeteroBench: Multi-kernel Benchmarks for Heterogeneous Systems
Hongzheng Tian, Alok Mishra 0002, Rolando P. Hong Enriquez, Dejan S. Milojicic, Eitan Frachtenberg, Sitao Huang
ICPE1
2025 A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization
abstract
Path planning is a critical task for autonomous driving, aiming to generate smooth, collision-free, and feasible paths based on input perception and localization information. The planning task is both highly time-sensitive and computationally intensive, posing significant challenges to resource-constrained autonomous driving hardware. In this article, we propose an end-to-end framework for accelerating path planning on FPGA platforms. This framework focuses on accelerating quadratic programming (QP) solving, which is the core of optimization-based path planning and has the most computationally-intensive workloads. Our method leverages a hardware-friendly alternating direction method of multipliers (ADMM) to solve QP problems while employing a highly parallelizable preconditioned conjugate gradient (PCG) method for solving the associated linear systems. We analyze the sparse patterns of matrix operations in QP and design customized storage schemes along with efficient sparse matrix multiplication and sparse matrix-vector multiplication units. Our customized design significantly reduces resource consumption for data storage and computation while dramatically speeding up matrix operations. Additionally, we propose a multi-level dataflow optimization strategy. Within individual operators, we achieve acceleration through parallelization and pipelining. For different operators in an algorithm, we analyze inter-operator data dependencies to enable fine-grained pipelining. At the system level, we map different steps of the planning process to the CPU and FPGA and pipeline these steps to enhance end-to-end throughput. We implement and validate our design on the AMD ZCU102 platform. Our implementation achieves state-of-the-art performance in both latency and energy efficiency compared with existing works, including an average 1.48× speedup over the best FPGA-based design, a 2.89× speedup compared with the state-of-the-art QP solver on an Intel i7-11800H CPU, a 5.62× speedup over an ARM Cortex-A57 embedded CPU, and a 1.56× speedup over state-of-the-art GPU-based work. Furthermore, our design delivers a 2.05× improvement in throughput compared with the state-of-the-art FPGA-based design.
Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang
ACM Trans. Archit. Code Optim.3
2024 Accelerating Autonomous Path Planning on FPGAs with Sparsity-Aware HW/SW Co-Optimizations
abstract
Path planning is a critical task in autonomous driving systems, with quadratic programming being the most time-consuming component. Solving quadratic programming problems using a CPU not only takes a long time but can also lead to high power consumption and costs. In this work, we propose an FPGA-based acceleration method for quadratic programming based path planning problems. Our approach leverages an operator splitting solver for quadratic programs (OSQP) and employs the preconditioned conjugate gradient (PCG) method for solving linear equations, which proves to be more scalable and hardware-friendly than the original direct method. We propose optimizations for better memory management, and boost processing throughput and reduce execution time by task level and operator level parallelism with hardware pipelining. Our FPGA-based implementation achieves up to 1.8× speedup and 3.2× power reduction compared with the Intel i5 CPU, 3.1× speedup compared with ARM Cortex-A57.
Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang
FPGA4
2024 A Sparsity-Aware Autonomous Path Planning Accelerator with Algorithm-Architecture Co-Design
abstract
Path planning is a critical task in autonomous driving systems that is most susceptible to real-time constraints but often demands computationally intensive mathematical solvers, two contradictory goals. This conflict makes the computing of path planning a paramount challenge. At the heart of most path planners is the quadratic programming (QP) solver, which places excessive demands on the CPU in real-world autonomous driving applications. In this paper, we present an FPGA-based acceleration framework for path planning problems. Our approach leverages an operator splitting solver for quadratic programs (OSQP) and employs the preconditioned conjugate gradient (PCG) method for solving linear systems, which are customized to be more hardware-friendly than prior works. Specific memory management and parallel processing were tailored to the matrix pattern, and the incorporation of pipelining was executed to enhance throughput and execution speed. Our FPGA-based implementation achieves state-of-the-art performance against existing works, including an average 1.98× speedup compared with the state-of-the-art QP solver on Intel i7-11800H CPU, 3.90× speedup over an ARM Cortex-A57 embedded CPU, and 12.3× speedup over an NVIDIA RTX 3090 GPU.
Hongzheng Tian, Bo Yu 0014, Shaoshan Liu, Sitao Huang
ICCAD4