Zhanhua Zhang

dblp:337/4838 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DGC-GS: Enhancing geometric consistency in sparse-view 3D Gaussian splatting
Zhanhua Zhang, Xingjian Cao, Can Tian, Xiaobing Ding, Wancheng Ge
Neurocomputing3
2025 FreeTimeGS: Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene Reconstruction
abstract
This paper addresses the challenge of reconstructing dynamic 3D scenes with complex motions. Some recent works define 3D Gaussian primitives in the canonical space and use deformation fields to map canonical primitives to observation spaces, achieving real-time dynamic view synthesis. However, these methods often struggle to handle scenes with complex motions due to the difficulty of optimizing deformation fields. To overcome this problem, we propose FreeTimeGS, a novel 4D representation that allows Gaussian primitives to appear at arbitrary time and locations. In contrast to canonical Gaussian primitives, our representation possesses the strong flexibility, thus improving the ability to model dynamic 3D scenes. In addition, we endow each Gaussian primitive with an motion function, allowing it to move to neighboring regions over time, which reduces the temporal redundancy. Experiments results on several datasets show that the rendering quality of our method outperforms recent methods by a large margin. The code will be released for reproducibility.
Yifan Wang 0026, Peishan Yang, Zhen Xu 0008, Jiaming Sun 0002, Zhanhua Zhang, Hujun Bao, Sida Peng, Xiaowei Zhou 0001
CVPR5
2025 LiDAR-RT: Gaussian-based Ray Tracing for Dynamic LiDAR Re-simulation
abstract
This paper targets the challenge of real-time LiDAR re-simulation in dynamic driving scenarios. Recent approaches utilize neural radiance fields combined with the physical modeling of LiDAR sensors to achieve high-fidelity re-simulation results. Unfortunately, these methods face limitations due to high computational demands in large-scale scenes and cannot perform real-time LiDAR rendering. To overcome these constraints, we propose LiDAR-RT, a novel framework that supports real-time, physically accurate LiDAR re-simulation for driving scenes. Our primary contribution is the development of an efficient and effective rendering pipeline, which integrates Gaussian primitives and hardware-accelerated ray tracing technology. Specifically, we model the physical properties of LiDAR sensors using Gaussian primitives with learnable parameters and incorporate scene graphs to handle scene dynamics. Building upon this scene representation, our framework first constructs a bounding volume hierarchy (BVH), then casts rays for each pixel and generates novel LiDAR views through a differentiable rendering algorithm. Importantly, our framework supports realistic rendering with flexible scene editing operations and various sensor configurations. Extensive experiments across multiple public benchmarks demonstrate that our method outperforms state-of-the-art methods in terms of rendering quality and efficiency. Our code and data are available at https://github.com/zju3dv/LiDAR-RT.
Chenxu Zhou, Lvchang Fu, Sida Peng, Yunzhi Yan, Zhanhua Zhang, Jiazhi Xia, Xiaowei Zhou 0001
CVPR5
2025 An Optimization-Aware Prerouting Timing Prediction Framework Based on Multimodal Learning
abstract
Accurate and efficient prerouting timing estimation is particularly crucial during placement to alleviate time-consuming design iterations. Machine-learning (ML)-based methods have been introduced recently to predict the post-routing timing results at placement stage, but most of them neglect the impact of timing optimization during physical design, suffering from accuracy loss due to inconsistent circuit netlist. In this work, an optimization-aware prerouting timing prediction framework based on multimodal learning is proposed to calibrate the timing changes between placement and routing stages, where the local netlist and layout information are extracted by graph neural network (GNN) and convolutional neural network (CNN), respectively, while the global information along the path is further extracted by Transformer network. Based on the predicted post-routing timing results by the proposed framework, timing optimization guidance is generated to enhance traditional design flow with better physical implementation quality. Experimental results demonstrate that for the OpenCores benchmark circuits under TSMC 22nm process, the proposed framework achieves significant correlation and accuracy improvement with an average of 0.9219 in terms of R2 score and 2.12% of mean absolute percentage error (MAPE) as well as an average runtime acceleration of$645\times $compared with traditional design flow on testing designs. With the timing optimization guidance, significant worst negative slack (WNS) and total negative slack (TNS) improvement are achieved compared with traditional flow after placement and routing, respectively, without noticeable area, power, wire length, and the number of design rule check (DRC) violations increase.
Peng Cao 0002, Yusen Qin, Guoqing He, Zhanhua Zhang, Yuyang Ye 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 A Physical and Timing Aware Placement Optimization Framework Based on Graph Neural Network
abstract
Timing-driven placement is crucial in physical design flow with significant impact on later routability and ultimate manufacturability, which may deviate from finding the optimal solution and/or lead to unnecessary iterations, suffering from interleaved optimization steps and the corresponding inaccurate timing estimation. To solve this issue, we propose a Physical and Timing Aware framework with Graph Neural Network, PTA-GNN, which provides the candidate gate sizing and buffer insertion solutions as well as the timing constraint for potential violated paths as guidance to improve placement quality significantly. Experimental results on the OpenCores benchmarks with 22nm technology demonstrate that the proposed placement optimization framework achieves up to 89.09% worst negative slack (WNS), 55.47% total negative slack (TNS) improvement and 25.36% reduction on the number of violating paths (#VP). Our framework benefits the later routing stage with 2.19% wire-length decrease and 22% runtime reduction compared to standard physical design flow.
Zhanhua Zhang, Guoqing He, Peng Cao 0002
ICCAD2
2024 LAG-Sizer: A Novel Gate Sizer Based on Leak Generative Adversarial Network with Feature Fusion
abstract
Gate sizing is an NP-hard problem to achieve Performance, Power and Area (PPA) optimization. Recently proposed learning-based approaches struggle to overcome the runtime issue of traditional heuristics, but lack the consideration of the intrinsic features for candidate gates in library and could not address the inequality issue of candidate sizes for different gates properly, suffering from insufficient design space exploration and inaccurate sizing assignment. In this work, based on a variant of generative adversarial network, Leak Adversarial Generation (LAG), a novel LAG-Sizer is proposed to model gate sizing as sequence generation problem, which breaks the traditional adversarial network by leaking the discriminator feature information into the generator to guide sizing generation. Feature fusion technique is introduced to comprehensively consider circuit feature and cell library feature while a unified classification is proposed to perfectly solve the inequality issue for sizing. The proposed sizer was validated with IWLS2005 and Opencores benchmark circuits under 22nm process. Experimental results demonstrate that an average of 4.6% Total Negative Slack (TNS) improvement and 15.6% number of violating endpoints (NVE) reduction are achieved by this work with similar area and power consumption compared to commercial tools as well as significant runtime speedup of 47.8×.
Zhanhua Zhang, Guoqing He, Peng Cao 0002
ICCAD1
2023 RTCoInfer: Real-Time Collaborative CNN Inference for Stream Analytics on Ubiquitous Images
abstract
Emerging intelligent applications based on accurate and timely stream analytics require real-time CNN inference of massive data continuously generated at the pervasive end devices. Due to the resource constraints, neither computing locally at end devices nor transmitting to remote servers is competent for computation-intensive CNN inference on large-volume images in real-time. Therefore, Collaborative Inference (CI), which conducts inference sequentially from the local device to the remote server with compressed intermediate inference data, is rapidly promoted. Due to the essential communication in collaboration, the CI efficiency is sensitive to network conditions, and will degrade under the unpredictable network fluctuations in practice, which may cause a severe delay in CI and degrade the responsiveness of stream analytics. For accurate and timely stream analytics in practical fluctuating networks, we present RTCoInfer, the real-time CI framework with run-time transmission adaption considering the network conditions. Specifically, we propose a novel Switchable CNN integrating CNNs with different compression rates on the partition layer for the run-time transmission adjustment, and construct a real-time controller determining the compression rate to maintain the real-time CI for stream analytics. Extensive experiments show that, compared with state-of-the-art methods, RTCoInfer achieves better efficiency and unprecedented resilience in real-time stream analytics.
Zhanhua Zhang, Shusen Yang, Cong Zhao 0001, Xuebin Ren, Hanqiao Yu, Siyan Guo
IEEE J. Sel. Areas Commun.1
2023 Efficient and Accurate ECO Leakage Optimization Framework With GNN and Bidirectional LSTM
abstract
Engineering change order (ECO) plays an important role in design flow to perform leakage optimization with gate-sizing and$V_{\mathrm{ th}}$assignment approaches. Unfortunately, it is extremely time consuming due to the iterative nature of cell swap and timing check. Many learning-based methods, especially, graph neural networks (GNNs), have been utilized in leakage optimization to predict$V_{\mathrm{ th}}$assignment, but most of them treat the cells and their neighborhood cells uniformly when aggregating cell-level topology information to gather design-level information and discard the path-level information, suffering from accuracy loss, which could be exploited by bidirectional long short-term memory (BiLSTM) network. In this work, a GNN-BiLSTM-based framework is proposed to perform commercial-quality$V_{\mathrm{ th}}$assignment for leakage optimization by learning design-level and path-level information and is validated with the benchmarks from Opencores and IWLS 2005 under TSMC 28 nm technology. The experimental results demonstrate that the proposed framework achieves the most accurate$V_{\mathrm{ th}}$assignment prediction compared with the competitive models with F1-score ranging from 0.954 to 0.975 for seen designs and from 0.945 to 0.965 for unseen designs, respectively. The divergence between the leakage optimization results of this work and the commercial tool is limited to be between 8.5% and 26.1%, which is reduced by at least$2.2\times $compared with prior works. Owing to efficient training convergence and inference speed, our approach achieves significant runtime improvement by up to$10\times $over commercial tool with similar leakage optimization results.
Peng Cao 0002, Guoqing He, Zhanhua Zhang, Jun Yang 0006
IEEE Trans. Very Large Scale Integr. Syst.4
2022 CNNPC: End-Edge-Cloud Collaborative CNN Inference With Joint Model Partition and Compression
abstract
Edge Intelligence (EI) aims at addressing concerns like response latency risen by the conflict between predominating Cloud-based deployments of computationally intensive AI applications and the expensive uploading of explosive end data. Convolutional Neural Networks (CNNs) leading the latest flourish of AI inevitably suffer from the aforementioned conflict. There emerge increasing EI-driven attempts on fast CNN inference with high accuracy in the End-Edge-Cloud (EEC) collaborative computing paradigm, where, however, neither model compression approaches for on-device inference nor collaborative inference methods across devices can effectively achieve the trade-off between latency and accuracy of End-to-End (E2E) inference. In this article, we present CNNPC that jointly partitions and compresses CNNs for fast inference with high accuracy in collaborative EEC systems. We implemented CNNPC (source code available athttps://github.com/IoTDATALab/CNNPC) and evaluated its performance within extensive real-world EEC scenarios. Experimental results demonstrate that, compared with state-of-the-art single-end and collaborative approaches, without obvious accuracy loss, collaborative inference based on CNNPC is up to$1.6\times$and$5.6\times$faster, and requires as low as$4.30\%$and$6.48\%$communications, respectively. Besides, when determines the optimal strategy, CNNPC requires as low as$0.1\%$actual compression operations that the traversal method (the only viable method providing the theoretically optimal strategy) requires.
Shusen Yang, Zhanhua Zhang, Cong Zhao 0001, Siyan Guo
IEEE Trans. Parallel Distributed Syst.2