Xiuhao Zhang

dblp:415/5240 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0000-0003-0489-0197ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 50% Hardware accelerators and domain-specific architectures · 50%
Computer graphics and multimedia
1 paper
Rendering · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › neural rendering accelerator
3d gaussian splatting accelerator
0.912025
GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization · DAC 2025
GPUs and heterogeneous computing
graphics accelerator
0.912025
GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization · DAC 2025
Rendering › gaussian splatting
3d gaussian splatting
0.312025
GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization · DAC 2025
Rendering › neural rendering
radiance field rendering
0.312025
GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization · DAC 2025

Methods — techniques the papers use, named apart from their topics

gaussian-centric rasterization · 1.7depth speculation · 1.7
YearPublicationVenuePosition
2026 Metal Stack Exploration for Front- Versus Back- Side Clock and Signal Allocation for Advanced CMOS PPA Improvements
abstract
This study explores the potential of back-side contact (BSC) technology for enhancing power, performance, and area (PPA) in integrated circuits (ICs). Through metal stack exploration experiments on a 32-bit RISC-V core, the research demonstrates significant PPA gains by optimizing front- and back-side metal (BSM) stack and layer purpose allocation. Key findings include a 2.7% power reduction per added back-side layer for clock routing, a 5.5% frequency improvement with combined front- and back-side clock routing, and up to 14.71% power reduction or 10% frequency increase through metal stack optimization for low-power and high-performance targets, respectively. A similar study on an FPGA was also conducted by optimizing the layer stack configuration to achieve maximum frequency gains of 55.34% or 8.95% in total power consumption, respectively, compared to the baseline of an eight front-side-only layer stack configuration.
Sirish Oruganti, S. S. Teja Nibhanupudi, Anup Ashok Kedilaya, Xiuhao Zhang, Jaydeep P. Kulkarni
IEEE Trans. Very Large Scale Integr. Syst.5
2025 GSAcc: Accelerate 3D Gaussian Splatting via Depth Speculation and Gaussian-centric Rasterization
abstract
D Gaussian Splatting (3DGS) has emerged as a promising real-time photorealistic radiance field rendering technique. Existing GPU and hardware accelerators face limitations due to insufficient parallelism in sequential rendering pipeline stages and the memory overhead associated with interim results. This paper presents GSAcc, a hardware accelerator co-designed with dataflow to render compressed 3DGS models on edge platforms efficiently. GSAcc enhances 3DGS rendering performance through several key innovations. First, it introduces Gaussian depth speculation, parallelizing preprocessing and sorting tasks. Second, GSAcc adopts a Gaussian-centric dataflow that interleaves preprocessing and rasterization, allowing all rendering steps to execute concurrently without storing intermediate results. Finally, it employs dedicated hardware acceleration to address sorting and rasterization bottlenecks within the optimized dataflow. We implemented and synthesized GSAcc using Intel16 PDK and evaluated its performance on real-world 3DGS scenes. Compared with desktop GPUs, GSAcc achieves up to $1.66 \times 10^{4} \mathrm{x}$ Power-Performance-Area (PPA) improvement as well as 48.7 x energy savings. Additionally, GSAcc outperforms the state-of-the-art hardware accelerator GSCore with up to 2.3x PPA improvement and 2.9x energy savings.
Mengtian Yang, Yipeng Wang 0017, Chieh-Pu Lo, Xiuhao Zhang, Sirish Oruganti, Jaydeep P. Kulkarni
DAC4