Haochuan Wan

dblp:241/6866 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-0996-6325ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ZeroBlade: A Spatial Similarity-aware HiSparse MLP Engine for Neural Volume Rendering
Antong Li, Haochuan Wan, Xiangyu Zhang 0002, Xin Lou 0001
ISCAS2
2026 SCOPE-3D: An Energy Efficient Accelerator for Implicit Neural Representation-based Sparse-view Computed Tomography Reconstruction
Haochuan Wan, Xin Li 0245, Qing Wu 0001, Yuhan Gu, Wenyan Su, Yuyao Zhang 0005, Xin Lou 0001
ISCAS1
2026 A Real-Time Neural Representation via Algorithm-Hardware Synergy for Sparse-View CT Reconstruction
abstract
Sparse-view computed tomography (SVCT) is an advancement in computed tomography (CT) technology that aims to reduce the radiation dose during imaging. Reconstructing high-quality images from sparse-view (SV) projections is an ill-posed inverse problem. Recently, implicit neural representations (INRs) as a self-supervised paradigm for solving underdetermined inverse problems have demonstrated excellent performance in SVCT reconstruction. However, since INR-based approaches rely on subject-specific training, they require a significant investment of time to optimize from scratch. Consequently, previous INR methods have not been able to meet the requisite timeliness of reconstruction. In our work, we propose RTSyner, an algorithm-hardware collaboration framework that facilitates the real-time efficiency of CT reconstruction. On the algorithmic side, we introduce an efficient coordinate-based feature module that exploits the local latent features as a positional external condition, leveraging the limited structural information of corrupted images derived from the sensory domain. By fusing latent features and coordinate information, the model learns a neural representation of the final tomographic image. On the hardware side, we design a dedicated hardware architecture with a customized algorithm flow to improve reconstruction speed and reduce power consumption. Furthermore, we improve the efficiency of model inference through model quantization, which also facilitates the subsequent deployment of hardware. Our extensive experimental results demonstrate that the RTSyner based on neural representation has achieved real-time SVCT reconstruction through the synergistic acceleration of the algorithm and hardware. We further explore its application potential via volume reconstructions under more complex acquisition geometries.
Xin Li 0245, Haochuan Wan, Kangjie Long, Qing Wu 0001, Chenhe Du, Xin Lou 0001, Yuyao Zhang 0005
IEEE Trans. Neural Networks Learn. Syst.2
2026 An Energy-Efficient Edge Coprocessor for Neural Rendering With Explicit Data Reuse Strategies
abstract
Neural radiance fields (NeRFs) have transformed 3-D reconstruction and rendering, facilitating photorealistic image synthesis from sparse viewpoints. This work introduces an explicit data reuse neural rendering (EDR-NR) architecture, which reduces frequent external memory accesses (EMAs) and cache misses by exploiting the spatial locality from three phases, including rays, ray packets (RPs), and samples. The EDR-NR architecture features a four-stage scheduler that clusters rays on the basis of$Z$-order, prioritize lagging rays when ray divergence happens, reorders RPs based on spatial proximity, and issues samples out-of-orderly (OoO) according to the availability of on-chip feature data. In addition, a four-tier hierarchical RP marching (HRM) technique is integrated with an axis-aligned bounding box (AABB) to facilitate spatial skipping (SS), reducing redundant computations and improving throughput. Moreover, a balanced allocation strategy for feature storage is proposed to mitigate SRAM bank conflicts. Fabricated using a 40-nm process with a die area of 10.5 mm2, the EDR-NR chip demonstrates a$2.41\times $enhancement in normalized energy efficiency, a$1.21\times $improvement in normalized area efficiency, a$1.20\times $increase in normalized throughput, and a 53.42% reduction in on-chip SRAM consumption compared with state-of-the-art accelerators.
Binzhe Yuan, Xiangyu Zhang 0002, Yuefeng Zhang, Haochuan Wan, Zhechen Yuan, Junsheng Chen, Yunxiang He, Junran Ding, Chaolin Rao, Wenyan Su, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001
IEEE Trans. Very Large Scale Integr. Syst.5
2024 ZeroTetris: A Spacial Feature Similarity-based Sparse MLP Engine for Neural Volume Rendering
abstract
Neural Volume Rendering (NVR), a novel paradigm for the longstanding problem of photo-realistic rendering of virtual worlds, has developed explosively in the past three years. The unique and substantial computational requirements of NVR pose challenge on deploying NVR to existing dedicated accelerator for neural networks. In this work, we propose ZeroTetris, a spacial feature similarity-based sparse multilayer perceptron (MLP) hardware accelerator for NVR. By leveraging the unique similarity-based sparsity between adjacent sampling points in NVR models, ZeroTetris efficiently bypass the computation of zero activations, thereby enhancing energy efficiency. Evaluation results affirm the effectiveness of the proposed design, showcasing ZeroTetris's superior performance in both area and power efficiency compared to other dedicated sparse matrix multiplication or MLP accelerator designs.
Haochuan Wan, Linjie Ma, Antong Li, Pingqiang Zhou, Jingyi Yu 0001, Xin Lou 0001
DAC1
2023 Digital-Twin Prediction of Metamorphic Object Transportation by Multi-Robots With THz Communication Framework
abstract
To predict the transportation process of a class of metamorphic object, e.g., deformable interlink linear object (DLO), the system particularity and coupling complexity is much increased due to external metamorphic constraints compared with classic multi-robot manipulation system. To explore the coupling effect, a virtual DLO model described by wave PDE equations with velocity feedback is proposed for replacing external metamorphic constraints. Independent stability controller is designed for feedback control of virtual DLO from perspective of MSR considering model stabilization. With the integration of proposed virtual DLO model, a digital-twin prediction prototype system of metamorphic object transportation is established. To bridge the communication gap between physical and digital world, a THz communication-based system framework is proposed to implement the digital-twin prediction for extremely security-sensitive system. Both numeric simulation and industrial experiment prove the availability and feasibility.
Lin Zhang 0024, Haochuan Wan, Xianhua Zheng, Meizi Tian, Lingling Su
IEEE Trans. Intell. Transp. Syst.2
2022 ICARUS: A Specialized Architecture for Neural Radiance Fields Rendering
abstract
The practical deployment of Neural Radiance Fields (NeRF) in rendering applications faces several challenges, with the most critical one being low rendering speed on even high-end graphic processing units (GPUs). In this paper, we present ICARUS, a specialized accelerator architecture tailored for NeRF rendering. Unlike GPUs using general purpose computing and memory architectures for NeRF, ICARUS executes the complete NeRF pipeline using dedicated plenoptic cores (PLCore) consisting of a positional encoding unit (PEU), a multi-layer perceptron (MLP) engine, and a volume rendering unit (VRU). A PLCore takes in positions & directions and renders the corresponding pixel colors without any intermediate data going off-chip for temporary storage and exchange, which can be time and power consuming. To implement the most expensive component of NeRF, i.e., the MLP, we transform the fully connected operations to approximated reconfigurable multiple constant multiplications (MCMs), where common subexpressions are shared across different multiplications to improve the computation efficiency. We build a prototype ICARUS using Synopsys HAPS-80 S104, a field programmable gate array (FPGA)-based prototyping system for large-scale integrated circuits and systems design. We evaluate the power-performancearea (PPA) of a PLCore using 40nm LP CMOS technology. Working at 400 MHz, a single PLCore occupies 16.5 mm 2 and consumes 282.8 mW, translating to 0.105 uJ/sample. The results are compared with those of GPU and tensor processing unit (TPU) implementations.
Chaolin Rao, Huangjie Yu, Haochuan Wan, Jindong Zhou, Yueyang Zheng, Minye Wu, Anpei Chen, Binzhe Yuan, Pingqiang Zhou, Xin Lou 0001, Jingyi Yu 0001
ACM Trans. Graph.3