Haoyang Sang

dblp:352/7810 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-3959-5051ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2026 SVRoM: A 9.52-mW Video Understanding Smart Vision SoC With On-Chip Sensing and Similarity-Aware SRAM/ROM CIM Macro
abstract
The rapid growth of virtual reality (VR), augmented reality (AR), extended reality (XR), and intelligent surveillance systems has driven increasing demand for video understanding tasks on edge devices. However, supporting such tasks on edge devices remains challenging due to the limited computational resources, memory constraints, power budgets, and high data transition latency. To address these challenges, this work proposes an ultralow-power smart vision System-on-Chip (SoC) with the following features: 1) a bitline segmented, parallel-charging read-only memory (ROM) compute-in-memory (CIM) macro with power gating and coupled coding; 2) heterogeneous similarity-aware hybrid SRAM/ROM cores to exploit temporal similarity with high core utilization; 3) an intracore and intercore pipeline with hierarchical dataflow for low-power data transmission; 4) a tilewise mixed-precision weight quantization and mapping scheme for weight data compression; and 5) a hierarchical multimodal system trigger mechanism utilizing an on-chip CMOS imager with near sensor caching to skip unnecessary inferences. The proposed ROM CIM macro achieves an area efficiency of 0.753–1.673 TOPS/mm2and a storage density of 4706 Kb/mm2, achieving a$3.2\!\!-\!\!14.83\times $improvement in density Figure of Merit (FoM) over the state-of-the-art (SoTA) ROM CIM designs. The proposed SoC also shows an ultralow always-on power of$0.13~\mu $W and an average active power of 9.52 mW, which demonstrates a high energy efficiency of 13.6–18.3 TOPS/W (ResNet-20 with fixed-point 8-bit activation and 10-bit weight) on video understanding tasks, showing a$1.8-3.0\times $improvement over the existing smart vision SoCs.
Haoyang Sang, Ningchao Lin, Guangshu Zhao, Man Kay Law
IEEE Trans. Very Large Scale Integr. Syst.1
2025 BROCA: A Low-power and Low-latency Conversational Agent RISC-V System-on-Chip for Voice-interactive Mobile Devices
Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Haoyang Sang, Dongseok Im, Sangyeob Kim, Chaeyun Jeong, Yujin Moon, Hoi-Jun Yoo
HCS5
2025 BAGNet: A Boundary-Aware Graph Neural Network for SRAM Yield Analysis in Post-LayoutSimulation
abstract
Yield analysis has grown in significance with the increasing integration of SRAM arrays. The post-layout simulation introduces strong inter-column correlations in SRAM caused by parasitic parameters, thereby complicating yield analysis. However, most existing methods only consider the pre-layout simulation of SRAM circuits. In this paper, we present BAGNet: a boundary-aware Graph Neural Network (GNN) for SRAM yield analysis in post-layout simulation. We introduce a GNN module that learns the graph representations of SRAM arrays while generating feature vectors. We then construct an accurate surrogate model by the Multilayer Perceptron (MLP) to provide predictions for circuit performances. Given that delineating failure boundaries is vital for yield estimation, we propose an innovative nonlinear mapping strategy and an adaptive iterative strategy integrated with BAGNet, thus endowing our model with boundary-aware capability. After the model is built, we employ the importance sampling (IS) method on our surrogate model to deliver efficient and accurate yield estimation without time-consuming circuit simulations. Experimental results demonstrate that BAGNet outperforms the state-of-the-art method with 1.823.52x speedup, without losing accuracy.
Haoyang Sang, Changhao Yan, Zhaori Bi, Keren Zhu 0001, Xuan Zeng 0001
ICCAD1
2025 Ultra Low Power Video Understanding Smart Vision SoC with On-Chip Sensing and Hybrid Similarity-Aware SRAM/ROM CIM Macro
abstract
The proliferation of augmented reality (AR), virtual reality (VR), and extended reality (XR) edge devices drives the demand for video understanding, which imposes substantial demands on computational resources, memory, and energy efficiency. This work proposes an ultra-low power and highly compact smart vision SoC, composed of: 1) a bit-line (BL) segmented parallel charging ROM CIM macro with power gating and coupled coding; 2) heterogeneous similarity-aware hybrid SRAM/ROM cores for exploiting temporal similarity with high utilization; 3) a hierarchical multi-modal system trigger uses an on-chip CMOS imager with near-sensor caching for skipping unnecessary inference; and 4) a RISC-V core featuring dedicated ISA extensions to support flexible workload allocation and programmability. The proposed ROM CIM macro achieves a storage density of 4705 kb/mm2and an area efficiency of 0.753 TOPS/mm2, demonstrating a 3.2~14.83× improvement in density FoM compared to state-of-the-art (SOTA) CIM macros. Furthermore, the SoC achieves an energy efficiency of 13.6~18.3 TOPS/W on ResNet-20 (fixed-point 8-bit activation and 10-bit weight), representing a 1.8~3.0× improvement over existing smart vision SoCs.
Haoyang Sang, Ningchao Lin, Guangshu Zhao, Man Kay Law
ISCAS1
2025 A Token-Passing-Based Trigger-Prediction Methodology for Event-Driven ToF Sensors
abstract
This paper presents an ambient light robust methodology for event-driven (ED) time-of-flight (ToF) sensors, mainly targeting 3D detection under outdoor scenarios. Different from prior ED methodologies, we propose a trigger-prediction scheme with token-passing algorithm to effectively filter out ambient photons for improving the system performance under strong background light. Specifically, we determine the variation in distance by detecting the change in the width of timing windows in the trigger stage, while extracting the alteration of the number of photons in the multi-step prediction stage. This can lower the error rate while ensuring a fast imaging speed. Using a 180 nm CMOS process, we developed a behavior model including the device characteristics and environmental parameters for performance evaluation. Monte Carlo simulation results demonstrate that with a 10 m range, our proposed method can achieve a depth accuracy of less than 5 cm under 10 klux, and less than 25 cm under 80 klux, respectively, and an event response time of less than 384 µs under dynamic scenes.
Yifei Xiang, Haoyang Sang, Man Kay Law
ISCAS3
2024 Space-Mate: A 303.5mW Real-Time NeRF SLAM Processor with Sparse-Mixture-of-Experts-based Acceleration
abstract
NeRF-based SLAM for robotic applications face computation barrier
Seokchan Song, Haoyang Sang, Dongseok Im, Donghyeon Han, Sangyeob Kim, Hongseok Lee, Hoi-Jun Yoo
HCS2