Zhangcheng Huang 0001

dblp:254/6073-1 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-2551-7044ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 10 since 2021
YearPublicationVenuePosition
2026 A Temperature-Adaptive Bias Generator with Threshold-Voltage Compensation from 3.6 K to 410 K
Yingzhe Sha, Jiaxuan Weng, Zhangcheng Huang 0001, Qi Liu 0010
ISCAS4
2025 Hierarchical Integration of Reinforcement Learning and Optimization Algorithms for Time-Efficient Design Automation of Complex Analog Circuit
abstract
Design automation of complex analog circuits (CAC) with multiple sub-blocks is challenging mainly due to large design search space, uncertain intermediate subgoal creation, and lengthy CAC simulation runtime. In this work, we propose a hierarchical and heterogeneous integration framework as a fully automated and time-efficient CAC design optimization solution. In Particularly, we (i) decompose CAC into two levels hierarchically and for the first time introduce hierarchical RL agents with hindsight and subgoal testing to automate the subgoal creation between these two levels. The subgoal converges to the optimal value through algorithm interactions. (ii) We enable high-level design space dimensionality reduction, minimize CAC simulation runs through a buffer hold, and employ low-level sub-block execution parallelization to reduce overall runtime. (iii) We construct a heterogeneous integration of different RL algorithms and black-box optimization algorithms in hierarchy to further boost the speed by benefiting both from the hierarchical structure and the advantages of each different algorithm. Experiments on four CAC topologies demonstrate that this framework achieves a maximum of 11.4× speed up compared to existing methods at the desired figure-of-merit. This work opens up a time efficient design automation route for complex analog circuits and systems.
Xingwei Feng, Yifan Xu 0026, Zhangcheng Huang 0001, Wuyi Xu, Zhaori Bi, Fan Yang 0001, Xuan Zeng 0001, Ye Lu 0005
ACM Trans. Design Autom. Electr. Syst.3
2024 CEDAR: Computing-in-pixel Edge-aware Detection and Reconstruction Architecture for High-resolution 3D Imaging
abstract
Large-format single-photon avalanche diode (SPAD)-based direct time of flight (dToF) sensors are expected to be widely applied in future L5 full driving automation. However, the high-power in-pixel TDCs and the huge amount of data generated by multiframe histogram sampling impose limitations on the pixel format of SPAD-based dToF sensors. To tackle this challenge, we proposed the Computing-in-pixel Edge-aware Detection and Reconstruction (CEDAR) architecture. In this architecture, edge pixels are recognized by charge-domain convolution (CDC) computing, and noise pixels are eliminated by in-memory denoising (IMD). Only few TDCs in these edge pixels are activated, resulting in significant power and data savings. Afterward, the full-format image is reconstructed by a U-Net using the obtained depth information from these edge pixels. For the first time, we proposed a high-resolution 512 × 512 SPAD-based dToF sensor with a low power of 83.3 mW, a distance accuracy of 0.9 cm, and a frame rate of 60 fps. The high-resolution 3D image can be reconstructed by only 3.5% sparse edge pixels, achieving a PSNR of 35.2 dB. The CEDAR architecture can achieve 16× pixel format and image resolution improvement under the same constraint of power dissipation.
Bu Chen, Zhangcheng Huang 0001, Qi Zheng 0004, Weiyi Tang, Hankun Lv, Chixiao Chen, Jianlu Wang, Qi Liu 0010
DAC2
2024 CAMPER: Exploring the Potential of Content Addressable Memory for 3D Point Cloud Efficient Range Search
abstract
The use of Light Detection and Ranging (LiDAR) for sensing has continuously improved the precision and performance of autonomous driving. At the same time, the large number of high-precision point clouds generated by LiDAR require real-time processing, and the range search is the key part of the processing pipeline. Content-Addressable Memory (CAM) has proven its efficiency for search tasks on switches and routers, but so far, there is still a lack of exploration on its application in point cloud range search. In this work, we propose CAMPER, a CAM-centered accelerator, aiming to explore the potential of CAM for point cloud range search. We developed a ripple comparison 13T (RC-13T) CAM cell for distance comparison, designed a spatial approximation search algorithm based on Chebyshev distance, and discussed the flexibility and scalability of the architecture. The results show that in the range search task of 64k@64k points, CAMPER achieves a latency of 0.83ms and a power consumption of 114.6mW. Compared with GPU, the throughput is increased by 10.4×; compared with SOTA accelerator, the energy efficiency is increased by about 228×.
Jiapei Zheng, Lizhou Wu, Yutong Su, Zhangcheng Huang 0001, Chixiao Chen, Qi Liu 0010
DAC5
2024 A 128×128 CMOS SPAD Receiver for 500Mbps Free Space Optical Communication with Column-wise Decoding and Fast Spot Tracking
abstract
This work presents a 128×128 pixel array receiver based on single-photon avalanche diode (SPAD) for free space optical communication (FSOC). Each pixel incorporates an active quenching circuit and a delay-time-adjusting circuit to reduce the afterpulsing effect and the dead time. To address the challenge of high-speed transmission of massive data in a large-format SPAD array, a column-wise decoding circuit with reduced bus parasitic capacitance and voltage-sensitive discrimination is proposed, which significantly reduces the latency of data transmission. Additionally, the receiver includes cluster engines with highly parallelized computation capabilities for tracking the central addresses of a laser spot. The chip has been designed using a 130nm CMOS technology. Simulation results of the receiver indicate that a bit error rate (BER) of 3 × 10−4can be achieved at 500Mbps with a sensitivity of -41dBm, under random NRZOOK bitstreams. Furthermore, the chip demonstrates 100% accuracy in tracking the laser spot at a rate of 100kHz during 1000 transceiver simulations.
Bu Chen, Zhangcheng Huang 0001, Qi Liu 0010
ISCAS2
2024 Multiagent Based Reinforcement Learning (MA-RL): An Automated Designer for Complex Analog Circuits
abstract
Despite the effort of analog circuit design automation, currently complex analog circuit design still requires extensive manual iterations, making it labor intensive and time-consuming. Recently, reinforcement learning (RL) algorithms have been demonstrated successfully for the analog circuit design optimization. However, a robust and highly efficient RL method to design analog circuits with complex design space has not been fully explored yet. In this work, inspired by multiagent planning theory as well as human expert design practice, we propose a multiagent based RL (MA-RL) framework to tackle this issue. Particularly, we (i) partition the complex analog circuits into several sub-blocks based on topology information and effectively reduce the complexity of design search space; (ii) leverage MA-RL for the circuit optimization, where each agent corresponds to a single sub-block, and the interactions between agents delicately mimic the best design tradeoffs between circuit sub-blocks by human experts; (iii) introduce and compare three different multiagent RL algorithms and corresponding frameworks to demonstrate the effectiveness of the MA-RL method. (iv) employing twin-delayed techniques and proximal policy to further boost training stability and accomplish higher performances. (v) The impacts of different reward function definitions as well as different state settings of MA-RL agents are investigated to further improve the robustness of this framework. (vi) Experiments on three different complex analog circuit topologies (GBA, DLL and SAR ADC) and knowledge transfers between two technology nodes are demonstrated. It’s shown that MA-RL framework can achieve the best FoM for complex analog circuits’ design. This work shines the light for future large scale analog circuit system design automation.
Jiarui Bao, Zhangcheng Huang 0001, Zhaori Bi, Xingwei Feng, Xuan Zeng 0001, Ye Lu 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Automated Design of Complex Analog Circuits with Multiagent based Reinforcement Learning
abstract
Despite the effort of analog circuit design automation, currently complex analog circuit design still requires extensive manual iterations, making it labor intensive and time-consuming. Recently, reinforcement learning (RL) algorithms have been demonstrated successfully for the analog circuit design optimization. However, a robust and highly efficient RL method to design analog circuits with complex design space has not been fully explored yet. In this work, inspired by multiagent planning theory as well as human expert design practice, we propose a multiagent based RL (MA-RL) framework to tackle this issue. Particularly, we (i) partition the complex analog circuits into several sub-blocks based on topology information and effectively reduce the complexity of design search space; (ii) leverage MA-RL for the circuit optimization, where each agent corresponds to a single sub-block, and the interactions between agents delicately mimic the best design tradeoffs between circuit sub-blocks by human experts; (iii) introduce the multiagent twin-delayed techniques to further boost training stability and accomplish higher performances. Experiments on two different analog circuit topologies and knowledge transfers between two technology nodes are demonstrated. It’s shown that MA-RL framework can achieve the best FoM for complex analog circuits design. This work shines the light for future large scale analog circuit system design automation.
Jiarui Bao, Zhangcheng Huang 0001, Xuan Zeng 0001, Ye Lu 0005
DAC3
2023 TiPU: A Spatial-Locality-Aware Near-Memory Tile Processing Unit for 3D Point Cloud Neural Network
abstract
Energy-efficient 3D point cloud neural network accelerators are desired for autonomous driving and AR/VR applications. This paper proposes TiPU, a spatial-locality-aware near-memory tile processing unit where the point clouds are partitioned into tiles to process spatial features locally. Intra-tile farthest point sampling and cross-tile neighbor search are employed to avoid unnecessary distance computing. To efficiently facilitate the tile operations, TiPU architecture consists of a tile-based unified distance computing unit, a near-CAM feature extractor, and a near-SRAM-computing MLP engine. The experimental results show that, compared to GPU implementation, TiPU achieves 15.7× processing speed and reduces 7308× energy consumption.
Jiapei Zheng, Hao Jiang 0024, Xinkai Nie, Zhangcheng Huang 0001, Chixiao Chen, Qi Liu 0010
DAC4
2023 A 10b 1.25GS/s Residue Post-Amplified Pipelined-SAR ADC with Supply-and-Temperature Stabilized Open-Loop Residue Amplifier
abstract
This paper presents a single-channel two-stage pipelined-SAR ADC with ping-pong switched half of capacitor-digital-to-analog-converter (CDAC) moving the residue amplification into 2ndstage to lighten the timing burden of the 1ststage. Besides, a dynamic open-loop residue amplifier (RA) is employed to improve the energy efficiency and amplification speed. In order to enhance the ADC robustness, the gain variation under the supply and temperature drift are suppressed through the supply-and-temperature compensation bias. The simulation results shows gain variation is less than 5% under the temperature range from -20 °C to 100 °C and supply voltage variation of$\pm \mathbf{5}\%$, which ensure the above 56 dB ADC SNDR under process-voltage-temperature (PVT) variation. The ADC is simulated with a 40 nm CMOS process and 1V supply voltage, it achieves 59 dB SNDR with Nyquist input frequency at 1.25 GS/s sampling rate. The power consumption is 5.6mW, leading to a Walden FoM (FoMw) of 6.2 fJ/conversion-step and a Schreier FoM (FoMs) of 169.5 dB.
Maosong Shi, Zhangcheng Huang 0001, Chixiao Chen, Wenning Jiang
ISCAS5
2023 Noise Model of Large-Format Readout Integrated Circuit for Infrared Focal Plane Array
abstract
With the upscaling of the pixel format and the downscaling of the pixel pitch of infrared focal plane arrays (FPAs), the demand for ultra-low-noise design is increasing drastically. Existing methods of noise modelling lack transistor-level analysis and are unable to provide directly effective guidance for circuit design because of extremely complicated mathematical expressions. In this work, a comprehensive noise model of large-format readout integrated circuit (ROIC) is presented, for the first time elucidating the noise mechanism with a full research framework, including transistor-, circuit-, block-, and chip-level analysis. By utilizing a novel approximation method, simplified analytical formulas with high accuracy are obtained for flicker noise and other noise. These simplified formulas clear up the confusions of how the various functional circuit blocks influence noise performance. Moreover, based on the presented noise model, the impacts of crucial parameters on FPA noise are analyzed in detail, and strategies regarding the trade-off between the signal-to-noise ratio (SNR) and other aspects of performance are proposed for low-noise design. Taking an astronomical application with a very long integration time as an example, the calculated results show that an ultralow noise of${5}{\,e^{-}}$can be achieved by leveraging the above approaches.
Zhangcheng Huang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.1