Po-Tsang Huang

dblp:12/713 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-8679-2755ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 4 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Energy-Efficient All-Digital Computation-in-Memory Macro Design with CFeFET-based Booth Decoder
Ting-Wei Yang, Yuan-Yu Huang, Pin Su, Po-Tsang Huang
ISCAS5
2026 Defect-based Testing for SRAM Address Decoders
Ho-Jie Hsu, Hsien-Chen Lee, Chun-Yu Shen, Po-Tsang Huang, Shih-Chieh Lin, Yung-Jheng Wang, Ying-Yen Chen, Chien-Yuan Pao, Hung-Yu Lee, Mango Chia-Tso Chao
VTS4
2025 Irregular Operation-Unit-based Compression for Non-Volatile Computation-in-Memory Accelerator
abstract
Irregular pruning significantly enhances the sparsity of deep neural networks (DNNs) and reduces computational demands. However, the irregular sparse weight matrices can diminish the benefits of pruning when implemented in non-volatile computational-in-memory (CIM) accelerators, which relay on dense matrix-vector multiplications. In this work, we propose an irregular operation-unit-based (OU-based) compression method for nonvolatile CIM. We transform the compression problem into a fixed-size clustering task by clustering zero-columns through the rearrangement of matrix rows, followed by their elimination. This approach significantly reduces the number of required non-volatile memory (NVM) macros by up to 2.3x, while compacting the irregular non-zero data for efficient mapping onto non-volatile CIM. We achieve a compression ratio of up to 89% for irregularly pruned weights. Furthermore, we evaluate the mapping of the compacted weights onto nonvolatile CIM accelerator with OU size of 2x128 and 4x128. This work can achieve area efficiency of 0.166 TOPS/mm2and improve area efficiency up to 3.1x compare to state-of-art works.
Liang-Te Huang, Hung-Ming Chen, Po-Tsang Huang
ISCAS4
2025 An Energy-Area-Efficient 3D Interleaved-Ring Accelerator With INT/FP Pipelined PE Array and 3D-SRAM Cube for On-Device CNN Training and DM Inference
abstract
Deep neural networks (DNNs) have demonstrated exceptional performance in image-related artificial intelligence (AI) applications. However, the inference of generative models, such as Diffusion Models (DMs), and the training of convolutional neural networks (CNNs) are computationally intensive tasks, requiring extensive floating-point (FP) operations to maintain accuracy and high-quality results. These tasks are also memory-bound, generating large intermediate data that result in significant external memory access (EMA), thus complicating the deployment of image-related DNNs on edge devices. While 3D-stacked SRAM with through-silicon via (TSV) technology offers promising solutions to alleviate EMA, the complexities of 3D interconnect architectures can introduce substantial overhead in intra-chip communication, potentially degrading overall efficiency. In this paper, we propose a flexible 3D interconnect architecture, termed the 3D interleaved-ring, which utilizes multiple 3D interleaved rings to connect the pipelined integer (INT) and floating-point (FP) processing element (PE) arrays with a 3D-SRAM cube, effectively mitigating the overhead caused by 3D interconnection. We design$3\times 3$micro-routers with dedicated channels and embedded adders for the proposed 3D interconnection architecture. This 3D ring-based architecture reduces the power consumption of on-chip data movement by a factor of 4.5 and decreases the number of required TSVs by a factor of 4.2, compared to conventional 3D mesh-based interconnects. Additionally, we introduce efficient mixed-bit precision dataflows that incorporates dynamic workload distribution to optimize data reuse, reduce bandwidth demands on the 3D-SRAM cube, and improve PE utilization. The proposed work achieves over 90% PE utilization and reduces DRAM accesses by more than$7\times $across various DNN models. Overall, the proposed 3D accelerator improves energy and area efficiency by up to$8.4\times $and$3.4\times $, respectively, compared to state-of-the-art DNN accelerators and processors.
Hung-Ming Chen, Po-Tsang Huang
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 A 28nm Energy-Area-Efficient Row-based pipelined Training Accelerator with Mixed FXP4/FP16 for On-Device Transfer Learning
abstract
Training deep convolutional neural networks (DNNs) requires significantly more computational capacity, complex dataflow, memory accesses, and data movement among processing elements (PEs), as well as higher bit precision for back propagation (BP), which demands more power and area overhead than DNN inference. For mobile/edge devices, energy and area efficiency are critical concerns. This research proposes a row-based pipelined DNN training accelerator that employs three techniques to improve energy and area efficiency for resource-constrained edge/mobile devices. The first technique involves freezing weight updates in convolution and batch normalization layers. The second technique involves decomposing the simulated quantization for convolutional layers and reorganizing the operations of batch normalization layers. The mathematical demonstration shows that FP convolution operations can be completed using fixed point (FXP) calculations. FXP MACs with dequantizer can replace the original FP MACs for convolutional layers. Additionally, a row-based FXP/FP pipelined training accelerator is designed for layers pipeline, convolution, and batch normalization layers to increase the FXP and FP resource utilization. The third method uses multi-bank buffer management to prevent data conflicts and reduce the need for on-chip buffers by up to 3.5 times. The proposed accelerator was implemented using the TSMC 28nm CMOS process and achieved an energy efficiency of 2.19 TFLOPS/W and an area efficiency of 85.32 GFLOPS/mm2. It outperforms state-of-the-art works with 6.8 times the area efficiency and 3.7 times the energy efficiency.
Han-Hsiang Pei, Jheng-Rong Yu, Hung-Ming Chen, Po-Tsang Huang
ISCAS5
2023 Reshaping System Design in 3D Integration: Perspectives and Challenges
abstract
In this paper, we depict modern system design methodologies via 3D integration along with the advance of packaging, considering system prototyping, interconnecting, and physical implementation. The corresponding challenges are presented as well.
Hung-Ming Chen, Chu-Wen Ho, Shih-Hsien Wu, Po-Tsang Huang, Hao-Ju Chang, Chien-Nan Jimmy Liu
ISPD5
2021 Rotational motion-aware beam refinement for high-throughput mmWave communications
Tourangbam Harishore Singh, Shabirahmed Badashasab Jigalur, Po-Tsang Huang
Wirel. Networks3
2019 A 7.5-mW 10-Gb/s 16-QAM wireline transceiver with carrier synchronization and threshold calibration for mobile inter-chip communications in 16-nm FinFET
abstract
A compact energy-efficient 16-QAM wireline transceiver with carrier synchronization and threshold calibration is proposed to leverage high-density fine-pitch interconnects. Utilizing frequency-division multiplexing, the transceiver transfers four-bit data through one RF band to reduce intersymbol interferences. A forwarded clock is also transmitted through the same interconnect with the data simultaneously to enable low-power PVT-insensitive symbol clock recovery. A carrier synchronization algorithm is proposed to overcome nontrivial current and phase mismatches by including DC offset calibration and dedicated I/Q phase adjustments. Along with this carrier synchronization, a threshold calibration process is used for the transceiver to tolerate channel and circuit variations. The transceiver implemented in 16-nm FinFET occupies only 0.006-mm2 and achieves 10 Gb/s with 0.75-pJ/bit efficiency and <2.5-ns latency.
Jieqiong Du, Chien-Heng Wong, Yo-Hao Tu, Wei-Han Cho, Yilei Li, Yuan Du, Po-Tsang Huang, Sheau Jiung Lee, Mau-Chung Frank Chang
NOCS7
2018 SMEM++: A Pipelined and Time-Multiplexed SMEM Seeding Accelerator for DNA Sequencing
abstract
The advent of next-generation sequencing has made a great impact on many applications from precision medicine to new drug discovery, leading to an explosion in sequencing of individual genomes. This motivates the research of FPGA acceleration for genome sequencing algorithms to complement the computation capabilities of conventional CPU systems. The recently developed SMEM seeding algorithm, which is based on FMD-index, becomes a time-consuming computation kernel in genome sequencing, but it has not been well studied. The fundamental challenge of accelerating the SMEM algorithm is to handle its large volume of random memory accesses. While the state-of-the-art SMEM accelerator attempts to achieve high memory bandwidth by sacrificing the performance of individual processing elements to maximize the task-level parallelism, this design methodology suffers serious inefficiency of resource utilization and does not scale well for future technology advances. To resolve these impediments, we propose SMEM++, a pipelined and time-multiplexed FPGA accelerator for the SMEM algorithm. SMEM++ features a fully pipelined processing element design that significantly improves the efficiency of FPGA on-chip resource utilization. Moreover, we design a communication interface adapter to make the accelerator compatible to the designated CPU-FPGA platform, increasing its portability. Our experiments on the Intel HARPv2 platform show that SMEM++ outperforms CPU by 24x, and outperforms the state-of-the-art SMEM accelerator design by 6.3x, even with 43% less logic resource consumption.
Jason Cong, Licheng Guo, Po-Tsang Huang, Peng Wei 0004, Tianhe Yu
FCCM3
2018 SMEM++: A Pipelined and Time-Multiplexed SMEM Seeding Accelerator for Genome Sequencing
abstract
Next-generation sequencing motivates the researchof FPGA acceleration for genome sequencing algorithms. Therecently developed quadratic-time SMEM seeding algorithmbecomes a time-consuming computation kernel in genomesequencing, but it has not been well studied. The fundamentalchallenge of accelerating the SMEM algorithm is to handle itslarge volume of random memory accesses. While the state-ofthe-art SMEM accelerator attempts sacrifices the performanceof individual processing elements to maximize the task-levelparallelism, this methodology suffers a serious resource underutilizationissue. Therefore, we propose SMEM++, a pipelinedand time-multiplexed FPGA accelerator for SMEM algorithm.SMEM++ adopts the canonical non-blocking pipelinemethodology and implements a fully pipelined acceleratorwith initiation interval equal to one. Moreover, we designa communication interface adapter to make the acceleratorcompatible to the target platform interface and increase itsportability. Experiments on the Intel HARPv2 platform showthat SMEM++ outperforms the original software by 24x, andoutperforms the state-of-the-art SMEM accelerator design by6.3x, with 43% less logic resource usage.
Jason Cong, Licheng Guo, Po-Tsang Huang, Peng Wei 0004, Tianhe Yu
FPL3
2017 Exploration and evaluation of low-dropout linear voltage regulator with FinFET, TFET and hybrid TFET-FinFET implementations
abstract
This paper investigates and evaluates analog and digital low-dropout linear voltage regulators (LDO) with FinFET, TFET and hybrid TFET-FinFET implementations. We utilize Sentaurus physics-based atomistic 3D TCAD mixed-mode simulations for device characteristics and HSPICE with look-up tables based on Verilog-A models calibrated with TCAD simulation results. Frequency response, load regulation and power supply rejection ratio (PSRR) are evaluated for analog LDOs under low, medium and high bias-current conditions. The results indicate that for analog implementations, TFET-LDO and hybrid-LDO provide better loop-gain and PSRR than FinFET-LDO under low and medium operating currents, whereas at higher operating current, FinFET implementation would outperform. As operating voltage is reduced, the performances of analog implementations degrade, and digital implementations become favorable for VIN below around 0.55V. We further show that for digital LDO, all FinFET implementation provides superior performance over all TFET and hybrid TFET-FinFET implementations.
Chia-Ning Chang, Yin-Nien Chen, Po-Tsang Huang, Pin Su, Ching-Te Chuang
ISCAS3
2017 An implantable 128-channel wireless neural-sensing microsystem using TSV-embedded dissolvable μ-needle array and flexible interposer
abstract
For implanted neural-sensing devices, one of the remaining challenges is to transmit stable power/data (P/D) transmission for high spatiotemporal resolution neural data. This paper presents a miniaturized implantable 128-channel wireless neural-sensing microsystem using TSV-embedded dissolvable μ-needle array, a flexible interposer and 4 dies by 2.5D/3D TSV heterogeneous SiP technology. The 4 dies are 2 neural-signal acquisition ICs implemented by 90nm CMOS, 1 neural-signal processor by 40nm CMOS and 1 wireless P/D transmission circuitry by 0.18μm CMOS. Thus, the proposed wireless microsystem realizes 128-channel neural-signal sensing within the area of 5mm × 5mm, neural feature extraction and wireless P/D transmission using an on-interposer inductor. The overall average power of the circuits in this microsystem is only 9.85mW.
Po-Tsang Huang, Yu-Chieh Huang, Shang-Lin Wu, Yu-Chen Hu, Ming-Wei Lu, Ting-Wei Sheng, Fung-Kai Chang, Chun-Pin Lin, Nien-Shang Chang, Hung-Lieh Chen, Chi-Shi Chen, Jeng-Ren Duann, Tzai-Wen Chiu, Wei Hwang, Kuan-Neng Chen, Ching-Te Chuang, Jin-Chern Chiou
ISCAS1
2016 The SMEM Seeding Acceleration for DNA Sequence Alignment
abstract
The advance of next-generation sequencing technology has dramatically reduced the cost of genome sequencing. However, processing and analyzing huge amounts of data collected from sequencers introduces significant computation challenges, these have become the bottleneck in many research and clinical applications. For such applications, read alignment is usually one of the most compute-intensive steps. Billions of reads generated from the sequencer need to be aligned to the long reference genome. Recent state-of-the-art software read aligners follow the seed-andextend model. In this paper we focus on accelerating the first seeding stage, which generates the seeds using the supermaximal exact match (SMEM) seeding algorithm. The two main challenges for accelerating this process are 1) how to process a huge number of short reads with high throughput, and 2) how to hide the frequent and long random memory access when we try to fetch the value of the reference genome. In this paper, we propose a scalable array-based architecture, which is composed by many processing engines (PEs) to process large amounts of data simultaneously for the demand of high throughput. Furthermore, we provide a tight software/hardware integration that realizes the proposed architecture on the Intel-Altera HARP system. With a 16-PE accelerator engine, we accelerate the SMEM algorithm by 4x, and the overall SMEM seeding stage by 26% when compared with 16-thread CPU execution. We further analyze the performance bottleneck of the design due to extensive DRAM accesses and discuss the possible improvements that are worthwhile to be explored in the future.
Mau-Chung Frank Chang, Yuting Chen 0003, Jason Cong, Po-Tsang Huang, Chun-Liang Kuo, Cody Hao Yu
FCCM4
2016 An ultra-high-density 256-channel/25mm2 neural sensing microsystem using TSV-embedded neural probes
abstract
Highly integrated neural sensing microsystems are crucial to capture accurate signals for brain function investigations. In this paper, a 256-channel/25 mm2 neural sensing microsystem is presented based on through-silicon-via (TSV) 2.5D integration. This microsystem composes of dissolvable μ-needles, TSV-embedded μ-probes, 256-channel neural amplifiers, 11-bit area-power-efficient SAR ADCs and serializers. Based on the dissolvable μ-needles and TSV 2.5D integration, this microsystem can detect 256 ECoG/LFP signals within the small area of 5mm × 5mm. Additionally, the neural amplifier realizes 57.8dB gain with only 9.8μW for each channel, and the 9.7-bit ENOB of the SAR ADC at 32kS/s can be achieved with 0.42μW and 0.036 mm2. The overall power of this microsystem is only 3.79mW for 256-channel neural sensing.
Yu-Chieh Huang, Po-Tsang Huang, Shang-Lin Wu, Yu-Chen Hu, Yan-Huei You, Yan-Yu Huang, Hsiao-Chun Chang, Yen-Han Lin, Jeng-Ren Duann, Tzai-Wen Chiu, Wei Hwang, Kuan-Neng Chen, Ching-Te Chuang, Jin-Chern Chiou
ISCAS2
2014 Energy-efficient configurable discrete wavelet transform for neural sensing applications
abstract
Highly integrated neural sensing microsystems are crucial to capture accurate signals for brain function investigations. In this paper, an energy-efficient configurable lifting-based discrete wavelet transform (DWT) is proposed for a high-density neural sensing microsystems to extract the features of neural signals by filtering the signals into different frequency bands. Based on the lifting-based DWT algorithm, the area and power consumption can be reduced by decreasing the computation circuits. Additionally, both the time window and mother wavelets can be adjusted via the configurable datapth. Moreover, the power-gating and clock-gating techniques are utilized to further reduce the energy consumption for the energy-limited bio-systems. The proposed configurable DWT is designed and implemented using TSMC 65nm CMOS low power process with total area of 0.11 mm2and power consumption of 26 μW. Moreover, this proposed DWT is also implemented in Lattice MachXO2-1200 FPGA and integrated in a 2.5D heterogeneously integrated high-density neural-sensing microsystem with the power consumption of 211.2 μW.
Tang-Hsuan Wang, Po-Tsang Huang, Kuan-Neng Chen, Jin-Chern Chiou, Kuo-Hua Chen, Chi-Tsung Chiu, Ho-Ming Tong, Ching-Te Chuang, Wei Hwang
ISCAS2
2012 Substrate noise suppression technique for power integrity of TSV 3D integration
abstract
In this paper, a substrate noise suppression technique is proposed for the power integrity of TSV 3D integrations. This substrate noise suppression technique reduces both substrate and TSV coupling noises using active substrate decouplers (ASDs) to absorb the substrate noise current. Additionally, the ASD placing is also presented to suppress noises effectively for different 3D structures. For a processor-memory stacking integration, the ground bouncing noises can be reduced by 44.1% via the noise suppression technique. The proposed substrate noise suppression technique can enhance the power integrity of TSV 3D-ICs by reducing the coupling substrate noises.
Po-Jen Yang, Po-Tsang Huang, Wei Hwang
ISCAS2
2008 A 5.2mW all-digital fast-lock self-calibrated multiphase delay-locked loop
abstract
A 333MHz-1GHz all-digital multiphase delay-locked loop with precise multi-phase output has been designed with TSMC 130nm CMOS technology model. A modified binary search algorithm is proposed to match up a linear approximate delay element (LADE). The LADE property of linearity and insensitive to PVT variations is good for digitally-controlled delay element. The lock-in time could be reduced down to 14 reference clock cycles, and enhance the operation range based on LADE/binary search algorithm co-operate effort. The timing error caused by process mismatch is further reduced by proposed rapid self-calibration (RSC) algorithm. A calibration unit is designed based on RSC algorithm, which reduces the maximum timing error to less than 9ps when DLL is operating at 500MHz. The entire calibration unit could be turned off after calibration procedure is complete to reduce power consumption. The total power dissipation of the all-digital self-calibrated multiphase delay-locked loop is 5.2mW at 1GHz with a 1.2V power supply.
Li-Pu Chuang, Ming-Hung Chang, Po-Tsang Huang, Chih-Hao Kan, Wei Hwang
ISCAS3
2008 "Green" micro-architecture and circuit co-design for ternary content addressable memory
abstract
In this paper, an energy-efficient and high performance ternary content addressable memory (TCAM) are presented. It employs the concept of "green" microarchitecture and circuit co-design. For achieving energy-efficient TCAM architecture, hierarchy search-line scheme and butterfly match-line scheme are proposed. Moreover, the match-lines are also implemented by noise-tolerant XOR-based conditional keeper and don't-care based power gating scheme to reduce not only search time but power consumption. In order to reduce increasing leakage power with advanced technologies, furthermore, the proposed TCAM design employs super cut-off power gating technique and multi-mode data-retention power gating technique to reduce leakage currents without reducing search time and destroying noise margin. An energy-efficient 256times144 TCAM array is implemented in TSMC 0.13 um and designed in 65 nm Berkeley Predictive Technology Model, respectively. The simulation results show the leakage power reduction is 70.7% and energy metric of TCAM macro is 0.047 fJ/bit/search.
Po-Tsang Huang, Shu-Wei Chang, Wen-Yen Liu, Wei Hwang
ISCAS1
2008 Low Power and Reliable Interconnection with Self-Corrected Green Coding Scheme for Network-on-Chip
Po-Tsang Huang, Wei-Li Fang, Wei Hwang
NOCS1
2006 2-level FIFO architecture design for switch fabrics in network-on-chip
abstract
The network-on-chip (NoC) architecture provides the integrated solution for system-on-chip (SoC) design. The buffer architecture and sizes, however, dominate the performance of NoC and influence on the design of arbiters in the switch fabrics. The 2-level FIFO architecture is proposed. It simplifies the design of the arbitration algorithm and gets better performance than other buffer architectures without increasing the buffer sizes. The concept of the shared memory mechanism and multiple accesses for the buffers are developed. The FIFO architecture is implemented and simulated with TSMC 0.13/spl mu/m network-on-chip by HSPICE and Verilog. The operation frequency of the 2-level FIFO reaches 400MHz.
Po-Tsang Huang, Wei Hwang
ISCAS1