Chieh-Fang Teng

dblp:209/6915 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0002-3965-842XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Device Noises Resilient Training and Inference Framework for Smart Sensing on Analog Computing In Memory
abstract
Neural networks have demonstrated superior performance over rule-based and model-based approaches in processing noisy sensing data. However, their substantial computational and energy demands hinder deployment in battery-powered embedded systems. Computing-in-Memory (CIM) devices offer a promising alternative by significantly reducing energy consumption. Prior work [2] achieves this by leveraging a non-von Neumann architecture, which minimizes data movement between memory and compute units, thereby mitigating the memory wall bottleneck. Despite these advantages, analog CIM (ACIM) systems face several key challenges, including analog noise, limited numerical precision, and increased hardware complexity. While cloud-based neural networks are still dominant, emerging applications increasingly demand real-time, privacy-preserving inference on-device. For instance, facial authentication requires local execution on edge devices to ensure low-latency responsiveness and to protect user privacy. CIM architectures are particularly well-suited to these scenarios due to their tightly integrated memory-compute structure, offering low-latency and energy-efficient inference capabilities.
Xin-You Liu, Chi-Sheng Shih 0001, Tsung-Te Liu, Pei-Kuei Tsung, Chih-Wei Chen, Chieh-Fang Teng
CASES6
2024 A 40nm 24.6TOPS/W Scalable EfficientDet Processor for Object Detection
abstract
Object detection is a crucial technology used to identify and locate objects in a wide range of applications. Google’s EfficientDet, a scalable solution, employs a compound scaling method to systematically adjust the network’s depth, width, and input resolution, meeting different resource constraints on edge devices. This paper presents the first dedicated processor for EfficientDet, featuring three key elements: 1) an adaptive channel/input-wise (CIW) mapper to improve hardware utilization by applying distinct mapping strategies for layers with varying data shapes, 2) a tri-mode activation compression (AC) engine to reduce external memory access (EMA) by leveraging the sparsity level of activations, and 3) a unified aggregation core (AggrCore) to flexibly handle different computations. The chip is fabricated using TSMC 40nm CMOS technology and achieves a maximum energy efficiency of 24.6TOPS/W. Compared to the state-of-the-art object detection processor, our chip demonstrates 3.7× and 2.1× improvements in energy and area efficiencies, respectively.
Yu-Chuan Chuang, Ming-Guang Lin, Chi-Tse Huang, Chieh-Fang Teng, Cheng-Yang Chang, Yi-Ta Chen, An-Yeu Wu
ISCAS4
2023 Joint Optimization of Dimension Reduction and Mixed-Precision Quantization for Activation Compression of Neural Networks
abstract
Recently, deep convolutional neural networks (CNNs) have achieved eye-catching results in various applications. However, intensive memory access of activations introduces considerable energy consumption, resulting in a great challenge for deploying CNNs on resource-constrained edge devices. Existing research utilizes dimension reduction (DR) and mixed-precision (MP) quantization separately to reduce computational complexity without paying attention to their interaction. Such naïve concatenation of different compression strategies ends up with suboptimal performance. To develop a comprehensive compression framework, we propose an optimization system by jointly considering DR and MP quantization, which is enabled by independent groupwise learnable MP schemes. Group partitioning is guided by a well-designed automatic group partition mechanism that can distinguish compression priorities among channels, and it can deal with the tradeoff between model accuracy and compressibility. Moreover, to preserve model accuracy under low bit-width quantization, we propose a dynamic bit-width searching technique to enable continuous bit-width reduction. Our experimental results show that the proposed system reaches 69.03%/70.73% with average 2.16/2.61 bits per value on Resnet18/MobileNetV2, while introducing only approximately 1% accuracy loss of the uncompressed full-precision models. Compared with individual activation compression schemes, the proposed joint optimization system reduces 55%/9% (−2.62/−0.27 bits) memory access of DR and 55%/63% (−2.60/−4.52 bits) memory access of MP quantization, respectively, on Resnet18/ MobileNetV2 with comparable or even higher accuracy.
Yu-Shan Tai, Cheng-Yang Chang, Chieh-Fang Teng, Yi-Ta Chen, An-Yeu Wu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2022 Compression-Aware Projection with Greedy Dimension Reduction for Convolutional Neural Network Activations
abstract
Convolutional neural networks (CNNs) achieve remarkable performance in a wide range of fields. However, intensive memory access of activations introduces considerable energy consumption, impeding deployment of CNNs on resource-constrained edge devices. Existing works in activation compression propose to transform feature maps for higher compressibility, thus enabling dimension reduction. Nevertheless, in the case of aggressive dimension reduction, these methods lead to severe accuracy drop. To improve the trade-off between classification accuracy and compression ratio, we propose a compression-aware projection system, which employs a learnable projection to compensate for the reconstruction loss. In addition, a greedy selection metric is introduced to optimize the layer-wise compression ratio allocation by considering both accuracy and #bits reduction simultaneously. Our test results show that the proposed methods effectively reduce 2.91×~5.97× memory access with negligible accuracy drop on MobileNetV2/ResNet18/VGG16.
Yu-Shan Tai, Chieh-Fang Teng, Cheng-Yang Chang, An-Yeu Wu
ICASSP2
2021 Convolutional Neural Network-Aided Bit-Flipping for Belief Propagation Decoding of Polar Codes
abstract
Known for their capacity-achieving abilities, polar codes have been selected as the control channel coding scheme for 5G communications. To satisfy high throughput and low latency, belief propagation (BP) is ideal as the decoding algorithm due to its nature of parallel processing. However, the error performance of BP is in general worse than that of enhanced successive cancellation (SC). Recently, bit-flipping (BF) mechanism is applied to BP decoding to lower the error rate. However, its trial-and-error process results in longer latency. In this work, we propose a convolutional neural network-aided bit-flipping (CNN-BF) mechanism to further enhance BP decoding. With carefully designed input data and model architecture, our proposed CNN-BF can achieve better error correction capability with less flipping attempts than prior works. It also achieves a lower block error rate (BLER) than SC list (SCL).
Chieh-Fang Teng, Andrew Kuan-Shiuan Ho, Chen-Hsi Derek Wu, Sin-Sheng Wong, An-Yeu Wu
ICASSP1
2021 A 7.8-13.6 pJ/b Ultra-Low Latency and Reconfigurable Neural Network-Assisted Polar Decoder With Multi-Code Length Support
abstract
Polar codes have been officially selected as the channel coding in 5G standard. To meet the requirements of enhanced mobile broadband (eMBB), most published polar decoder chips aim to improve throughput rate and error-correction performance. However, to meet with the requirements of another two 5G new radio (NR) application scenarios, ultra-reliable low-latency communications (URLLC), and massive machine-type communications (mMTC), the design features of low latency and energy efficiency are also desirable. In this article, we present a 7.8-13.6 pJ/b ultra-low latency and energy-efficient polar decoder fabricated in 40nm CMOS technology. By adopting the decoding algorithm of recurrent neural network-assisted belief propagation (RNN-BP), the learned scaling parameters can improve the convergence rate by 8 times with reasonable hardware and memory overhead. Then, by taking advantage of BP's regular structure, we propose a fully-reconfigurable RNN-BP decoder architecture to support multiple code lengths with negligible hardware complexity. It contributes to 2- 8× improved hardware utilization rate while providing a flexible adjustment between throughput and error-correction performance. At the architectural level, two optimization techniques for the design of the processing element (PE) are proposed to jointly reduce the chip's area and power by 73% and 67%, respectively. From the measurement results, our reconfigurable RNN-BP polar decoder chip has 2.3×, 2.3×, and 10.0× enhancement over prior designs in terms of latency, throughput rate, and energy efficiency. Consequently, our reconfigurable design has great potential to meet various 5G NR applications.
Chieh-Fang Teng, An-Yeu Wu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2020 Low-Complexity LSTM-Assisted Bit-Flipping Algorithm For Successive Cancellation List Polar Decoder
abstract
Polar codes have attracted much attention in the past decade due to their capacity-achieving performance. The higher decoding capacity is required for 5G and beyond 5G (B5G). Although the cyclic redundancy check (CRC)- assisted successive cancellation list bit-flipping (CA-SCLF) decoders have been developed to obtain a better performance, the solution to error bit correction (bitflipping) problem is still imperfect and hard to design. In this work, we leverage expert knowledge in communication systems and adopt deep learning (DL) techniques to obtain a better solution. A low-complexity long short-term memory network (LSTM)-assisted CASCLF decoder is proposed to further improve the performance of conventional CA-SCLF and avoid complexity and memory overhead. Our test results show that we can effectively improve the BLER performance by 0.11dB compared to prior work and reduce the complexity and memory overhead by over 30% of the network.
Chun-Hsiang Chen, Chieh-Fang Teng, An-Yeu Wu
ICASSP2
2019 Low-complexity Recurrent Neural Network-based Polar Decoder with Weight Quantization Mechanism
abstract
Polar codes have drawn much attention and been adopted in 5G New Radio (NR) due to their capacity-achieving performance. Recently, as the emerging deep learning (DL) technique has breakthrough achievements in many fields, neural network decoder was proposed to obtain faster convergence and better performance than belief propagation (BP) decoding. However, neural networks are memory-intensive and hinder the deployment of DL in communication systems. In this work, a low-complexity recurrent neural network (RNN) polar decoder with codebook-based weight quantization is proposed. Our test results show that we can effectively reduce the memory overhead by 98% and alleviate computational complexity with slight performance loss.
Chieh-Fang Teng, Chen-Hsi Derek Wu, Andrew Kuan-Shiuan Ho, An-Yeu Wu
ICASSP1