Ming-Guang Lin

dblp:322/6367 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0001-7748-1189ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 WS-CIM: Enabling Fast and Simultaneous Update for Multi-Macro Compute-in-Memory Architecture Using Weight Sharing Technique
abstract
Compute-in-memory CIM) architecture has been widely studied to accelerate deep neural networks (DNNs). CIM improves energy and area efficiency by performing multiply-accumulate (MAC) operations within memory array. However, the limited size of individual CIM macros requires frequent weight updates for large DNN models, leading to significant latency and energy overhead. In this paper, we propose WS-CIM, a novel framework to enable fast and simultaneous weight updates for multi-macro CIM architectures. Specifically, WS-CIM adopts fine-grained weight sharing technique considering the CIM architecture to minimize redundant write operations, guided by a distance loss function to maintain accuracy. A workload balance mechanism is further introduced for multi-macro architecture to prevent bottlenecks from any single macro during simultaneous weight updates. Experiments on ResNet50 and DeiT-B show that WS-CIM achieves 34.8% latency reduction with 1.29% accuracy loss for ResNet50 and 35.8% latency reduction with 0.70% accuracy degradation for DeiT-B, respectively. These results demonstrate the scalability and efficiency of WS-CIM for deploying DNNs on advanced CIM hardware platforms.
Yan-Ding Shieh, Ming-Guang Lin, Hung-Yu Wang, An-Yeu Wu
ISCAS2
2024 A 40nm 24.6TOPS/W Scalable EfficientDet Processor for Object Detection
abstract
Object detection is a crucial technology used to identify and locate objects in a wide range of applications. Google’s EfficientDet, a scalable solution, employs a compound scaling method to systematically adjust the network’s depth, width, and input resolution, meeting different resource constraints on edge devices. This paper presents the first dedicated processor for EfficientDet, featuring three key elements: 1) an adaptive channel/input-wise (CIW) mapper to improve hardware utilization by applying distinct mapping strategies for layers with varying data shapes, 2) a tri-mode activation compression (AC) engine to reduce external memory access (EMA) by leveraging the sparsity level of activations, and 3) a unified aggregation core (AggrCore) to flexibly handle different computations. The chip is fabricated using TSMC 40nm CMOS technology and achieves a maximum energy efficiency of 24.6TOPS/W. Compared to the state-of-the-art object detection processor, our chip demonstrates 3.7× and 2.1× improvements in energy and area efficiencies, respectively.
Yu-Chuan Chuang, Ming-Guang Lin, Chi-Tse Huang, Chieh-Fang Teng, Cheng-Yang Chang, Yi-Ta Chen, An-Yeu Wu
ISCAS2
2024 Similarity-Aware Fast Low-Rank Decomposition Framework for Vision Transformers
abstract
Vision transformers (ViTs) have shown success in computer vision tasks. However, their high computational and memory demands limit on-device implementation. Low-rank decomposition (LRD) is widely acknowledged as an effective method for reducing computational complexity. Nevertheless, previous approaches relying on heuristic-based or network architecture search (NAS) methods consume extensive searching and training time to uncover the optimal ranks. To efficiently discover rank distributions, this paper introduces a Similarity-Aware Fast-LRD framework, leveraging a greedy selection metric to optimize the compression ratio within each weight matrix by considering cosine similarity and rank reduction. Notably, our proposed framework restores accuracy in 15-25 epochs of model fine-tuning, surpassing the 30-300 epochs needed in prior work. Furthermore, our automated rank search process consumes fewer than 4 GPU hours. Experimental results show that our proposed Similarity-Aware Fast-LRD framework reduces FLOPs and associated parameters by 49.6% and 38.9% for DeiT-B and Swin-S, respectively, with merely 0.83% and 0.72% accuracy degradation on ImageNet.
Yuan-June Luo, Yu-Shan Tai, Ming-Guang Lin, An-Yeu Wu
ISCAS3
2024 Retraining-free Constraint-aware Token Pruning for Vision Transformer on Edge Devices
abstract
Vision transformer (ViT) and its variants have demonstrated great potential in various computer vision tasks. However, intensive computation requirements with respect to the token size hinder ViT from being deployed on edge devices with diverse computation resources. Recently, token pruning has been proven to be a promising method to exploit the redundancy of tokens. However, it often requires a laborious retraining process to meet different resource constraints. In this paper, we introduce Fisher information (FI) from tokens to evaluate token importance across different transformer blocks and propose a Retraining-free Constraint-aware Token Pruning (RCTP) framework. RCTP employs a two-step process to obtain the optimal pruning thresholds without retraining under different FLOPs constraints. In the first step, a candidate threshold table and a FLOPs-Fisher table are constructed through a three-stage pipeline to record the trade-off between FLOPs and FI loss of each candidate threshold. In the second step, a modified Viterbi algorithm determines optimal threshold sets with minimum overall FI loss under different FLOPs-constraints in one shot. Our experiment illustrates that RCTP attains better accuracy-FLOPs trade-off than prior pruning-based approaches.
Yun-Chia Yu, Mao-Chi Weng, Ming-Guang Lin, An-Yeu Wu
ISCAS3
2023 TSPTQ-ViT: Two-Scaled Post-Training Quantization for Vision Transformer
abstract
Vision transformers (ViTs) have achieved remarkable performance in various computer vision tasks. However, intensive memory and computation requirements impede ViTs from running on resource-constrained edge devices. Due to the non-normally distributed values after Softmax and GeLU, post- training quantization on ViTs results in severe accuracy degradation. Moreover, conventional methods fail to address the high channel-wise variance in LayerNorm. To reduce the quantization loss and improve classification accuracy, we propose a two-scaled post-training quantization scheme for vision transformer (TSPTQ-ViT). We design the value-aware two-scaled scaling factors (V-2SF) specialized for post- Softmax and post-GeLU values, which leverage the bit sparsity in non-normal distribution to save bit-widths. In addition, the outlier-aware two-scaled scaling factors (O-2SF) are introduced to LayerNorm, alleviating the dominant impacts from outlier values. Our experimental results show that the proposed methods reach near-lossless accuracy drops (<0.5%) on the ImageNet classification task under 8-bit fully quantized ViTs.
Yu-Shan Tai, Ming-Guang Lin, An-Yeu Wu
ICASSP2
2022 Automated Quantization Range Mapping for DAC/ADC Non-linearity in Computing-In-Memory
abstract
Computing-in-memory (CIM) has demonstrated the great potential of analog computing in improving the energy efficiency of matrix-vector multiplications for deep learning applications. Albeit low-power feature of CIM, the non-linearity of digital-to-analog converters (DACs)/analog-to-digital converters (ADCs) causes deviation between the computed outputs and desired values, thus degrading classification accuracy. This paper proposes Automated Quantization Range Mapping (A-QRM) mechanism to mitigate the negative effect of non-linearity on model accuracy. Instead of fixing the quantization range for quantized deep learning models, the proposed A-QRM automatically finds a better quantization range that balances the model capability and quantization errors caused by the non-linearity. Experimental results show that our proposed A-QRM achieves 89.02% and 86.93% of top-1 accuracy in ResNet20 and VGG8 on Cifar-10, respectively, under the non-linearity of DACs/ADCs.
Chi-Tse Huang, Yu-Chuan Chuang, Ming-Guang Lin, An-Yeu Wu
ISCAS3