VLDB 2026 Research / reviewers in the wild / expert
An-Yeu Wu
dblp:94/2234 · also An-Yeu Andy Wu
· DBLP profile ↗
106ranked-venue papers
8as first author
35since 2021 · last 2026
0000-0003-4731-8633ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 68 · 2 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High Energy Efficiency TCAM In-Memory Search Architecture with Pre-Determined Short-Circuit Current Cutoff
Wei-Chieh Lee, Chen-Ming Lee, Chia-Wei Su, Chun-Fu Chen 0006, I-Chieh Hsu, I-Chyn Wey, Tee Hui Teo, An-Yeu Wu |
ISCAS | 8 |
| 2026 | Energy-Efficient In-Memory Vector Similarity Search With Fidelity-Aware Range EncodingabstractVector similarity search (VSS) is a crucial operation in applications of machine learning, but it often incurs high energy consumption due to frequent memory accesses. Previous works have adopted ternary content-addressable memories (TCAMs) to perform parallel VSS within memory. Among these approaches, Exact-Match TCAM (EX-TCAM) combined with range encoding has shown promise for executing in-memory VSS under theL∞norm distance. Existing EX-TCAM-based approaches initiate the search from the positions of query vectors and iteratively expand the search ranges to identify the closest stored vectors. However, this initialization strategy leads to excessive search iterations and long codewords. To address these challenges, we present FORE, an EX-TCAM-based framework that significantly improves latency, energy efficiency, and accuracy for in-memory VSS. To reduce redundant search iterations, we first initialize the search from ranges based on theL∞norm distance. To shorten the codeword length, we then develop a range encoding scheme that supports range-to-range matching. In addition, we introduce a novel metric, “range fidelity,” to evaluate the quality of range encoding. Building on the insight that a certain degree of loss in range fidelity is tolerable in EX-TCAM-based VSS, we further propose a lossy range encoding scheme that yields more compact codewords without significantly compromising accuracy. Finally, FORE incorporates aL∞norm distance-based training mechanism that further reduces search iterations and enhances classification accuracy in EX-TCAM-based VSS. Experimental results demonstrate that FORE improves energy efficiency by 45.25× to 59.72× and reduces latency by 4.47× to 5.66× compared to previous EX-TCAM-based methods within 1% accuracy loss. Moreover, FORE outperforms Best-Match TCAM (Best-TCAM)-based approaches by 6.9× in energy efficiency under non-ideal conditions. Chi-Tse Huang, Hsiang-Yun Cheng, An-Yeu Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | Efficient and Reliable Vector Similarity Search Using Asymmetric Encoding with NAND-Flash for Many-Class Few-Shot LearningabstractWhile memory-augmented neural networks (MANNs) offer an effective solution for few-shot learning (FSL) by integrating deep neural networks with external memory, the capacity requirements and energy overhead of data movement become enormous due to the large number of support vectors in many-class FSL scenarios. Various in-memory search solutions have emerged to improve the energy efficiency of MANNs. NAND-based multi-bit content addressable memory (MCAM) is a promising option due to its high density and large capacity. Despite its potential, MCAM faces limitations such as a restricted number of word lines, limited quantization levels, and non-ideal effects like varying string currents and bottleneck effects, which lead to significant accuracy drops. To address these issues, we propose several innovative methods. First, the Multi-bit Thermometer Code (MTMC) leverages the extensive capacity of MCAM to enhance vector precision using cumulative encoding rules, thereby mitigating the bottleneck effect. Second, the Asymmetric Vector Similarity Search (AVSS) reduces the precision of the query vector while maintaining that of the support vectors, thereby minimizing the search iterations and improving efficiency in many-class scenarios. Finally, the Hardware-Aware Training (HAT) method optimizes controller training by modeling the hardware characteristics of MCAM, thus enhancing the reliability of the system. Our integrated framework reduces search iterations by up to 32×, and increases overall accuracy by 1.58% to 6.94%. Hao-Wei Chiang, Chi-Tse Huang, Hsiang-Yun Cheng, Po-Hao Tseng, Ming-Hsiu Lee, An-Yeu Wu |
ASP-DAC | 6 |
| 2025 | Energy-Efficient Large-Scale Vector Similarity Search in NAND-Flash via Hybrid MatchingabstractVector similarity search (VSS) is crucial in many AI applications, such as few-shot learning (FSL) and approximate nearest-neighbor search (ANNS), but it demands significant memory capacity and incurs substantial energy costs for data transfers during large-scale comparisons. Various in-memory search technologies have been developed to improve energy efficiency, with NAND-based multi-bit content-addressable memory (MCAM) standing out as a promising solution for its high density and large capacity. MCAM can operate in exact-search (ES) mode, supporting only perfect matches with low energy cost, or in approximate-search (AS) mode, enabling flexible VSS. However, AS mode incurs significant energy waste when comparing queries with non-target stored vectors. To address this issue, we propose Hybrid-M, a 3D NAND-based in-memory VSS architecture that integrates both modes into a single hybrid matching process, using ES mode as a filter to reduce redundant searches for AS mode. We apply three techniques to optimize this integration: range encoding for multi-level cells (MLC) to enhance filtering, search voltage shifts to mitigate the impact on AS accuracy and reduce matching currents, and a filtering-aware training method to further improve reliability and energy efficiency. Results show that Hybrid-M achieves comparable accuracy while reducing energy consumption by 67% to 83% compared to MACM-based VSS using only AS mode, across various many-class FSL and ANNS workloads. Chih-Yu Hu, Chi-Tse Huang, Hao-Wei Chiang, Hsiang-Yun Cheng, Po-Hao Tseng, Ming-Hsiu Lee, An-Yeu Wu |
DAC | 7 |
| 2025 | Segmented Angular Pre-Processing for Accurate and Efficient In-Memory Vector Similarity SearchabstractVector similarity search (VSS) is a fundamental operation in modern AI applications, including few-shot learning (FSL) and approximate nearest neighbor search (ANNS). Cosine similarity is widely regarded as the optimal metric for VSS. However, VSS incurs substantial energy and computational overhead, primarily due to frequent vector transfers and the complexity of cosine similarity calculations in high-dimensional spaces. Prior research has explored the use of ternary content addressable memories (TCAMs) for parallel in-memory VSS to reduce vector movement. Exact-Match TCAM (EX-TCAM) enables exact bitmatching, and Best-Match TCAM (Best-TCAM) supports Hamming distance calculations, both of which are spatial metrics and computationally efficient. As a result, existing TCAM-based VSS approaches have focused on developing frameworks to efficiently support more complex spatial metrics such as the $L_{\infty}$ and $L_{1}$ norms. However, these spatial metrics exhibit notable discrepancies compared to angular metrics like cosine similarity. To overcome this limitation, we propose Seg-Cos, a TCAM-based framework that directly approximates cosine similarity within TCAM for angular VSS. Seg-Cos introduces a dedicated preprocessing technique and encoding scheme that segments vectors and encodes them as circular ranges based on their angles and magnitudes. Seg-Cos is the first angular VSS framework compatible with both EX-TCAM and Best-TCAM, enabling accurate and energy-efficient VSS in the angular domain. Simulation results demonstrate that Seg-Cos improves energy efficiency by $1.41 \times$ and achieves up to $2.2 \%$ higher accuracy over prior EX-TCAMbased methods in FSL. In ANNS, Seg-Cos enhances recall rate by $10 \%$ to $52 \%$ and improves energy efficiency by $2 \times$ compared to previous Best-TCAM approaches with $L_{1}$ norm. Chi-Tse Huang, Jen-Chieh Wang, Hsiang-Yun Cheng, An-Yeu Wu |
DAC | 4 |
| 2025 | Deep Unfolding WMMSE Algorithm and Architecture Co-design for Relay-assisted Carrier Aggregation and MU-MIMO SystemsabstractLow latency and high throughput are critical to enhance immersive Extended Reality (XR) experiences. To overcome the limitations of XR devices’ battery and antenna constraints, we investigate a relay-assisted carrier aggregation (RACA) system that transmits data across two frequency bands. However, the iterative weighted minimum mean square error (WMMSE) algorithm used for maximizing data rate incurs long latency. To address this, we propose a deep unfolding WMMSE (UWMMSE) with trainable step sizes in each layer to accelerate convergence via gradient descent. Furthermore, we reschedule serial updates into parallel updates, omit a less-sensitive precoder update for balanced computation, and develop hardware-friendly approximations to simplify matrix inversion and square root operations. To satisfy real-time communication, we implement the first dual-mode UWMMSE hardware engine with folding and pipeline techniques that reduces area by 75% compared to direct mapping and supports MU-MIMO precoder optimization. This engine achieves 1.3x higher hardware efficiency than prior generalized eigenvalue decomposition (GEVD)-based MU-MIMO processors, completing 4 iterations of UWMMSE within 0.4 µs. Chi-Wei Chen, Shu-Kae Liu, An-Yeu Wu |
ISCAS | 3 |
| 2025 | WS-CIM: Enabling Fast and Simultaneous Update for Multi-Macro Compute-in-Memory Architecture Using Weight Sharing TechniqueabstractCompute-in-memory CIM) architecture has been widely studied to accelerate deep neural networks (DNNs). CIM improves energy and area efficiency by performing multiply-accumulate (MAC) operations within memory array. However, the limited size of individual CIM macros requires frequent weight updates for large DNN models, leading to significant latency and energy overhead. In this paper, we propose WS-CIM, a novel framework to enable fast and simultaneous weight updates for multi-macro CIM architectures. Specifically, WS-CIM adopts fine-grained weight sharing technique considering the CIM architecture to minimize redundant write operations, guided by a distance loss function to maintain accuracy. A workload balance mechanism is further introduced for multi-macro architecture to prevent bottlenecks from any single macro during simultaneous weight updates. Experiments on ResNet50 and DeiT-B show that WS-CIM achieves 34.8% latency reduction with 1.29% accuracy loss for ResNet50 and 35.8% latency reduction with 0.70% accuracy degradation for DeiT-B, respectively. These results demonstrate the scalability and efficiency of WS-CIM for deploying DNNs on advanced CIM hardware platforms. Yan-Ding Shieh, Ming-Guang Lin, Hung-Yu Wang, An-Yeu Wu |
ISCAS | 4 |
| 2025 | PAT-ViT: Token Pruning-based Adversarial Tuning for Robust Vision TransformersabstractRecent studies have demonstrated that Vision Transformers (ViTs) are vulnerable to adversarial attacks. While adversarial training is a recognized strategy for enhancing model robustness, it demands substantial computational resources during the training process. Furthermore, the high computational complexity of ViTs during inference presents further challenges. To address these challenges, we propose token pruning-based adversarial tuning for robust ViTs (PAT-ViT). PAT-ViT enhances the robustness of ViTs by reducing their predictability to attackers and tuning them using adversarial data. Unlike conventional methods, PAT-ViT reduces training overhead by tuning pre-trained ViTs rather than training ViTs from scratch. Moreover, PAT-ViT improves inference efficiency by pruning partial input image. Compared to the state-of-the-art (SOTA) robust method, PAT-ViT boosts robust accuracy by 7.5% while also achieving a 10.3% increase in clean accuracy. During inference, PAT-ViT reduces FLOPs by 0.7× relative to SOTA. Additionally, it reduces the training time by 7.5× to 13.1× compared to prior works. Yun-Hao Yang, Yuan-June Luo, Wan-Jung Chen, An-Yeu Wu, Shih-Hsu Huang, Mladen Berekovic |
ISCAS | 4 |
| 2025 | Robust DWMMSE Framework for Energy-Efficient Relay-Assisted Carrier Aggregation (RACA) in Power-Constrained XR SystemsabstractExtended Reality (XR) applications demanding high data rates pose significant challenges for uplink transmissions from user equipment (UE) with limited power and antenna resources. Recently, relay-assisted carrier aggregation (RACA) systems have been proposed to enhance data rates by transmitting data across two frequency bands. This paper addresses the practical issues of power-constrained XR devices and imperfect channel state information (CSI) in RACA systems. We propose a robust Dinkelbach's-transformed weighted minimum mean square error (DWMMSE) framework to maximize energy efficiency (EE) and minimize mean square error (MSE) under Gaussian CSI errors. The framework adapts to content-intensive, powerconstrained, and reliability-critical XR scenarios through specific alternating optimization settings. Simulation results demonstrate that DWMMSE achieves superior EE with a 92% reduction in complexity compared to baseline methods. Moreover, it delivers lower MSE and bit error rate (BER) than non-robust schemes. Chi-Wei Chen, An-Yeu Wu |
VTC2025-Spring | 2 |
| 2025 | Message Passing-Initialized Approximated-Variance Expectation Propagation (MPI-AVEP) Detector for AFDM WaveformabstractHigh-mobility communications in 6G systems face challenges due to Doppler shifts that degrade the performance of traditional orthogonal frequency division multiplexing (OFDM). Affine frequency division multiplexing (AFDM) offers high spectral efficiency but requires efficient detection to exploit its diversity gain. We introduce expectation propagation (EP) as an effective detector for AFDM, improving decoding by estimating the joint posterior distribution of symbols. However, EP incurs high complexity due to matrix inversion, and the Neumannseries approximation (NSA), which is widely used to reduce matrix inversion complexity, fails in AFDM systems due to ill-conditioned channel matrices. To address this, we propose a message passing-initialized approximated-variance EP (MPI-AVEP). It exploits the quasi-banded structure of AFDM channel matrices to reduce the total complexity. The MPI-AVEP achieves performance similar to traditional EP while reducing complexity from cubic to linear in the number of subcarriers. Thus, it offers a practical solution for high-mobility communication systems. Shao-Chun Wang, Chi-Wei Chen, Yi-Ming Lee, An-Yeu Wu |
VTC2025-Spring | 4 |
| 2025 | AMP-ViT: Optimizing Vision Transformer Efficiency with Adaptive Mixed-Precision Post-Training QuantizationabstractVision transformers (ViTs) have revolutionized computer vision but face significant challenges due to their high computational and memory demands. Existing post-training quantization methods struggle to maintain performance at low bit-widths due to activation asymmetry and reliance on manual configurations. To overcome these challenges, we introduce SymAlign to address activation asymmetry and reduce clamping loss. Additionally, we propose AutoScale, an automatic and data-driven mechanism that adapts to variant activations. We incorporate the above-mentioned techniques and propose an adaptive mixed-precision post-training quantization framework for vision transformers (AMP-ViT). Our comprehensive approach addresses asymmetry, variant distribution, and uneven sensitivities, making it the first to tackle these challenges thoroughly. Our experiments on ViT, DeiT, and Swin demonstrate significant accuracy improvements compared with SOTA on the ImageNet dataset. Specifically, our proposed methods achieve accuracy improvements ranging from 0.90% to 23.35% on 4-bit ViTs with single-precision and from 3.82% to 78.14% on 5-bit fully quantized ViTs with mixed-precision. Yu-Shan Tai, An-Yeu Wu |
WACV | 2 |
| 2025 | EECS: Efficient End-Cloud Collaborative System With On-Demand Offloading Mechanism for Convolutional Neural NetworkabstractThe end-cloud synergy with neural networks offers significant promise for enhancing AI-IoT systems. Numerous studies have been conducted to effectively address the dynamic and resource-constrained nature of IoT; however, challenges persist in adaptability to constraints, horizontal integration between technologies, and system-level optimization. In this article, we propose an end-cloud collaborative system (EECS) with an on-demand offloading mechanism for AI-IoT. We enhance three critical areas previously narrowly explored: 1) model pruning to fit into the end device’s constraints; 2) dynamic data transmission between end and cloud; and 3) an effective offloading policy. We pay particular attention to the constraints imposed by multiply-and-add (MAC) operations and bandwidth (BW) at the end, with a focus on image classification tasks. Our contributions are threefold. First, we introduce a two-stage trainable pruning method that can automatically adjust the end model to end constraints and optimize from a system-wide perspective. Second, we propose an adaptive mechanism that accommodates fluctuating BW conditions based on our trainable pruning, integrating static pruning with dynamic inference. Third, we develop a smart offloading policy that enhances decision-making, thereby elevating system-level efficiency. Finally, our EECS shows substantial improvements, achieving$2.6\times $ Yi-Cheng Lo, Cheng-Lin Hsieh, Hao-Wei Chiang, An-Yeu Wu |
IEEE Internet Things J. | 4 |
| 2024 | BFP-CIM: Data-Free Quantization with Dynamic Block-Floating-Point Arithmetic for Energy-Efficient Computing-In-Memory-based AcceleratorabstractConvolutional neural networks (CNNs) are known for their exceptional performance in various applications; however, their energy consumption during inference can be substantial. Analog Computing-In-Memory (CIM) has shown promise in enhancing the energy efficiency of CNNs, but the use of analog-to-digital converters (ADCs) remains a challenge. ADCs convert analog partial sums from CIM crossbar arrays to digital values, with high-precision ADCs accounting for over 60% of the system’s energy. Researchers have explored quantizing CNNs to use low-precision ADCs to tackle this issue, trading off accuracy for efficiency. However, these methods necessitate data-dependent adjustments to minimize accuracy loss. Instead, we observe that the first most significant toggled bit indicates the optimal quantization range for each input value. Accordingly, we propose a range-aware rounding (RAR) for runtime bit-width adjustment, eliminating the need for pre-deployment efforts. RAR can be easily integrated into a CIM accelerator using dynamic block-floating-point arithmetic. Experimental results show that our methods maintain accuracy while achieving up to 1.81 × and 2.08 × energy efficiency improvements on CIFAR-10 and ImageNet datasets, respectively, compared with state-of-the-art techniques. Cheng-Yang Chang, Chi-Tse Huang, Yu-Chuan Chuang, Kuang-Chao Chou, An-Yeu Wu |
ASPDAC | 5 |
| 2024 | BORE: Energy-Efficient Banded Vector Similarity Search with Optimized Range Encoding for Memory-Augmented Neural NetworkabstractMemory-augmented neural networks (MANNs) in-corporate external memories to address the significant issue of catastrophic forgetting in few-shot learning applications. MANNs rely on vector similarity search (VSS), which incurs substantial energy and computational overhead due to frequent data transfers and complex cosine similarity calculations. To tackle these challenges, prior research has proposed adopting ternary content addressable memories (TCAMs) for parallel VSS within memory. One promising approach is to use Exact-Match TCAM (EX-TCAM) with range encoding to find the vector with the minimum$L_{\infty}$distance, avoiding the need for sensing circuit modifications as required by Best-Match TCAM (Best-TCAM), However, this method demands multiple search iterations and longer code words, limiting its practicality. In this paper, we propose an energy-efficient EX-TCAM-based design called BORE. BORE skips redundant search iterations and reduces code word length through performing Banded$L_{\infty}$distance search with Optimized Range Encoding. Additionally, we consider the characteristics of the similarity metric and develop a distance-based training mechanism aimed at improving classification accuracy. Simulation results demonstrate that BORE enhances energy efficiency by$9.35\times$to$11.84\times$and accuracy by 2.95 % to 4.69 % compared to previous EX-TCAM-based approaches. Furthermore, BORE improves energy efficiency by$1.04\times$to$1.63\times$over prior works of Best-TCAM-based VSS. Chi-Tse Huang, Cheng-Yang Chang, Hsiang-Yun Cheng, An-Yeu Wu |
DATE | 4 |
| 2024 | A 40nm 24.6TOPS/W Scalable EfficientDet Processor for Object DetectionabstractObject detection is a crucial technology used to identify and locate objects in a wide range of applications. Google’s EfficientDet, a scalable solution, employs a compound scaling method to systematically adjust the network’s depth, width, and input resolution, meeting different resource constraints on edge devices. This paper presents the first dedicated processor for EfficientDet, featuring three key elements: 1) an adaptive channel/input-wise (CIW) mapper to improve hardware utilization by applying distinct mapping strategies for layers with varying data shapes, 2) a tri-mode activation compression (AC) engine to reduce external memory access (EMA) by leveraging the sparsity level of activations, and 3) a unified aggregation core (AggrCore) to flexibly handle different computations. The chip is fabricated using TSMC 40nm CMOS technology and achieves a maximum energy efficiency of 24.6TOPS/W. Compared to the state-of-the-art object detection processor, our chip demonstrates 3.7× and 2.1× improvements in energy and area efficiencies, respectively. Yu-Chuan Chuang, Ming-Guang Lin, Chi-Tse Huang, Chieh-Fang Teng, Cheng-Yang Chang, Yi-Ta Chen, An-Yeu Wu |
ISCAS | 7 |
| 2024 | Similarity-Aware Fast Low-Rank Decomposition Framework for Vision TransformersabstractVision transformers (ViTs) have shown success in computer vision tasks. However, their high computational and memory demands limit on-device implementation. Low-rank decomposition (LRD) is widely acknowledged as an effective method for reducing computational complexity. Nevertheless, previous approaches relying on heuristic-based or network architecture search (NAS) methods consume extensive searching and training time to uncover the optimal ranks. To efficiently discover rank distributions, this paper introduces a Similarity-Aware Fast-LRD framework, leveraging a greedy selection metric to optimize the compression ratio within each weight matrix by considering cosine similarity and rank reduction. Notably, our proposed framework restores accuracy in 15-25 epochs of model fine-tuning, surpassing the 30-300 epochs needed in prior work. Furthermore, our automated rank search process consumes fewer than 4 GPU hours. Experimental results show that our proposed Similarity-Aware Fast-LRD framework reduces FLOPs and associated parameters by 49.6% and 38.9% for DeiT-B and Swin-S, respectively, with merely 0.83% and 0.72% accuracy degradation on ImageNet. Yuan-June Luo, Yu-Shan Tai, Ming-Guang Lin, An-Yeu Wu |
ISCAS | 4 |
| 2024 | Retraining-free Constraint-aware Token Pruning for Vision Transformer on Edge DevicesabstractVision transformer (ViT) and its variants have demonstrated great potential in various computer vision tasks. However, intensive computation requirements with respect to the token size hinder ViT from being deployed on edge devices with diverse computation resources. Recently, token pruning has been proven to be a promising method to exploit the redundancy of tokens. However, it often requires a laborious retraining process to meet different resource constraints. In this paper, we introduce Fisher information (FI) from tokens to evaluate token importance across different transformer blocks and propose a Retraining-free Constraint-aware Token Pruning (RCTP) framework. RCTP employs a two-step process to obtain the optimal pruning thresholds without retraining under different FLOPs constraints. In the first step, a candidate threshold table and a FLOPs-Fisher table are constructed through a three-stage pipeline to record the trade-off between FLOPs and FI loss of each candidate threshold. In the second step, a modified Viterbi algorithm determines optimal threshold sets with minimum overall FI loss under different FLOPs-constraints in one shot. Our experiment illustrates that RCTP attains better accuracy-FLOPs trade-off than prior pruning-based approaches. Yun-Chia Yu, Mao-Chi Weng, Ming-Guang Lin, An-Yeu Wu |
ISCAS | 4 |
| 2024 | BFP-CIM: Runtime Energy-Accuracy Scalable Computing-in-Memory-Based DNN Accelerator Using Dynamic Block-Floating-Point ArithmeticabstractConvolutional neural networks (CNNs) are known for their exceptional performance in various applications; however, their energy consumption during inference can be substantial. Analog Computing-In-Memory (CIM) has shown promise in enhancing the energy efficiency of CNNs, but the use of analog-to-digital converters (ADCs) remains a challenge. In analog CIM-based accelerators, ADCs convert analog partial sums from CIM crossbar arrays to digital values, with high-precision ADCs accounting for over 60% of the system’s energy consumption. To prevent ADCs from damaging the energy efficiency benefits of CIM, researchers have explored quantizing CNNs to use low-precision ADCs, trading off accuracy for energy efficiency. However, these approaches often necessitate data-dependent adjustments to minimize accuracy loss. Instead, we observe that the first most significant toggled bit indicates the optimal quantization range for each input value. Accordingly, we propose a range-aware rounding (RAR) method for runtime bit-width adjustment, eliminating the need for pre-deployment efforts. RAR can be easily integrated into a CIM accelerator using dynamic block-floating-point arithmetic. We also seamlessly incorporate a bit-level zero-skipping mechanism by dynamically forming input blocks. Experimental results demonstrate that our methods maintain accuracy while achieving up to 1.81$\bm{\times }$and 2.08$\bm{\times }$energy efficiency improvements on the CIFAR-10 and ImageNet datasets, respectively, compared with state-of-the-art techniques. Cheng-Yang Chang, Chi-Tse Huang, Yu-Chuan Chuang, Kuang-Chao Chou, An-Yeu Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | BWA-NIMC: Budget-based Workload Allocation for Hybrid Near/In-Memory-ComputingabstractTo enable efficient computation for convolutional neural networks, in-memory-computing (IMC) is proposed to perform computation within memory. However, the non-ideality significantly degrades the accuracy of IMC. In this work, we leverage a hybrid near/in-memory-computing architecture (NIMC) that allocates sensitive weights to error-free NMC and computes remained weights with high-efficient IMC. We further propose a Budget-based Workload Allocation for NIMC (BWA-NIMC). Specifically, we consider the resource difference between NMC and IMC to effectively allocate workloads under a targeted resource budget. Simulation results show that BWA-NIMC improves the accuracy by 18.38-48.54% under limited budgets (e.g., energy and latency) compared with prior works. Chi-Tse Huang, Cheng-Yang Chang, Yu-Chuan Chuang, An-Yeu Wu |
DAC | 4 |
| 2023 | TSPTQ-ViT: Two-Scaled Post-Training Quantization for Vision TransformerabstractVision transformers (ViTs) have achieved remarkable performance in various computer vision tasks. However, intensive memory and computation requirements impede ViTs from running on resource-constrained edge devices. Due to the non-normally distributed values after Softmax and GeLU, post- training quantization on ViTs results in severe accuracy degradation. Moreover, conventional methods fail to address the high channel-wise variance in LayerNorm. To reduce the quantization loss and improve classification accuracy, we propose a two-scaled post-training quantization scheme for vision transformer (TSPTQ-ViT). We design the value-aware two-scaled scaling factors (V-2SF) specialized for post- Softmax and post-GeLU values, which leverage the bit sparsity in non-normal distribution to save bit-widths. In addition, the outlier-aware two-scaled scaling factors (O-2SF) are introduced to LayerNorm, alleviating the dominant impacts from outlier values. Our experimental results show that the proposed methods reach near-lossless accuracy drops (<0.5%) on the ImageNet classification task under 8-bit fully quantized ViTs. Yu-Shan Tai, Ming-Guang Lin, An-Yeu Wu |
ICASSP | 3 |
| 2023 | Compressive Channel Estimation for IRS-Aided Millimeter-Wave Systems via Two-Stage Lamp NetworkabstractIn millimeter-wave (mmWave) systems aided by intelligent reflecting surfaces (IRSs), accurate channel estimation under low pilot overhead is challenging because of the large number of passive IRS elements. By exploiting the low-rank nature of mmWave channels in the virtual angular domain (VAD) and the powerful learned approximate message passing (LAMP) network, we propose a two-stage LAMP network with row compression (RCTS-LAMP). Specifically, the two LAMP networks jointly recover the VAD channel by solving two low-dimensional sparse signal recovery problems. Moreover, row compression is adopted between the two networks to further reduce the complexity according to the row sparsity structure. Numerical results show that the estimation performance is increased while the computational complexity can be significantly reduced, which achieves a better trade-off between the accuracy and the complexity. Wen-Chiao Tsai, Chi-Wei Chen, An-Yeu Wu |
ICASSP | 3 |
| 2023 | H-RIS: Hybrid Computing-in-Memory Architecture Exploring Repetitive Input SharingabstractComputing-in-memory (CIM) has become a potential trend for accelerating convolutional neural networks (CNNs). Ongoing research, e.g., Repetitive Input Sharing (RIS), focuses on removing redundant matrix-vector multiplication (MVM) by exploiting computational reuse for higher energy efficiency. However, we argue that the RIS neglects the extra overheads of the computation reuse scheme. Moreover, analog CIM is inherently vulnerable to noise. Consequently, reusing the noisy MVM results may lead to severe accuracy degradation. To address the above issues, we first evaluate the extra buffer overheads resulting from the computation reuse scheme for storing repetitive MVM results in the buffer. Based on our evaluation, we find an optimal RIS reuse ratio that balances between buffer costs and the efficiency gain from computation reuse, leading to more energy reduction. In addition, we introduce the RIS-based Hybrid-CIM (H-RIS), which mixes up the analog CIM and digital near-memory-computing (NMC) at the pattern level to maintain accuracy. Based on the above techniques, when we set the RIS ratio to 25%, H-RIS increases 18% accuracy compared with the pure analog CIM and also reduces 97% energy compared with the pure digital NMC. Cheng-Yang Chang, An-Yeu Wu |
ISCAS | 3 |
| 2023 | Machine-aided PPG Signal Quality Assessment (SQA) for Multi-mode Physiological Signal MonitoringabstractPhotoplethysmography (PPG) is a non-invasive technique for recording human vital signs. PPG is normally recorded by wearable devices that are prone to artifacts. This results in signal corruption that decreases measurement accuracy. Thus, a signal quality assessment (SQA) system is essential in obtaining reliable measurements. Conventionally, SQA is mainly driven by human-knowledge and supervised through experts’ annotations. However, they are not tailored for the particularities of the domain applications. Hence, we propose a machine-aided SQA framework that generates respective quality criteria for applications. By using the proposed approach, quality criteria can be easily trained for different applications. Then, quality assessment can be applied to several PPG-based physiological signals telemonitoring. Compared with conventional approaches, the proposed system has a higher rejection rate for high-error signals and a lower mean absolute error is achieved when estimating heart rate (-3.06 BPM), determining respiration rate (–1.36 BPM), and predicting hypertension (+24%). The proposed method enhances accuracy in monitoring physiological signals and thus is suitable for healthcare applications. Win-Ken Beh, Yu-Chia Yang, Yi-Cheng Lo, Yun-Chieh Lee, An-Yeu Wu |
ACM Trans. Comput. Heal. | 5 |
| 2023 | Joint Optimization of Dimension Reduction and Mixed-Precision Quantization for Activation Compression of Neural NetworksabstractRecently, deep convolutional neural networks (CNNs) have achieved eye-catching results in various applications. However, intensive memory access of activations introduces considerable energy consumption, resulting in a great challenge for deploying CNNs on resource-constrained edge devices. Existing research utilizes dimension reduction (DR) and mixed-precision (MP) quantization separately to reduce computational complexity without paying attention to their interaction. Such naïve concatenation of different compression strategies ends up with suboptimal performance. To develop a comprehensive compression framework, we propose an optimization system by jointly considering DR and MP quantization, which is enabled by independent groupwise learnable MP schemes. Group partitioning is guided by a well-designed automatic group partition mechanism that can distinguish compression priorities among channels, and it can deal with the tradeoff between model accuracy and compressibility. Moreover, to preserve model accuracy under low bit-width quantization, we propose a dynamic bit-width searching technique to enable continuous bit-width reduction. Our experimental results show that the proposed system reaches 69.03%/70.73% with average 2.16/2.61 bits per value on Resnet18/MobileNetV2, while introducing only approximately 1% accuracy loss of the uncompressed full-precision models. Compared with individual activation compression schemes, the proposed joint optimization system reduces 55%/9% (−2.62/−0.27 bits) memory access of DR and 55%/63% (−2.60/−4.52 bits) memory access of MP quantization, respectively, on Resnet18/ MobileNetV2 with comparable or even higher accuracy. Yu-Shan Tai, Cheng-Yang Chang, Chieh-Fang Teng, Yi-Ta Chen, An-Yeu Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | S-QRD-ELM: Scalable QR-Decomposition-Based Extreme Learning Machine Engine Supporting Online Class-Incremental Learning for ECG-Based User IdentificationabstractUser identification enables secure access to data and machines in smart factories. Compared with other modalities, ECG-based user identification is rising due to its intrinsic liveness proof and invulnerability to spoofing without contact. On the other hand, as new employees are registered at the factory, the ECG-based user identification system needs to be updated based on the new coming data. This scenario can be defined as an online class-incremental learning (O-CIL) problem. By exploiting hardware-software co- design, this work presents a Scalable QR-decomposition-based extreme learning machine (S-QRD-ELM) engine that can effectively and efficiently support O-CIL for ECG-based user identification. At the software level, we apply the concept of “the others” class and inversion-free QR-decomposition (QRD) recursive least squares to the S-QRD-ELM. This makes S-QRD-ELM achieve 79.7% higher accuracy in the O-CIL scenario compared with the neural network trained with back-propagation (BP-NN). At the hardware level, a one-dimensional diagonally-mapped linear array (1D-DMLA) is proposed to efficiently compute the QRD and back-substitution (BS) operations inside the S-QRD-ELM, reducing 98.5% of the silicon area. Moreover, the integrated processing element (PE) design with the unified COordinate Rotation DIgital Computer (u-CORDIC) further reduces 15.3% of the area and 22.4% of the power consumption. This engine is fabricated in 40nm CMOS technology with a$1.33\times 1.33$mm2 die area. The chip achieves$0.02\mu \text{J}$/sample and$2.47\mu \text{J}$/sample inferencing and learning energy efficiency, respectively, which is$6.4\times $and$28.5\times $than the state-of-the-art. To the best of our knowledge, the proposed highly energy-efficient S-QRD-ELM engine is the first chip to meet the requirements of O-CIL for ECG-based user identification. Yi-Ta Chen, Yu-Chuan Chuang, Li-Sheng Chang, An-Yeu Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Robust PPG-Based Mental Workload Assessment System Using Wearable DevicesabstractHeart rate variability (HRV) has been used in assessing mental workload (MW) level. Compared with ECG, photoplethysmogram (PPG) provides convenient in assessing MW with wearable devices, which is more suitable for daily usage. However, PPG collected by smartwatches are prone to suffer from artifacts. Those signal corruptions cause invalid Inter-beat Intervals (IBI), making it challenging to evaluate the HRV feature. Hence, the PPG-based MW assessment system is difficult to obtain a sustainable and reliable assessment of MW. In this paper, we propose a pre- and post- processing technique, called outlier removal and uncertainty estimation, respectively, to reduce the negative influences of invalid IBIs. The proposed method helps to acquire accurate HRV features and evaluate the reliability of incoming IBIs, rejecting possibly misclassified data. We verified our approach in two open datasets, which are CLAS and MAUS. Experiment results show proposed method achieved higher accuracy (66.7% v.s. 74.2%) and lower variance (11.3% v.s. 10.8%) among users, which has comparable performance to an ECG-based MW system. Win-Ken Beh, Yi-Hsuan Wu, An-Yeu Wu |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Noise-Level Aware Compressed Analysis Framework for Robust Electrocardiography Signal MonitoringabstractCompressed sensing (CS) has drawn much attention in electrocardiography (ECG) signal monitoring for its effectiveness in reducing the transmission power of wireless sensor systems. Compressed analysis (CA) is an improved methodology to further elevate the system's efficiency by directly performing classification on the compressed data at the back-end of the monitoring system. However, conventional CA lacks of considering the effect of noise, which is an essential issue in practical applications. In this work, we observe that noise causes an accuracy drop in the previous CA framework, thus discovering that different signal-to-noise ratios (SNRs) require different sizes of CA models. We propose a two-stage noise-level aware compressed analysis framework. First, we apply the singular value decomposition to estimate the noise level in the compressed domain by projecting the received signal into the null space of the compressed ECG signal. A transfer-learning-aided algorithm is proposed to reduce the long-training-time drawback. Second, we select the optimal CA model dynamically based on the estimated SNR. The CA model will use a predictive dictionary to extract features from the ECG signal, and then imposes a linear classifier for classification. A weight-sharing training mechanism is proposed to enable parameter sharing among the pre-trained models, thus significantly reducing storage overhead. Lastly, we validate our framework on the atrial fibrillation ECG signal detection on the NTUH and MIT-BIH datasets. We show improvement in the accuracy of 6.4% and 7.7% in the low SNR condition over the state-of-the-art CA framework. Yi-Cheng Lo, Win-Ken Beh, Chiao-Chun Huang, An-Yeu Wu |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Compression-Aware Projection with Greedy Dimension Reduction for Convolutional Neural Network ActivationsabstractConvolutional neural networks (CNNs) achieve remarkable performance in a wide range of fields. However, intensive memory access of activations introduces considerable energy consumption, impeding deployment of CNNs on resource-constrained edge devices. Existing works in activation compression propose to transform feature maps for higher compressibility, thus enabling dimension reduction. Nevertheless, in the case of aggressive dimension reduction, these methods lead to severe accuracy drop. To improve the trade-off between classification accuracy and compression ratio, we propose a compression-aware projection system, which employs a learnable projection to compensate for the reconstruction loss. In addition, a greedy selection metric is introduced to optimize the layer-wise compression ratio allocation by considering both accuracy and #bits reduction simultaneously. Our test results show that the proposed methods effectively reduce 2.91×~5.97× memory access with negligible accuracy drop on MobileNetV2/ResNet18/VGG16. Yu-Shan Tai, Chieh-Fang Teng, Cheng-Yang Chang, An-Yeu Wu |
ICASSP | 4 |
| 2022 | Face Recognition for Fisheye ImagesabstractFace recognition often suffers from severe degradation in accuracy when applied to images captured by fisheye cameras. One way to resolve the issue is to rectify the fisheye images before classification; however, it can only achieve local optimum. In this paper, we present an end-to-end model with global optimum. To tackle the challenges due to intra-class variance and diversity of fisheye transformations, we propose 1) a structural correction to guide the model learning, and 2) a spatial-transformer-networks embedded model to compensate for the non-linear distortion of fisheye lenses. We test the proposed model on the CelebA dataset and a real image dataset and achieve an average accuracy of 98.7% and 98.0%, respectively, which represent improvements of 4.05% and 5.72% over the state-of-the-art results. Yi-Cheng Lo, Chiao-Chun Huang, Yueh-Feng Tsai, I-Chan Lo, An-Yeu Wu, Homer H. Chen |
ICIP | 5 |
| 2022 | Automated Quantization Range Mapping for DAC/ADC Non-linearity in Computing-In-MemoryabstractComputing-in-memory (CIM) has demonstrated the great potential of analog computing in improving the energy efficiency of matrix-vector multiplications for deep learning applications. Albeit low-power feature of CIM, the non-linearity of digital-to-analog converters (DACs)/analog-to-digital converters (ADCs) causes deviation between the computed outputs and desired values, thus degrading classification accuracy. This paper proposes Automated Quantization Range Mapping (A-QRM) mechanism to mitigate the negative effect of non-linearity on model accuracy. Instead of fixing the quantization range for quantized deep learning models, the proposed A-QRM automatically finds a better quantization range that balances the model capability and quantization errors caused by the non-linearity. Experimental results show that our proposed A-QRM achieves 89.02% and 86.93% of top-1 accuracy in ResNet20 and VGG8 on Cifar-10, respectively, under the non-linearity of DACs/ADCs. Chi-Tse Huang, Yu-Chuan Chuang, Ming-Guang Lin, An-Yeu Wu |
ISCAS | 4 |
| 2022 | An Effective Entropy-Assisted Mind-Wandering Detection System Using EEG Signals of MM-SART DatabaseabstractMind-wandering (MW), which is usually defined as a lapse of attention has negative effects on our daily life. Therefore, detecting when MW occurs can prevent us from those negative outcomes resulting from MW. In this work, we first collected a multi-modal Sustained Attention to Response Task (MM-SART) database for MW detection. Eighty-two participants' data were collected in our dataset. For each participant, we collected measures of 32-channels electroencephalogram (EEG) signals, photoplethysmography (PPG) signals, galvanic skin response (GSR) signals, eye tracker signals, and several questionnaires for detailed analyses. Then, we propose an effective MW detection system based on the collected EEG signals. To explore the non-linear characteristics of the EEG signals, we utilize entropy-based features. The experimental results show that we can reach 0.712 AUC score by using the random forest (RF) classifier with the leave-one-subject-out cross-validation. Moreover, to lower the overall computational complexity of the MW detection system, we propose correlation importance feature elimination (CIFE) along with AUC-based channel selection. By using two most significant EEG channels, we can reduce the training time of the classifier by 44.16%. By applying CIFE on the feature set, we can further improve the AUC score to 0.725 but with only 14.6% of the selection time compared with the recursive feature elimination (RFE). Finally, we can apply the current work to educational scenarios nowadays, especially in remote learning systems. Yi-Ta Chen, Hsing-Hao Lee, Ching-Yen Shih, Zih-Ling Chen, Win-Ken Beh, Su-Ling Yeh, An-Yeu Wu |
IEEE J. Biomed. Health Informatics | 7 |
| 2021 | Convolutional Neural Network-Aided Bit-Flipping for Belief Propagation Decoding of Polar CodesabstractKnown for their capacity-achieving abilities, polar codes have been selected as the control channel coding scheme for 5G communications. To satisfy high throughput and low latency, belief propagation (BP) is ideal as the decoding algorithm due to its nature of parallel processing. However, the error performance of BP is in general worse than that of enhanced successive cancellation (SC). Recently, bit-flipping (BF) mechanism is applied to BP decoding to lower the error rate. However, its trial-and-error process results in longer latency. In this work, we propose a convolutional neural network-aided bit-flipping (CNN-BF) mechanism to further enhance BP decoding. With carefully designed input data and model architecture, our proposed CNN-BF can achieve better error correction capability with less flipping attempts than prior works. It also achieves a lower block error rate (BLER) than SC list (SCL). Chieh-Fang Teng, Andrew Kuan-Shiuan Ho, Chen-Hsi Derek Wu, Sin-Sheng Wong, An-Yeu Wu |
ICASSP | 5 |
| 2021 | A Scalable Extreme Learning Machine (S-ELM) for Class-Incremental ECG-Based User IdentificationabstractUser identification using electrocardiogram (ECG) is emerging due to the uniqueness and convenience of ECG signals. In addition, in real world applications, new subjects may enter the existed identification system and be authorized to access the private data. Therefore, we propose a scalable extreme learning machine (S-ELM) to meet the demand for class-incremental ECG-based user identification. We first prove that the output weight of the S-ELM learnt in a class-incremental manner is as same as that of a regular ELM which prior has the information of the number of total class. Therefore, in our experiment, S-ELM immunes from the catastrophic forgetting phenomenon, which is a common problem in class-incremental scenarios. Comparing to another class-incremental extreme learning machine such as progressive ELM, S-ELM outperforms progressive ELM by 7% accuracy in online dataset. Comparing to another commonly applied classifier, support vector machine (SVM) with linear and radial basis function (RBF) kernels, S-ELM shows its efficiency by 13.35% and 10.54% higher accuracy but only spends 5.09% and 3.48% of the inference time. Therefore, the proposed S-ELM is promising for the class-incremental ECG-based user identification. Cheng-Lin Lee, Yi-Ta Chen, An-Yeu Wu |
ISCAS | 3 |
| 2021 | MulTa-HDC: A Multi-Task Learning Framework For Hyperdimensional ComputingabstractBrain-inspired Hyperdimensional computing (HDC) has shown its effectiveness in low-power/energy designs for edge computing in the Internet of Things (IoT). Due to limited resources available on edge devices, multi-task learning (MTL), which accommodates multiple cognitive tasks in one model, is considered a more efficient deployment of HDC. However, as the number of tasks increases, MTL-based HDC (MTL-HDC) suffers from the huge overhead of associative memory (AM) and performance degradation. This hinders MTL-HDC from the practical realization on edge devices. This article aims to establish an MTL framework for HDC to achieve a flexible and efficient trade-off between memory overhead and performance degradation. For the shared-AM approach, we propose Dimension Ranking for Effective AM Sharing (DREAMS) to effectively merge multiple AMs while preserving as much information of each task as possible. For the independent-AM approach, we propose Dimension Ranking for Independent MEmory Retrieval (DRIMER) to extract and concatenate informative components of AMs while mitigating interferences among tasks. By leveraging both mechanisms, we propose a hybrid framework of Multi-Tasking HDC, called MulTa-HDC. To adapt an MTL-HDC system to an edge device given a memory resource budget, MulTa-HDC utilizes three parameters to flexibly adjust the proportion of the shared AM and independent AMs. The proposed MulTa-HDC is widely evaluated across three common benchmarks under two standard task protocols. The simulation results of ISOLET, UCIHAR, and MNIST datasets demonstrate that the proposed MulTa-HDC outperforms other state-of-the-art compressed HD models, including SparseHD and CompHD, by up to 8.23% in terms of classification accuracy. Cheng-Yang Chang, Yu-Chuan Chuang, En-Jui Chang, An-Yeu Wu |
IEEE Trans. Computers | 4 |
| 2021 | A 7.8-13.6 pJ/b Ultra-Low Latency and Reconfigurable Neural Network-Assisted Polar Decoder With Multi-Code Length SupportabstractPolar codes have been officially selected as the channel coding in 5G standard. To meet the requirements of enhanced mobile broadband (eMBB), most published polar decoder chips aim to improve throughput rate and error-correction performance. However, to meet with the requirements of another two 5G new radio (NR) application scenarios, ultra-reliable low-latency communications (URLLC), and massive machine-type communications (mMTC), the design features of low latency and energy efficiency are also desirable. In this article, we present a 7.8-13.6 pJ/b ultra-low latency and energy-efficient polar decoder fabricated in 40nm CMOS technology. By adopting the decoding algorithm of recurrent neural network-assisted belief propagation (RNN-BP), the learned scaling parameters can improve the convergence rate by 8 times with reasonable hardware and memory overhead. Then, by taking advantage of BP's regular structure, we propose a fully-reconfigurable RNN-BP decoder architecture to support multiple code lengths with negligible hardware complexity. It contributes to 2- 8× improved hardware utilization rate while providing a flexible adjustment between throughput and error-correction performance. At the architectural level, two optimization techniques for the design of the processing element (PE) are proposed to jointly reduce the chip's area and power by 73% and 67%, respectively. From the measurement results, our reconfigurable RNN-BP polar decoder chip has 2.3×, 2.3×, and 10.0× enhancement over prior designs in terms of latency, throughput rate, and energy efficiency. Consequently, our reconfigurable design has great potential to meet various 5G NR applications. Chieh-Fang Teng, An-Yeu Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2020 | Low-Complexity LSTM-Assisted Bit-Flipping Algorithm For Successive Cancellation List Polar DecoderabstractPolar codes have attracted much attention in the past decade due to their capacity-achieving performance. The higher decoding capacity is required for 5G and beyond 5G (B5G). Although the cyclic redundancy check (CRC)- assisted successive cancellation list bit-flipping (CA-SCLF) decoders have been developed to obtain a better performance, the solution to error bit correction (bitflipping) problem is still imperfect and hard to design. In this work, we leverage expert knowledge in communication systems and adopt deep learning (DL) techniques to obtain a better solution. A low-complexity long short-term memory network (LSTM)-assisted CASCLF decoder is proposed to further improve the performance of conventional CA-SCLF and avoid complexity and memory overhead. Our test results show that we can effectively improve the BLER performance by 0.11dB compared to prior work and reduce the complexity and memory overhead by over 30% of the network. Chun-Hsiang Chen, Chieh-Fang Teng, An-Yeu Wu |
ICASSP | 3 |
| 2020 | Low-Complexity Compressed Alignment-Aided Compressive Analysis for Real-Time Electrocardiography TelemonitoringabstractIn order to implement a real-time electrocardiogram (ECG) telemonitoring, compressed sensing (CS) is a new technology that reduces the power consumption of biosensors and data transmission. Unfortunately, limited label data and computing resources hinder the real-time ECG telemonitoring. Prior experiments have shown that aligning ECG signals is a good way to solve the problem of limited label data. However, the reconstructed learning (RL) framework requires a lot of computing resources, and the compressed learning (CL) framework makes alignment difficult. In this paper, we propose a new compressed alignment-aided compressive analysis (CA-CA) framework that enables simple alignment and low-complexity requirements. From simulation results, we have demonstrated that our technology can maintain more than 95% accuracy while reducing training data (labeled data) by 70%. Therefore, compared to RL, the computation time and memory overhead of CA-CA are reduced by 6.6 times and 2.45 times, respectively. Compared with CL, the inference accuracy with a small amount of labeled data is improved by 13.5%. Yo-Woei Pua, Ching-Yao Chou, An-Yeu Wu |
ICASSP | 3 |
| 2020 | Weighted Pulse Decomposition Analysis of Fingertip Photoplethysmogram Signals for Blood Pressure AssessmentabstractA weighted pulse decomposition analysis is proposed for cuffless blood pressure assessment. Five Gaussian waves are used for decomposition. The weighted least squares criterion is adopted for optimization. Instead of applying weights to certain crest and trough points of the pulse signal from fingertip photoplethysmogram, we apply the weights to the pulse segment that is recognized informative for vascular age, vessel stiffness, and blood pressure. In addition, the boundary constraints of the Gaussian parameters are carefully set so that the normalized root mean square error (NRMSE) between the original PPG and synthesized PPG can be kept acceptable. From the results, we can see that weighting makes the decomposition stable and the variances of Gaussian parameters reduced. Besides, the modified boundary constraints improve the NRMSE. Furthermore, the correlation between the blood pressures and the features from weighted pulse decomposition analysis is enhanced and thus the proposed approach is helpful to blood pressure assessment. Chiu-Hua Huang, Jia-Wei Guo 0001, Yu-Chia Yang, Pei-Yun Tsai 0001, An-Yeu Wu, Hung-Ju Lin, Tzung-Dau Wang |
ISCAS | 5 |
| 2019 | Scattering Multi-connectivity Estimation for Indoor mmWave Small Cells under Limited Training StepsabstractMulti-connectivity, connections among multiple small cells (SCs) simultaneous, is used to be a promising solution to achieve high reliability in indoor millimeter-wave (mmWave) networks. This work further improves the multi-connectivity estimation under scenarios with strictly-limited training steps. To quickly estimate multiple links connected to different SCs, we propose a novel scattering multi-beam codebook (SMBC). Then, we develop a scattering multi-connectivity estimation (SMCE) measuring effective possible links instead of all possible links. From our simulation results, our method can complete multi-connectivity in only 20% of the training steps in the uplink exhaustive measurement (UEM). Hung-Yi Cheng, Ching-Chun Liao, An-Yeu Wu |
ICASSP | 3 |
| 2019 | Low-Complexity Compressive Analysis in Sub-Eigenspace for ECG Telemonitoring SystemabstractCompressive sensing (CS) is attractive in long-term electrocardiography (ECG) telemonitoring to extend life-time for resource-limited wireless wearable sensors. Moreover, health monitoring has emphasized the need for edge computing to process real-time data without the bandwidth costs. However, the reconstructed analysis (RA) and the compressed learning (CL) frameworks have extremely high memory and computational overhead, cost-prohibitive for online usage at resource-constrained edge device. In this paper, to efficiently analyze the received CS measurements with different levels of compression, we propose a low-complexity framework of Compressive Analysis in Sub-Eigenspace (CA-SE) based on subspace-based representation. The dictionary is used for sifting the sub-eigen information from the CS measurements online, and it is built by eigenspace learning offline. The framework can reduce the memory overhead with a single light-weight machine learning model and multiple small filter matrices, and the computational complexity with sifting by matrix-vector product rather than sparse coding. CA-SE is implemented in ECG-based atrial fibrillation detection. The memory overhead of CA-SE is 13 and 39 times fewer compared with RA and CL, respectively, and the computational complexity of CA-SE is 42 and 10 times fewer compared with RA and CL, respectively. Ching-Yao Chou, An-Yeu Wu |
ICASSP | 2 |
| 2019 | Low-complexity Recurrent Neural Network-based Polar Decoder with Weight Quantization MechanismabstractPolar codes have drawn much attention and been adopted in 5G New Radio (NR) due to their capacity-achieving performance. Recently, as the emerging deep learning (DL) technique has breakthrough achievements in many fields, neural network decoder was proposed to obtain faster convergence and better performance than belief propagation (BP) decoding. However, neural networks are memory-intensive and hinder the deployment of DL in communication systems. In this work, a low-complexity recurrent neural network (RNN) polar decoder with codebook-based weight quantization is proposed. Our test results show that we can effectively reduce the memory overhead by 98% and alleviate computational complexity with slight performance loss. Chieh-Fang Teng, Chen-Hsi Derek Wu, Andrew Kuan-Shiuan Ho, An-Yeu Wu |
ICASSP | 4 |
| 2019 | Low-Complexity Compressed-Sensing-Based Watermark Cryptosystem and Circuits Implementation for Wireless Sensor NetworksabstractThe emerging compressed sensing (CS) technique provides lightweight data compression with zero-cost encryption. Therefore, CS enables reduced-complexity designs for sensor nodes and saves transmission power in wireless sensor networks (WSNs). However, CS's linear encoding process makes it vulnerable to several attacks, which may lead to privacy leakage issues. In this article, leveraging the characteristic that CS reconstruction is sensitive to measurement noise, we propose a CS-based watermark cryptosystem for WSNs. In the front-end sensor, a low-dimension watermark is embedded in measurements. In the back-end solver, we present a CS-based watermark decryption/reconstruction engine for the Internet of Things (IoT) gateway. Without synchronization of the key, the proposed cryptosystem can resist ciphertext-only attack and known-plaintext attack effectively. Furthermore, leveraging watermarks as a digital signature, the proposed engine can detect denial of service attack effectively before signal reconstruction. For real-time signal processing, multiple-indices updating algorithm and VLSI architecture are applied to eliminate the throughput degradation from watermark removal. Finally, this CS decoder is fabricated in 40-nm CMOS, and it can support the simultaneous reconstruction of over 10 000 wireless sensors in real time while offering synchronization-free watermark decryption. Therefore, the proposed cryptosystem is suitable for the emerging IoT applications that need encryption strength with limited complexity. Ting-Sheng Chen, Kai-Ni Hou, Win-Ken Beh, An-Yeu Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | Bundle-Updatable SRAM-Based TCAM Design for OpenFlow-Compliant Packet ProcessorabstractStatic random-access memory (SRAM)-based ternary content-addressable memories (TCAMs) emulate TCAM functions with high throughput at low cost. However, the implementation of SRAMbased TCAM as a rule table in network switches tends to prolong updating latency, which can cause a false packet routing. This brief proposes a novel low-latency bundle-updatable TCAM (BU-TCAM) scheme that uses binary tree-based prefix encoding (BPE) to support singleand multiple-rule updating in a software-defined networking/OpenFlow network. The proposed encoding method transforms the original ternary rule data into a binary code word and determines the range of overlap in SRAM addresses to facilitate updating. This greatly decreases latency in cases where multiple rules are required to update on an SRAMbased TCAM. We implemented an emulated 64 × 32-bit TCAM of the proposed design on a Xilinx ZC-706 field-programmable gate array. The proposed scheme reduced updating latency by 79.6%, compared with a conventional updating structure, which had only 9.8% and 23% increases in LUTs and registers overhead, respectively. Ding-Yuan Lee, Ching-Che Wang, An-Yeu Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Low-Complexity Secure Watermark Encryption for Compressed Sensing-Based Privacy PreservingabstractThe emerging compressed sensing (CS) technique enables new reduced-complexity designs of sensor nodes and helps to save overall transmission power in wireless sensor network. Because of the linearity of its encoding process, CS is vulnerable to Ciphertext-Only Attack (COA) and Known-Plaintext Attack (KPA). The prior works use multiple sensing matrices as the shared secret key, however, the complexity overhead of frontend sensor and synchronization issue arising from multiple keys should be well considered. In this paper, by leveraging the characteristic of CS that is sensitive to destroyed sparsity, a low-dimension watermark is randomly chosen and embedded in measurement in front-end part. Then, in back-end solver, the proposed decrypting basis can decipher the encrypted signals without synchronization. Simulation results show that the proposed scheme achieves effective protection against COA and KPA with only 5% storage overhead. It furtherly eases the encryption complexity of front-end sensor by 98.8% under our experiments. Kai-Ni Hou, Ting-Sheng Chen, Hung-Chi Kuo, Tzu-Hsuan Chen, An-Yeu Wu |
ICASSP | 5 |
| 2018 | Error-Resilient Reconfigurable Boosting Extreme Learning Machine for ECG Telemonitoring SystemsabstractMachine learning models have gained popularity of realizing Electrocardiography (ECG) monitoring systems. For constructing an inference device, providing both low latency and high accuracy is of a great concern. Extreme Learning Machine (ELM) is a single layer neural network that provides an effective solution for fast inference. In addition, the use of Adaptive Boosting (AdaBoost) algorithm can aggregate ELMs to enhance the overall learning ability. However, these computing units may encounter reliability issues that result from CMOS technology scaling and lead to a severe decline in performance. Hence, improving error resilience for a machine learning engine becomes a new design issue. This work presents a reliability-aware scheme for AdaBoost-based ELM. By exploiting the inherent redundancy in AdaBoost algorithm, it can strengthen the combination between different ELM classifiers. In an ECG-based atrial fibrillation detection case, the experimental results show that the proposed method can restore 71.4% of accuracy degradation caused by injected random bit-flip rate of 4 × 10-4in computing units with small computational overhead. The classification engine is synthesized by TSMC 40nm CMOS technology, which can achieve extremely high classification rate. Sheng-Hui Wang, Huai-Ting Li, An-Yeu Wu |
ISCAS | 3 |
| 2018 | Dynamically Updatable Ternary Segmented Aging Bloom Filter for OpenFlow-Compliant Low-Power Packet Processing
Sheng-Chun Kao, Ding-Yuan Lee, Ting-Sheng Chen, An-Yeu Wu |
IEEE/ACM Trans. Netw. | 4 |
| 2017 | Compressive sensing based ECG monitoring with effective AF detectionabstractAtrial fibrillation (AF) patients need long-term electrocardiography (ECG) monitoring to detect occurrence of AF. We can acquire ECG signals under low power by compressive sensing based sensor and detect AF by existing algorithms. However, the compression ratio of AF signal is low when DWT basis is applied for CS reconstruction. On the other hand the complexity of AF detection algorithms is high. In this paper, we propose a CS-based ECG monitoring system with effective AF detection. We exploit dictionary learning to improve 2.5× better compression ratio than existing works. With built-in AF detection, we can detect AF with 96.0% sensitivity and 97.2% specificity from highly compressed data, without any complex detection algorithm. Hung-Chi Kuo, Yu-Min Lin, An-Yeu Wu |
ICASSP | 3 |
| 2017 | Path-Diversity-Aware Fault-Tolerant Routing Algorithm for Network-on-Chip SystemsabstractNetwork-on-Chip (NoC) is the regular and scalable design architecture for chip multiprocessor (CMP) systems. With the increasing number of cores and the scaling of network in deep submicron (DSM) technology, the NoC systems become subject to manufacturing defects and have low production yield. Due to the fault issues, the reduction in the number of available routing paths for packet delivery may cause severe traffic congestion and even to a system crash. Therefore, the fault-tolerant routing algorithm is desired to maintain the correctness of system functionality. To overcome fault problems, conventional fault-tolerant routing algorithms employ fault information and buffer occupancy information of the local regions. However, the information only provides a limited view of traffic in the network, which still results in heavy traffic congestion. To achieve fault-resilient packet delivery and traffic balancing, this work proposes a Path-Diversity-Aware Fault-Tolerant Routing (PDA-FTR) algorithm, which simultaneously considers path diversity information and buffer information. Compared with other fault-tolerant routing algorithms, the proposed work can improve average saturation throughput by 175 percent with only 8.9 percent average area overhead and 7.1 percent average power overhead. Yu-Yin Chen, En-Jui Chang, Hsien-Kai Hsin, Kun-Chih Chen, An-Yeu Wu |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | Variation-Aware Reliable Many-Core System Design by Exploiting Inherent Core RedundancyabstractReliability issues are more severe in multi/many-core systems because of the integration of more devices in advanced technology nodes. To achieve robust computing in nanoscale designs, many circuit-level and architecture-level redundancy techniques had been proposed, which pose large fixed silicon area overhead and a lack of flexibility. In recent years, some methods have exploited the “inherent core redundancy” of many-core systems to implicitly implement N-modular redundant (NMR) subsystems to achieve area-efficient fault-tolerant computing. However, while facing the different levels of soft error rate, task vulnerability, and task significance in the many-core system, existing core-level redundancy methods become ineffective. To achieve robust computation in many-core systems with intercore variations and mixed workloads, we propose a variation-aware core-level redundancy scheme. Two novel approaches are presented in this scheme: 1) we construct NMR tables that store the degree of redundancy using mathematical models for systems affected by these variations and 2) we dynamically allocate each replicated task to a proper core with variation-aware mapping algorithms to achieve high reliability. Based on a modified multicore simulator, Sniper-Transient Error Process Variation (TEVR), the experimental results show that the proposed scheme can increase the reliability by 47.92% and achieve the energy saving of 39% compared with conventional core-level redundancy methods. Huai-Ting Li, Ching-Yao Chou, Yuan-Ting Hsieh, Wei-Ching Chu, An-Yeu Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | A 1.96mm2 low-latency multi-mode crypto-coprocessor for PKC-based IoT security protocolsabstractIn this paper, we present the implementation of a multi-mode crypto-coprocessor, which can support three different public-key cryptography (PKC) engines (NTRU, TTS, Pairing) used in post-quantum and identity-based cryptosystems. The PKC-based security protocols are more energy-efficient because they usually require less communication overhead than symmetric-key-based counterparts. In this work, we propose the first-of-its-kind tri-mode PKC coprocessor for secured data transmission in Internet-of-Things (IoT) systems. For the purpose of low energy consumption, the crypto-coprocessor incorporates three design features, including 1) specialized instruction set for the multi-mode cryptosystems, 2) a highly parallel arithmetic unit for cryptographic kernel operations, and 3) a smart scheduling unit with intelligent control mechanism. By utilizing the parallel arithmetic unit, the proposed crypto-coprocessor can achieve about 50% speed up. Meanwhile, the smart scheduling unit can save up to 18% of the total latency. The crypto-coprocessor was implemented with AHB interface in TSMC 90nm CMOS technology, and the die size is only 1.96 mm2. Furthermore, our chip is integrated with an ARM-based system-on-chip (SoC) platform for functional verification. Cheng-Rung Tsai, Ming-Chun Hsiao, Wen-Chung Shen, An-Yeu Wu, Chen-Mou Cheng |
ISCAS | 4 |
| 2015 | Regional ACO-Based Cascaded Adaptive Routing for Traffic Balancing in Mesh-Based Network-on-Chip SystemsabstractThe regular topology of mesh-based network-on-chip (NoC) provides flexible and scalable architecture for chip multiprocessor (CMP) systems. However, as the complexity of network increases, routing problems become performance bottlenecks. In the field of wide area networks (WANs), ant colony optimization (ACO) has been applied to an adaptive routing for improving performance and achieving load balancing. Nevertheless, if we directly apply ACO to NoC systems, the implementation cost of ACO is excessively high. To overcome this problem, the ACO-based adaptive routing must be reformulated while considering both router cost and NoC efficiency. This work proposes the regional ACO-based cascaded adaptive routing (RACO-CAR) scheme with the following techniques: 1) table elimination by removing redundant information, 2) table sharing by grouping pheromone information to merge table content, and 3) cascaded routing that assigns traffic to different uncongested regions to balance traffic. Our experimental results demonstrate that the RACO-CAR scheme has an improvement of 3.9-36.84 percent in saturation throughput compared with existing adaptive routing schemes. The implementation cost of the RACO-CAR router is only 37.4 percent of that of the ACO-based router with full routing table. Therefore, the proposed RACO-CAR scheme has high area efficiency, defined as saturation throughput divided by the total cost of router. En-Jui Chang, Hsien-Kai Hsin, Chih-Hao Chao, Shu-Yen Lin, An-Yeu Wu |
IEEE Trans. Computers | 5 |
| 2015 | Ant Colony Optimization-Based Adaptive Network-on-Chip Routing Framework Using Network Information RegionabstractThe network-on-chip (NoC) system can provide more scalable and flexible on-chip interconnection compared with system bus. The performance of on-chip adaptive routing algorithms greatly relies on the adopted network information. To the best our knowledge, previous routing algorithms utilize either spatial or temporal network information to improve performance. However, few works have established a framework on analyzing the network information nor showed how to integrate the spatial and temporal network information. In this paper, we define the network information region (NIR) framework for NoC systems. The NIR can indicate arbitrary combinations of network information and corresponding routing algorithms. We demonstrate how to apply NIR on analyzing the adaptive routing algorithms. To further demonstrate how NIR can help to integrate the spatial or temporal network information, we propose the ACO-based pheromone diffusion (ACO-PhD) adaptive routing framework based on the NIR. By diffusing the pheromone outward, spatial and temporal network information can be exchanged among adjacent routers. The range (i.e., size and shape) of the NIR is controllable by setting the parameters in the ACO-PhD algorithm. We show that we can reconfigure the ACO-PhD algorithm to each routing algorithm in its NIR subsets by adjusting the parameter settings. Finally, we implement and analyze the hardware design of corresponding router architecture. The results show an improvement of 4.86-16.93 percent on network performance and the highest area efficiency is achieved by the proposed algorithm. Hsien-Kai Hsin, En-Jui Chang, Kuan-Yu Su, An-Yeu Wu |
IEEE Trans. Computers | 4 |
| 2015 | RC-Based Temperature Prediction Scheme for Proactive Dynamic Thermal Management in Throttle-Based 3D NoCsabstractThe three-dimensional Network-on-Chip (3D NoC) has been proposed to solve the complex on-chip communication issues in multicore systems using die stacking in recent days. Because of the larger power density and the heterogeneous thermal conductance in different silicon layers of 3D NoC, the thermal problems of 3D NoC become more exacerbated than that of 2D NoC and become a major design constraint for a high-performance system. To control the system temperature under a certain thermal limit, many Dynamic Thermal Managements (DTMs) have been proposed. Recently, for emergent cooling, the full throttling scheme is usually employed as the system temperature reaches the alarming level. Hence, the conventional reactiveDTMsuffers from significant performance impact because of the pessimistic reaction. In this paper, we propose a throttle-based proactiveDTM(T-PDTM) scheme to predict the future temperature through a newThermal RC-based temperature prediction (RCTP) model. TheRCTPmodel can precisely predict the temperature with heterogeneous workload assignment with low constant computational complexity. Based on the predictive temperature, the proposedT-PDTMscheme will assign the suitable clock frequency for each node of the NoC system to perform early temperature control through power budget distribution. Based on the experimental results, compared with the conventional reactive throttled-basedDTMs, theT-PDTMscheme can help to reduce 11.4∼80.3 percent fully throttled nodes and improves the network throughput by around 1.5∼211.8 percent. Kun-Chih Chen, En-Jui Chang, Huai-Ting Li, An-Yeu Wu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2014 | Robust decision feedback equalizer scheme by using sphere-decoding detectorabstractThe decision feedback equalizer (DFE) is an efficient scheme to suppress intersymbol interference (ISI) in various communication and magnetic recording systems. However, most DFE implementations suffer from the phenomenon of error propagation, which degrades its bit error rate (BER) performance. In this paper, We use sphere detector (SD) to achieve maximum likelihood (ML) detection and significantly reduce the system symbol error rate (SER). Simulations show that the proposed scheme with sphere detector decision feedback equalizer (SD-DFE) algorithm can efficiently reduce the SER. At SNR=28, the SER can be improved from 2.0 × 10−5(Ideal DFE) to 1.8 × 10−6(six-stage SD-DFE). Hung-Yi Cheng, Chun-Yuan Chu, Yen-Liang Chen, An-Yeu Wu |
ICASSP | 4 |
| 2014 | High-throughput QC-LDPC decoder with cost-effective early termination scheme for non-volatile memory systemsabstractThis paper presents a high-throughput layered min-sum quasi-cyclic LDPC (QC-LDPC) decoder for non-volatile memory systems (NVMs). A cost-effective column-based early termination (CB-ET) scheme is proposed to early terminate decoding process within iteration. The throughput improvement is 37.7% compared to the state-of-the-art early termination scheme when raw bit error rate of flash memory is 3×10−3. The QC-LDPC decoder with proposed early termination scheme is synthesized by TSMC 90nm CMOS technology, and the area overhead is only 2.20%. Yu-Min Lin, Yu-Hao Chen, Ming-Han Chung, An-Yeu Wu |
ISCAS | 4 |
| 2014 | Path-Congestion-Aware Adaptive Routing With a Contention Prediction Scheme for Network-on-Chip SystemsabstractNetwork-on-chip systems can achieve higher performance than bus systems for chip multiprocessor systems. However, as the complexity of the network increases, the channel and switch congestion problems become major performance bottlenecks. An effective adaptive routing algorithm can help minimize path congestion through load balancing. However, conventional adaptive routing schemes only use channel-based information to detect the congestion status. Due to the lack of switch-based information, channel-based information is difficult to reveal the real congestion status along the routing path. Therefore, in this paper, we remodel the path congestion information to show hidden spatial congestion information and improve the effectiveness of routing path selection. We propose a path-congestion-aware adaptive routing (PCAR) scheme based on the following techniques: 1) a path-congestion-aware selection strategy that simultaneously considers switch congestion and channel congestion, and 2) a contention prediction technique that uses the rate of change in the buffer level to predict possible switch contention. The experimental results show that the proposed PCAR scheme can achieve a high saturation throughput with an improvement of 15.4%-48.7% compared to existing routing schemes. The proposed PCAR method also includes a VLSI architecture, which has higher area efficiency with an improvement of 16%-35.7% compared with the other router designs. En-Jui Chang, Hsien-Kai Hsin, Shu-Yen Lin, An-Yeu Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2014 | Ant Colony Optimization-Based Fault-Aware Routing in Mesh-Based Network-on-Chip SystemsabstractThe advanced deep submicrometer technology increases the risk of failure for on-chip components. In advanced network-on-chip (NoC) systems, the failure constrains the on-chip bandwidth and network throughput. Fault-tolerant routing algorithms aim to alleviate the impact on performance. However, few works have integrated the congestion-, deadlock-, and fault-awareness information in channel evaluation function to avoid the hotspot around the faulty router. To solve this problem, we propose the ant colony optimization-based fault-aware routing (ACO-FAR) algorithm for load balancing in faulty networks. The behavior of an ant colony while facing an obstacle (failure in NoC) can be described in three steps: 1) encounter; 2) search; and 3) select. We implement the corresponding mechanisms as: 1) notification of fault information; 2) path searching mechanism; and 3) path selecting mechanism. With proposed ACO-FAR, the router can evaluate the available paths and detour packets through a less-congested fault-free path. The simulation results show that this paper has higher throughput than related works by 29.1%–66.5%. In addition, ACO-FAR can reduce the undelivered packet ratio to 0.5%–0.02% and balance the distribution of traffic flow in the faulty network. Hsien-Kai Hsin, En-Jui Chang, Chia-An Lin, An-Yeu Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2014 | Spatial-Temporal Enhancement of ACO-Based Selection Schemes for Adaptive Routing in Network-on-Chip SystemsabstractNetworks-on-Chip (NoC) provides a regular and scalable design architecture for chip multi-processor (CMP) systems. The Ant Colony Optimization (ACO) is a distributed algorithm. Applying ACO to selection models of adaptive routing can improve NoC performance. Currently, ACO-based selection only uses the historical traffic information. While additional temporal and spatial information provides better approximation of network status for global load-balancing. In this paper, we first consider the temporal enhancement of congestion information. We propose the Multi-Pheromone ACO-based (MP-ACO) selection scheme which adopts the concept of Exponential Moving Average (EMA) from stock market. We implement a novel ACO system where ants lay two kinds of pheromones with different evaporation rates. The temporal pheromone variation can help to capture hidden-state dependencies of upcoming congestion status. Secondly, to acquire the spatial range of congestion information, we propose Regional-Aware ACO-based (RA-ACO) selection to record historical buffer information from routers within two-hop of distances, which helps to extend spatial pheromone coverage. Information provided by the proposed two schemes improves the system performance. Simulation results show that MP-ACO and RA-ACO with Odd-Even routing algorithm yields an improvement in saturation throughput over OBL and NoP selection by 14.38 percent and 18.64 percent, respectively. The router architectures for the proposed schemes are also implemented and analyze with small hardware overhead. Hsien-Kai Hsin, En-Jui Chang, An-Yeu Wu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | Cost-effective scalable QC-LDPC decoder designs for non-volatile memory systemsabstractThis paper presents a cost-effective scalable quasi-cyclic LDPC (QC-LDPC) decoder architecture for non-volatile memory systems (NVMS). A re-arranged architecture is proposed to eliminate the first-in-first-out (FIFO) memory in conventional decoders, where the FIFO size is linearly proportional to the codeword size. The area reduction is 18.5% compared to the conventional decoder architecture. The scalable datapaths of the proposed decoder reduce the re-design cost and enable the flexibility of using QC-LDPC codes for NVMS. A prototyping decoder with maximum codeword size of 9280 bits is implemented in TSMC 90nm CMOS technology, and the core area is only 2.52mm2at 138.8MHz. Ming-Han Chung, Yu-Min Lin, Cheng-Zhou Zhan, An-Yeu Wu |
ICASSP | 4 |
| 2013 | Motion artifact elimination algorithm with eigen-based clutter filter for color Doppler processingabstractColor Doppler imaging is used to visualize the distribution of blood flow in the region of interest. Slight relative motion may cause severe image corruption and incorrect blood velocity estimation. In this work, we propose a velocity bias cancellation algorithm based on the autocorrelation technique widely used in color Doppler and eigen-based clutter filter to eliminate the motion artifact. The proposed algorithm assists clutter filter to suppress tissue noises effectively and compensates the biased blood velocity. It has more than 3-9 dB better performance and the error of blood velocity estimation can be reduced by more than 69%. Zih-Ling Liu, Yu-Hao Chen, Cheng-Zhou Zhan, An-Yeu Wu |
ICASSP | 4 |
| 2013 | Traffic- and Thermal-aware Adaptive Beltway Routing for three dimensional Network-on-Chip systemsabstractThe distribution of traffic and temperature in a high-performance three dimensional Network-on-Chip (3D NoC) system become more unbalanced because of chip stacking and applied minimal routing algorithms. To regulate the temperature under a certain thermal limit, the overheated nodes are usually throttled by run-time thermal management (RTM). Therefore, the network topology becomes a Non-Stationary Irregular Mesh (NSI-Mesh) and leads to heavy traffic congestion around the throttled nodes. Because of the traffic imbalance in the network, the system performance degrades sharply as temperature rises. In this paper, a Traffic- and Thermal-aware Adaptive Beltway Routing (TTABR) is proposed to balance both the distribution of the traffic and temperature in the network. The proposed TTABR can be applied to NSI-Mesh and regular mesh. The experimental results show that the proposed TTABR can achieve more balanced both traffic and temperature distribution, and the network throughput is improved by around 3.4~113% with less than 18% area overhead. Kun-Chih Chen, Che-Chuan Kuo, Hui-Shun Hung, An-Yeu Wu |
ISCAS | 4 |
| 2013 | VLSI implementation of real-time motion compensated beamforming in synthetic transmit aperture imagingabstractSynthetic transmit aperture (STA) has been widely investigated in ultrasound system recently due to its high frame rate and low cost characteristics. Since the highresolution image (HRI) of STA is formed by summation of low-resolution images (LRIs), it is susceptible to motion between firings. In this work, we propose a low-complexity two-dimensional motion compensation algorithm. The velocity and direction of motion can be evaluated by cross-correlation between specific beams according to geometry characteristics of STA. Compared to the uncompensated image, simulation results which used Field II program show that proposed method can improve the contrast ratio (CR) and contrast noise ratio (CNR) about 8.6 dB and 1.3 dB. The whole imaging system was implemented in TSMC 90nm technology. Operating at 125 MHz, the circuit with 11% hardware overhead for motion compensation can beamform 64 image lines consisting of 1024 complex samples at the rate of 45 frames per second. Kuan-Yu Ho, Yu-Hao Chen, Cheng-Zhou Zhan, An-Yeu Wu |
ISCAS | 4 |
| 2013 | Transport-layer-assisted routing for runtime thermal management of 3D NoC systems
Chih-Hao Chao, Kun-Chih Chen, Tsu-Chu Yin, Shu-Yen Lin, An-Yeu Wu |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2013 | Topology-Aware Adaptive Routing for Nonstationary Irregular Mesh in Throttled 3D NoC SystemsabstractThree-dimensional network-on-chip (3D NoC) has been proposed to solve the complex on-chip communication issues in future 3D multicore systems. However, the thermal problems of 3D NoC are more serious than 2D NoC due to chip stacking. To keep the temperature below a certain thermal limit, the thermal emergent routers are usually throttled. Then, the topology of 3D NoC becomes a Nonstationary Irregular Mesh (NSI-Mesh). To ensure the successful packet delivery in the NSI-Mesh, some routing algorithms had been proposed in the previous works. However, the network still suffers from extremely traffic imbalance among lateral and vertical logic layer. In this paper, we propose a Topology Aware Adaptive Routing (TAAR) to balance the traffic load for NSI-Mesh in 3D NoC. TAAR has three routing modes, which can be dynamically adjusted based on the topology status of the routing path. In addition to increasing routing flexibility, the TAAR also increases both vertical and lateral path diversity to balance the traffic load. Compared with the related adaptive routing methods, the experimental results show that the proposed TAAR can reduce 19 to 295 percent traffic loads in the bottom logic layer and improve around 7.7 to 380 percent network throughput. According to our proposed VLSI architecture, the TAAR only needs less than 24.8 percent hardware overhead compared with the previous works. Kun-Chih Chen, Shu-Yen Lin, Hui-Shun Hung, An-Yeu Wu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | Routing-Based Traffic Migration and Buffer Allocation Schemes for 3-D Network-on-Chip Systems With Thermal LimitabstractThe 3-D network-on-chip (NoC) router is a major source of thermal hotspots, limiting the performance gain of 3-D integration. Due to the varying cooling efficiency of different silicon layers in 3-D NoC, the optimal criteria of traditional load balancing design (LBD) scheme and temperature balancing design (TBD) scheme may not be satisfied. To analyze the tradeoff between performance and temperature, we provide a new analytical model. The model shows that the LBD scheme and the TBD scheme can be considered as two corner cases in the design space, and design cases can be categorized by comparing the bandwidth bound and the thermal-limited bound. To find the optimal design criteria between the LBD and the TBD schemes in 3-D NoC, we propose a new routing-based traffic migration, vertical-downward lateral-adaptive proactive routing (VDLAPR), and buffer allocation methods, vertical buffer allocation (VBA). The VDLAPR algorithm enables to tradeoff between the LBD and the TBD schemes. The proposed VBA method mitigates the traffic congestion caused by traffic migration. To reach the optimal configuration, we propose a systematic design flow, which assists in finding the best design parameters in the expanded space between LBD and TBD. Based on the traffic-thermal co-simulation experiments, the achievable throughput can be improved from 2.7% to 45.2% using the proposed design scheme. Chih-Hao Chao, Kun-Chih Chen, An-Yeu Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | Reconfigurable Adaptive Singular Value Decomposition Engine Design for High-Throughput MIMO-OFDM SystemsabstractSingular value decomposition (SVD) is an optimal method to obtain spatial multiplexing gain in multi-input multi-output (MIMO) channels. However, the high cost of implementation and high decomposing latency of the SVD restricts its usage in current wireless communication applications. In this paper, we present a complete adaptive SVD algorithm and a reconfigurable architecture for high-throughput MIMO-orthogonal frequency division multiplexing systems. There are several proposed architectural design techniques: reconfigurable scheme, division-free adaptive step size scheme, early termination scheme, and data interleaving scheme. The reconfigurable scheme can support all antenna configurations in a MIMO system. The division-free adaptive step size and early termination schemes are used to effectively reduce the decomposing latency and improve hardware utilization. The data interleaving scheme helps to deal with several channel matrices concurrently. Besides, we propose an orthogonal reconstruction scheme to obtain more accurate SVD outputs, and then the system performance will be greatly enhanced. We apply our SVD design to the IEEE 802.11 n applications. This design is implemented and fabricated in UMC 90 nm 1P9M CMOS technology. The maximum operating frequency is measured to be at 101.2 MHz, and the corresponding power dissipation is at 125 mW. The core size is 2.17 mm2and the die size occupies 4.93 mm2. The chip result shows that the average latency is only 0.33% of the wireless local area network coherence time. Hence, the proposed reconfigurable adaptive SVD engine design is very suitable for high-throughput wireless communication applications. Yen-Liang Chen, Cheng-Zhou Zhan, Ting-Jyun Jheng, An-Yeu Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2011 | Area-Efficient Scalable MAP Processor Design for High-Throughput Multistandard Convolutional Turbo DecodingabstractMost of advanced wireless standards, such as WiMAX and LTE, have adopted different convolutional turbo code (CTC) schemes with various block sizes and throughput rates. Thus, a reconfigurable and scalable hardware accelerator for multistandard CTC decoding is necessary. In this paper, we propose scalable maximum a posteriori algorithm (MAP) processor designs which can support both single-binary (SB) and double-binary (DB) CTC decoding, and handle arbitrary block sizes for high throughput CTC decoding. We first propose three combinations of parallel-window (PW) and hybrid-window (HW) MAP decoding. Moreover, the computational modules and storages of the dual-mode (SB/DB) MAP decoding are designed to achieve a high area utilization. To verify the proposed approaches, a 1.28 mm2dual-mode 2PW-1HW MAP processor is implemented in 0.13 μ m CMOS process. The prototyping chip achieves a maximum throughput rate of 500 Mb/s at 125 MHz with an energy efficiency of 0.19 nJ/bit and an area efficiency of 3.13 bits/mm2. For the multistandard systems, the expected throughput rates of the WiMAX and LTE CTC schemes is achieved by using five dual-mode 2PW-1HW MAP processors. An-Yeu Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Traffic- and Thermal-Aware Run-Time Thermal Management Scheme for 3D NoC SystemsabstractThree-dimensional network-on-chip (3D NoC), the combination of NoC and die-stacking 3D IC technology, is motivated to achieve lower latency, lower power consumption, and higher network bandwidth. However, the length of heat conduction path and power density per unit area increase as more dies stack vertically. Routers of NoC have comparable thermal impact as processors and contributes significant to overall chip temperature. High temperature increases the vulnerability of the system in performance, power, reliability, and cost. To ensure both thermal safety and less performance impact from temperature regulation, we propose a traffic- and thermal-aware run-time thermal management (RTM) scheme. The scheme is composed of a proactive downward routing and a reactive vertical throttling. Based on a validated traffic-thermal mutual-coupling co-simulator, our experiments show the proposed scheme is effective. The proposed RTM can be combined with thermal-aware mapping techniques to have potential for higher run-time thermal safety. Chih-Hao Chao, Kai-Yuan Jheng, An-Yeu Wu |
NOCS | 5 |
| 2009 | A 52-mW 8.29mm2 19-mode LDPC decoder chip for mobile WiMAX applicationsabstractThis paper presents a LDPC decoder chip supporting all 19 modes in Mobile WiMAX applications. An efficient IC design strategy is proposed to reduce 31.25% decoding latency, and enhance hardware utilization ratio from 50% to 75%. In addition, we propose a new early termination scheme that can dynamically adjust the iteration number. The multi-mode chip implemented in 8.29mm2die area can be maximally measured at 83.3MHz with only 52mW power consumption. Xin-Yu Shih, Cheng-Zhou Zhan, An-Yeu Wu |
ASP-DAC | 4 |
| 2009 | A Triple-mode LDPC Decoder Design for IEEE 802.11n SYSTEMabstractThis paper shows a triple-mode LDPC decoder design with two design techniques, the matrix reordering algorithm for multi-mode reconfiguration and the single-entry-multiple-data (SEMD) scheme for throughput enhancement. The matrix reordering algorithm can reduce the computational complexity from O(n!) to O(n3). The SEMD can enhance the throughput by m times with small area overhead. With TSMC 0.13 mum CMOS, the proposed design is synthesized in 1.99 mm2area at 172.4 MHz. Min-An Chao, Jen-Yang Wen, Xin-Yu Shih, An-Yeu Wu |
ISCAS | 4 |
| 2009 | A Scalable Built-in Self-test/Self-diagnosis Architecture for 2D-Mesh based Chip Multiprocessor SystemsabstractIn this paper, we proposed a scalable built-in self-test/self-diagnosis architecture, surrounding test ring (STR), to detect and locate faulty FIFOs and faulty MUXs for 2D-mesh based CMP systems. Proposed STR supports 97.79% fault coverage in FIFOs and MUXs and tests with 388 ~ 2886 test cycles in different testing methods and mesh sizes. Shu-Yen Lin, Chan-Cheng Hsu, An-Yeu Wu |
ISCAS | 3 |
| 2008 | High-performance scheduling algorithm for partially parallel LDPC decoderabstractIn this paper, we propose a new scheduling algorithm for the overlapped message passing decoding, which can be applied to general low-density parity check (LDPC) codes. The partially parallel LDPC architecture is commonly used for reducing the area cost of the processing units. The dependency of two kinds of processing units, check node unit (CNU) and bit node unit (BNU), should be considered to enhance the hardware utilization efficiency (HUE). Based on the properties of the parity check matrix of LDPC codes, the updating calculation of the CNU and BNU can be overlapped to reduce the decoding latency by enhancing the HUE with the matrix scheduling algorithm. By applying our proposed LDPC scheduling algorithm to a (1944, 972)-irregular LDPC code, we can get about 60% throughput gain in average without any performance degradation. Cheng-Zhou Zhan, Xin-Yu Shih, An-Yeu Wu |
ICASSP | 3 |
| 2008 | Cost-effective echo and NEXT canceller designs for 10GBASE-T ethernet systemabstractIn this paper, new echo and NEXT cancellers are proposed for echo and NEXT cancellation in full-duplex digital transmission over 10GBASE-T system. The proposed cancellation schemes inherit the concept of the adaptive interpolated FIR (AIFIR)-based crosstalk canceller, where the long portion is modeled by an adaptive sparse FIR filter. Furthermore, we also employ the channel shortening technique to shorten the impulse response of crosstalk interference. Hence, the complexity of echo and NEXT cancellers can be greatly reduced. Simulations show that compared with the conventional architecture, although the performance is degraded by 1.5 dB, it still meets the SNR requirement in 10GBASE-T system. The complexity reduction of the proposed echo and NEXT cancellation schemes in arithmetic is about 70% and 65% respectively. The saving of hardware complexity results in computationally efficient VLSI implementation of the echo and NEXT cancellers in 10GBASE-T system. Yen-Liang Chen, Cheng-Zhou Zhan, An-Yeu Wu |
ISCAS | 3 |
| 2008 | Low-power traceback MAP decoding for double-binary convolutional turbo decoderabstractConvolutional turbo decoding requires large data access and consumes large memories. To reduce the size of the metrics memory, the traceback MAP decoding is introduced for double-binary convolutional turbo codes without losing the correction performance. The traceback technique reduces the metrics memory size with no other checkers which prolong the decoding latency. Two proposed traceback structures have a tradeoff between the power and operating frequency. The traceback structures can achieve around 20% power reduction of the metrics memory and around 7% power reduction of the decoders for WiMAX standard. An-Yeu Wu |
ISCAS | 3 |
| 2008 | An efficient methodology to evaluate nanoscale circuit fault-tolerance performance based on belief propagationabstractAs silicon circuits quickly approach their physical limitations, researchers are actively looking for novel building blocks to develop nanocircuits. However, future nanoelectronic circuits are more error-prone than conventional CMOS designs because of their self-assembly design. To help design fault-tolerant nanoscale circuits, new circuit design and testing tools are needed. In this paper, an efficient methodology to evaluate nanoscale circuit fault tolerance based on belief propagation (BP) algorithm is proposed. Compared with existing approaches, the BP algorithm is more efficient in terms of memory requirements and CPU times. The proposed methodology can be easily run on multiple CPUs to achieve parallel processing and thus further reduces simulation time. Huifei Rao, Jie Chen 0002, Vicky H. Zhao, Woon Tiong Ang, I-Chyn Wey, An-Yeu Wu |
ISCAS | 6 |
| 2008 | Traffic-Balanced Routing Algorithm for Irregular Mesh-Based On-Chip NetworksabstractOn-chip networks (OCNs) have been proposed to solve the increasing scale and complexity of the designs in nanoscale multicore VLSI designs. The concept of irregular meshes is an important issue because IPs of different sizes may be supported by various vendors. In order to solve routing problems in irregular meshes, modified routing algorithms to detour oversized IPs (OIPs) are needed. However, directly applying fault-tolerant routing algorithms may cause two serious problems: 1) heavy traffic loads around OIPs and 2) unbalanced traffic loads in irregular meshes. In this paper, we propose an OIP avoidance prerouting (OAPR) algorithm to solve the aforementioned problems. The proposed OAPR can make traffic loads evenly spread on the networks and shorten the average paths of packets. Therefore, the networks using the OAPR have lower latency and higher throughput than those using fault- tolerant routing algorithms. In our experiments, four different cases are simulated to demonstrate that the proposed OAPR improves 13.3 percent to 100 percent sustainable throughputs than two previous fault-tolerant routing algorithms. Moreover, the hardware overhead of the OAPR is less than 1 percent compared to the cost of a whole router. Hence, the proposed OAPR algorithm has good performance and is practical for irregular mesh-based OCNs. Shu-Yen Lin, Chun-Hsiang Huang, Chih-Hao Chao, Keng-Hsien Huang, An-Yeu Wu |
IEEE Trans. Computers | 5 |
| 2008 | Unified Convolutional/Turbo Decoder Design Using Tile-Based Timing Analysis of VA/MAP KernelabstractTo satisfy the advanced forward-error-correction (FEC) standards, in which the Convolutional code and Turbo code may co-exit, a prototype design of a unified Convolutional/Turbo decoder is proposed. In this paper, we systematically analyze the timing charts of both the Viterbi algorithm and the MAP algorithm. Then, three techniques, including Distribution, Pointer, and Parallel schemes, are introduced; they can be used as flexible tools in timing-chart analysis to either reduce memory size or to increase throughput rate. Furthermore, we propose a tile-based methodology to analyze the key features of timing charts, such as computing/memory units and hardware utilization. On the basis of the timing analysis, we developed a VA/MAP timing chart that has three modes (VA mode, MAP mode, and concurrent VA/MAP mode) by complementing the idle time of both VA and MAP decoding procedures. The new combined timing analysis helps us for constructing a unified component decoder with near 100% utilization rate of the processing element (PE) in both VA/MAP decoding functions. According to the triple-mode VA/MAP timing chart, we construct a triple-mode FEC kernel that can perform both Convolutional/Turbo decoding functions seamlessly for different communication systems. By integrating the FEC kernel with different size of memory, we can construct four types of FEC decoders for different application scenarios, such as 1) standalone Convolutional decoder (VA mode); 2) standalone Turbo decoder (MAP mode); 3) dual- mode Convolutional/Turbo decoder (VA mode and MAP mode); and 4) triple-mode Convolutional/Turbo decoder (VA mode, MAP mode, and concurrent VA/MAP mode). Finally, a prototyping FEC kernel processor that is compliant to 3GPP standard is verified in TSMC 0.18-mum CMOS process in the type of triple-mode FEC decoder. Fan-Min Li, An-Yeu Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2008 | Design and Analysis of Isolated Noise-Tolerant (INT) Technique in Dynamic CMOS CircuitsabstractAlong with the progress of advanced VLSI technology, noise issues in dynamic circuits have become an imperative design challenge. The twin-transistor design, is the current state-of-the-art design to enhance the noise immunity in dynamic CMOS circuits. To achieve the high noise-tolerant capability, in this paper, we propose a new isolated noise-tolerant (INT) technique which is a mechanism to isolate noise tolerant circuits from noise interference. Simulation results show that the proposed 8-bitINTManchester adder can achieve 1.66times average noise threshold energy (ANTE) improvement. In addition, it can save 34% power delay product (PDP) in low signal-to-noise ratio(SNR)environments as compared with the 8-bit twin-transistor Manchester adder under TSMC 0.18-mu m process. I-Chyn Wey, You-Gang Chen, An-Yeu Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | A Power-Aware Reconfigurable Rendering Engine Design with 453MPixels/s, 16.4MTriangles/s PerformanceabstractThis paper presents a power-aware dynamically reconfigurable rendering engine design, which changes power as rendering throughput and image quality change. At algorithm level, a precision-aware shading scheme is proposed to improve the power efficiency of the conventional shading algorithm through the combination of precision detection and fraction masking techniques. At architecture level, a processing element (PE) based scalable architecture is combined with dynamic task scheduling and dispatching techniques to raise hardware utilization rate and reduce computation latency. Finally a prototyping design which delivers 453MPixels/s, 16.4MTriangles/s, 2.24MPixel/mJ is presented Chih-Hao Chao, Yen-Lin Kuo, An-Yeu Wu, Weber Chien |
ISCAS | 3 |
| 2007 | Ensemble Dependent Matrix Methodology for Probabilistic-Based Fault-tolerant Nanoscale Circuit DesignabstractTwo probabilistic-based models, namely the ensemble-dependent matrix model (Chen and Li, 2006), (Patel et al., 2003) and the Markov random field model (Chen et al., 2003), have been proposed to deal with faults in nanoscale system. The MRF design can provide excellent noise tolerance in nanoscale circuit design. However, it is complicated to be applied to model circuit behavior at system level. Ensemble dependent matrix methodology is more effective and suitable for CAD tools development and to optimize nanoscale circuit and system design. In this paper, we show that the ensemble-dependent matrices describe the actual circuit performances when signal errors are present. We then propose a new criterion to compare circuit error-tolerance capability. We also prove that the matrix model and the Markov model converge when signals are digital Huifei Rao, Jie Chen 0002, Changhong Yu, Woon Tiong Ang, I-Chyn Wey, An-Yeu Wu |
ISCAS | 6 |
| 2007 | Low-Latency Quasi-Synchronous Transmission Technique for Multiple-Clock-Domain IP ModulesabstractData transmission on multiple clock domains faces reliable problems. The conventional globally asynchronous locally synchronous (GALS) technique can resolve the problem but has a high latency problem. In this paper, we present a novel asynchronous transmission technique called quasi-synchronous with an adaptive phase mechanism to reduce the transmission latency. Compared with the conventional GALS techniques, the proposed technique saves 50% ~ 83% of latency. It is implemented on standard-cell library by using TSMC 0.18 mum 1P6M CMOS technology Jhao-Ji Ye, You-Gang Chen, I-Chyn Wey, An-Yeu Wu |
ISCAS | 4 |
| 2007 | A New Binomial Mapping and Optimization Algorithm for Reduced-Complexity Mesh-Based On-Chip NetworkabstractThis paper presents an efficient binomial IP mapping and optimization algorithm (BMAP) to reduce the hardware cost of on-chip network (OCN) infrastructure. The complexity of BMAP is O(N2log(N)). Based on our OCN system synthesis flow, the proposed algorithm provides more economic network component mapping in comparison with traditional OCN mapping algorithm. The experimental result shows total traffic on network is reduced by 37% and average network hop count is reduced by 46%. With further optimization, the hardware efficiency is enhanced therefore the total hardware cost of network infrastructure is reduced to 51%~85% Wein-Tsung Shen, Chih-Hao Chao, Yu-Kuang Lien, An-Yeu Wu |
NOCS | 4 |
| 2007 | Joint AGC-Equalization Algorithm and VLSI Architecture for Wirelined Transceiver DesignsabstractTraditional approaches of automatic gain control (AGC) involve estimating the average power or the peak amplitude over an extended time period, which results in high hardware complexity and a long processing time. Moreover, the accuracy of traditional approaches is seriously degraded by noise and intersymbol interference. In this paper, we propose a joint AGC and equalization (Joint AGC-EQ) scheme, in which the AGC circuitry comprises only one-tenth of the area of a traditional AGC. In addition, the total convergence time of the proposed Joint AGC-EQ is only half that of traditional blind equalization. The scheme is already silicon proven for the application of a Fast Ethernet transceiver using Faraday/UMC 0.18-mum cell libraries Jyh-Ting Lai, An-Yeu Wu, Chien-Hsiung Lee |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | DSP engine design for LINC wireless transmitter systemsabstractLinear amplification with nonlinear components (LINC) technique is a linearization technique for power amplifier designs. By using LINC, the nonlinear power amplifier with high power efficiency can be used to amplify the input signal in a transmitter system with high linearity performance. In this paper, we propose a DSP engine design for LINC wireless transmitter under the WCDMA specification. New algorithms and architectures are proposed, and they can reduce the hardware cost efficiently. We also utilize Agilent ADS to perform mixed-mode system simulation to verify our design. The DSP engine is implemented by Verilog and synthesized with UMC 0.18 /spl mu/m process. The DSP engine area is 0.155 mm/sup 2/ and the maximum clock frequency is 117 MHz. Finally, we use a FPGA to verify our design. Kai-Yuan Jheng, Yi-Chiuan Wang, An-Yeu Wu, Hen-Wai Tsao |
ISCAS | 3 |
| 2006 | A portable all-digital pulsewidth control loop for SOC applicationsabstractA cell-based all-digital PWCL is presented in this paper. To improve design effort as well as facilitate system-level integration, the new design can be developed in hardware description language (HDL) and implemented with standard-cell libraries, therefore, easily portable between technologies. In addition, a high-resolution architecture is designed to enhance pulsewidth precision. For different requirements of applications, the characteristic of scalable modulating range allows hardware decision in early stage. The proposed methodology has been proven at UMC 0.18mum CMOS technology. When operated at 350 MHz, the pulse width acquisition ranges from 10% to 85% with 0.9% steps Wei Wang 0252, I-Chyn Wey, Chia-Tsun Wu, An-Yeu Wu |
ISCAS | 4 |
| 2006 | A frequency estimation algorithm for ADPLL designs with two-cycle lock-in timeabstractThis paper presents a frequency-estimation algorithm for the ADPLL designs instead of traditional binary frequency-search algorithm. With the proposed ADPLL architecture and synchronization process, the lock time can be optimized to two cycles. As the reference clock varies or frequency multiplication switches, lock time holds in two reference clock cycles. An implementation of proposed ADPLL design is realized in UMC 0.18 mum 1P6M CMOS technology with core area of 520times530 mum2. The PLL has the frequency range of 140 MHz to 1030 MHz with 22ps DCO resolution Chia-Tsun Wu, Wei Wang 0252, I-Chyn Wey, An-Yeu Wu |
ISCAS | 4 |
| 2006 | Multi-Symbol-Sliced Dynamically Reconfigurable Reed-Solomon Decoder Design Based on Unified Finite-Field Processing ElementabstractReed-Solomon (RS) codes play an important role in providing the error correction and the data integrity in various communication/storage applications. For high-speed applications, most RS decoders are implemented as dedicated application-specified integrated circuits (ASICs) based on parallel architectures, which can deliver high data throughput rate. For lower-speed applications, the RS decoding operations are usually performed by using fine-grained processing elements (PE) controlled by a programmable digital signal processing (DSP) core, which provides high flexibility. In this paper, we propose a novel m-PE multi-symbol-sliced (MSS) RS datapath structure. The m-PE RS architecture is a highly scalable design and can be dynamically reconfigured at 1-PE, 2-PE,...,m/2-PE, and m-PE modes to deliver necessary data throughput rate. With the help of the gated-clock scheme to turn off the idle PEs, the proposed runtime configurable ASIC design provides good tradeoff between the data throughput rate and the power consumption. Hence, it can save energy to extend the battery life of the portable devices. We demonstrate a prototyping design using 4 PEs by using UMC 0.18-/spl mu/m CMOS technology. The design can be dynamically reconfigured to be operated at 1-PE, 2-PE, and 4-PE modes, with performance of 140 Mb/s at 18.91 mW, 280 Mb/s at 28.77 mW, and 560 Mb/s at 48.47 mW, respectively. Compared with existing RS designs, the proposed m-PE RS decoder has better normalized area/power efficiency than most DSP-type and ASIC-type RS designs. The reconfigurable feature makes our design a good candidate for the error control coding (ECC) unit of the storage system in power-aware portable devices. Huai-Yi Hsu, Jih-Chiang Yeo, An-Yeu Wu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2005 | Low cost decision feedback equalizer (DFE) design for Giga-bit systemsabstractThis paper addresses the design of a DFE for Gbits throughput rate in each communication path of a 10G base LX4 Ethernet system. It is well-known that the feedback loop within a DFE limits an upper bound of the achievable speed. For an L-tap feed-backward filter (FBF) with word-length W and M-PAM signal, previous authors have reformulated the FBF as a (log/sub 2/M)/sup L/- to-1 multiplexer. However, the overhead of extra adders and extra multiplexers are as large as (log/sub 2/M)/sup L/. The required hardware overhead would be more severe when the DFE is designed in parallel. In this paper, we propose two new approaches to implement the DFE when Gbit speed is required. The first approach is partial pre-computation, which can trade-off between hardware complexity and computation speed. The second approach is two-stage pre-computation, which can be applied to higher speed applications. We can reduce the hardware overhead to about 2(log/sub 2/M)/sup (-L/2)/ times that of the previous authors, and the iteration bound is 2(log/sub 2/W+2)/L+(log/sub 2/M) multiplexer-delays. Chih-Hsiu Lin, An-Yeu Wu |
ICASSP (3) | 2 |
| 2004 | High-performance VLSI architecture of adaptive decision feedback equalizer based on predictive parallel branch slicer (PPBS) schemeabstractAmong existing works of high-speed pipelined adaptive decision feedback equalizer (ADFE), the pipelined ADFE using relaxed look-ahead technique results in a substantial hardware saving than the parallel processing or Look-ahead approaches. However, it suffers from both the signal-to-noise ratio (SNR) degradation and slow convergence rate. In this paper, we employ the predictive parallel branch slicer (PPBS) to eliminate the dependencies of the present and past decisions so as to reduce the iteration bound of decision feedback loop of the ADFE. By adding negligible hardware complexity overheads, the proposed architecture can help to improve the output mean-square error (MSE) of the ADFE compared with the Relaxed Look-ahead ADFE architecture. Moreover, we show the superior performance of the proposed pipelined ADFE by using theoretical derivations and computer simulation results. A VLSI design example using Avant! 0.35-/spl mu/m CMOS standard cell library is also illustrated. From the post-layout simulation results, we can see that the PPBS scheme requires only 38.4% gate count overhead, but it can help to reduce the critical path from 7.06 to 4.69 ns so as to meet very high-speed data transmission systems. Meng-Da Yang, An-Yeu Wu, Jyh-Ting Lai |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2003 | Mixed-scaling-rotation CORDIC (MSR-CORDIC) algorithm and architecture for scaling-free high-performance rotational operationsabstractIn this paper, we propose the mixed-scaling-rotation CORDIC (MSR-CORDIC) algorithm which merges micro-rotation operation and scaling operation in conventional CORDIC algorithms to eliminate the overhead of the scaling operation. At the system architecture level, we propose the data-path-selection (DPS) strategy for the tradeoff between hardware complexity and quantization error performance. In general, the CORDIC algorithms suffer from the roundoff noise in fixed-wordlength implementations. We propose two schemes to control and reduce the impairment. Our simulation results show that MSR-CORDIC enhances the SQNR performance, computing speed (reducing the iteration number), and reduces the hardware complexity when compared with the newly proposed EEAS-CORDIC algorithm. Zhi-Xiu Lin, An-Yeu Wu |
ICASSP (2) | 2 |
| 2003 | Angle quantization approach for lattice IIR filter implementation and its trellis de-allocation algorithmabstractIn a multiplier-less digital filter implementation, the sign-power-of-two (SPT) scheme can significantly reduce the hardware complexity but this may seriously degrade the filter performance due to the limited number of SPT terms. Since the CORDIC algorithm is similar to the SPT scheme in that they are both constructed by several shift-and-add operations, the performance would also be significantly affected by the number of these operations. In this paper, we propose the modified angle rotator (MAR) scheme; it provides a systematic solution to enhance the precision of quantized angle without additional hardware overhead. Furthermore, we also apply an appropriate optimization procedure, the trellis de-allocation algorithm, in the angle domain to further reduce the unnecessary operations. Our simulation results show that we can save 40% of the number of adders compared with the direct coefficient quantization approach in normalized lattice filter implementation. An-Yeu Wu, I-Hsien Lee, Cheng-Shing Wu |
ICASSP (2) | 1 |
| 2001 | Cost-efficient multiplier-less FIR filter structure based on modified DECOR transformationabstractWe propose a new design approach to implement an FIR filter using canonical sign digit (CSD) multipliers based on modified decorrelating transformation (MDECOR). The direct CSD approach will introduce serious quantization errors since the distribution of CSD numbers is very non-uniform. The proposed MDECOR transformation provides a systematic solution to reduce the dynamic range effectively. By combining the proposed MDECOR transformation followed by CSD quantization, we can avoid the aforementioned quantization problem. As a result, we do not need to employ additional non-zero bits to compensate for the distortion caused by direct CSD quantization, which helps to save the number of adders in VLSI implementations. Furthermore, the MDECOR transformation offers more freedom in the filter design. It can achieve high-precision performance under the same hardware complexity as the direct CSD approach. Our simulation results show that we can save 20% of the number of adders compared with the direct CSD approach. I-Hsien Lee, Cheng-Shing Wu, An-Yeu Wu |
ICASSP | 3 |
| 2001 | A novel trellis-based searching scheme for EEAS-based CORDIC algorithmabstractThe CORDIC algorithm is a well-known iterative method for the computation of vector rotation. For applications that require forward rotation (or vector rotation) only, the extended elementary angle set (EEAS) Scheme provides a relaxed approach to speed up the operation of the CORDIC algorithm. When determining the parameters of EEAS-based CORDIC algorithm, two optimization problems are encountered. In the previous work, the greedy algorithm is suggested to solve these optimization problems. However, for an application that requires high-precision rotation operation, the results generated by the greedy algorithm may not be applicable. We propose a novel searching algorithm to overcome the aforementioned problem, called the trellis-based searching (TBS) algorithm. Compared with the greedy algorithm used in the conventional EEAS-based CORDIC algorithm, the proposed TBS algorithm yields apparent performance improvement. Moreover, the derivation of the error boundary as well as computer simulations are provided to support our arguments. Cheng-Shing Wu, An-Yeu Wu |
ICASSP | 2 |
| 2001 | A unified design framework for vector rotational CORDIC family based on angle quantization processabstractVector rotation is the key operation employed extensively in many digital signal processing applications. We introduce a new design concept called angle quantization (AQ). It can be used as a design index for vector rotational operation, where the rotational angle is known in advance. Based on the AQ process, we establish a unified design framework for cost-effective low-latency rotational algorithms and architectures. Several existing works, such as conventional CORDIC, AR-CORDIC, MVR-CORDIC, and EEAS-based CORDIC, can be fitted into the design framework, forming a vector rotational CORDIC family. Based on the new design framework, we can realize high-speed/low complexity rotational VLSI circuits, whereas without degrading the precision performance in fixed-point implementations. An-Yeu Wu, Cheng-Shing Wu |
ICASSP | 1 |
| 2001 | A cost-effective TEQ algorithm for ADSL systemsabstractThe discrete multitone (DMT) modulation/demodulation scheme is the physical-layer standard of the asymmetric digital subscriber line (ADSL) system. In the DMT transceiver, channel equalization is completed through two steps, namely, time-domain equalization (TEQ) and frequency-domain equalization (FEQ). The TEQ is introduced to shorten the channel response to a pre-defined length. On the contrary, FEQ is used to compensate for the magnitude and phase distortion caused by channel. A cost-effective on-line TEQ training algorithm is proposed. Simulation results show that the proposed TEQ algorithm has comparable performance with other existing on-line algorithms, and it achieves the lowest computational complexity among these TEQ algorithms. Chih-Chi Wang, An-Yeu Wu, Bor-Ming Wang |
ICC | 2 |
| 2000 | Design methodology for Booth-encoded Montgomery module design for RSA cryptosystemabstractIn this paper, a design methodology for the design of a Montgomery module is proposed. We summarize the result in pseudo C-like codes and call it Booth-encoded Montgomery modular multiplication algorithm. Using this algorithm, iteration number is reduced to about n/2 in each Montgomery operation. In addition, we apply the folding and unfolding techniques to shorten the critical path. Finally, we propose the 4 bit-digit-serial pipelined architecture to process RSA encryption/decryption in a more efficient way. The speed of the proposed algorithm is approximately 1.7 times that of most RSA VLSI designs based on original Montgomery modular multiplication algorithm. Jye-Jong Leu, An-Yeu Wu |
ISCAS | 2 |
| 2000 | Modified vector rotational CORDIC (MVR-CORDIC) algorithm and its application to FFTabstractThe CORDIC algorithm is a well-known iterative method for the computation of vector rotation. However, the major disadvantage is its relatively slow computational speed. For applications that require forward rotation (or vector rotation) only, we propose a new scheme, the Modified Vector Rotational CORDIC (MVR-CORDIC) algorithm, to improve the speed performance of CORDIC algorithm. The basic idea of the proposed scheme is to reduce the iteration number directly while maintaining the SQNR performance. This can be achieved by modifying the basic microrotation procedure of the CORDIC algorithm. In addition, three searching algorithms are suggested to find the corresponding directional and rotational sequences. In the example of a 128-point FFT, we have shown that by using the proposed MVR-CORDIC algorithm, the hardware complexity is only 33% compared with conventional CORDIC-based FFT. Meanwhile, the SQNR performance is 7 dB better than the conventional CORDIC approach. Cheng-Shing Wu, An-Yeu Wu |
ISCAS | 2 |
| 1999 | A novel multirate adaptive FIR filtering algorithm and structureabstractA new class of FIR filtering algorithms and VLSI architectures based on the multirate approach were recently proposed. They not only reduce the computational complexity in FIR filtering, but also retain attractive implementation-related properties such as regularity and multiply-and-accumulate structure. In addition, the multirate feature can be applied to low-power/high-speed VLSI implementation. These properties make the multirate FIR filtering very attractive in many DSP and communication applications. In this paper, we propose a novel adaptive filter based on this new class of multirate FIR filtering structures. The proposed adaptive filter inherits the advantages of the multirate structures such as low computational complexity and low-power/high-speed applications. Moreover, the multirate feature helps to improve the convergence property of the adaptive filters. Cheng-Shing Wu, An-Yeu Wu |
ICASSP | 2 |
| 1998 | Cost-efficient parallel lattice VLSI architecture for the IFFT/FFT in DMT transceiver technologyabstractThe discrete multitone (DMT) modulation/demodulation scheme is the standard transmission technique in the application of asymmetric digital subscriber lines (ADSL). Although the DMT can achieve a higher data rate compared with other modulation/demodulation schemes, its computational complexity is too high for cost-efficient implementations. For example, it requires 512-point IFFT/FFT as the modulation/demodulation kernel. The large block size results in heavy computational load in running programmable DSP processors. It also makes VLSI implementation not feasible. We derive the parallel lattice structure for the IFFT/FFT based on the time-recursive approach. The resulting architectures are regular, modular, and without global communications so that they are very suitable for VLSI implementation. Also, the proposed structure requires only 11% of the multipliers and 9% of the adders compared with the direct implementation approach. An-Yeu Wu, Tsun-Shan Chan |
ICASSP | 1 |
| 1998 | Algorithm-based low-power and high-performance multimedia signal processingabstractLow power and high performance are the two most important criteria for many signal-processing system designs, particularly in real-time multimedia applications. There have been many approaches to achieve these two design goals at many different implementation levels ranging from very-large-scale-integration fabrication technology to system design. We review the works that have been done at various levels and focus on the algorithm-based approaches for low-power and high-performance design of signal processing systems. We present the concept of multirate computing that originates from filterbank design, then show how to employ it along with the other algorithmic methods to develop low-power and high-performance signal processing systems. The proposed multirate design methodology is systematic and applicable to many problems. We demonstrate that multirate computing is a powerful tool at the algorithmic level that enables designers to achieve either significant power reduction or high throughput depending on their choice. Design examples on basic multimedia processing blocks such as filtering, source coding, and channel coding are given. A digital signal-processing engine that is an adaptive reconfigurable architecture is also derived from the common features of our approach. Such an architecture forms a new generation of high-performance embedded signal processor based on the adaptive computing model. The goal of this paper is to demonstrate the flexibility and effectiveness of algorithm-based approaches and to show that the multirate approach is an effective and systematic design methodology to achieve low-power and high throughput signal processing at the algorithmic and architectural level. K. J. Ray Liu, An-Yeu Wu, Arun Raghupathy, Jie Chen 0002 |
Proc. IEEE | 2 |
| 1998 | System architecture of an adaptive reconfigurable DSP computing engineabstractIn this paper, we present the system architecture of an adaptive reconfigurable DSP computing engine for numerically intensive front-end audio/video communications. The proposed system is a massively parallel architecture that is capable of performing most low-level computationally intensive data processing including finite impulse response/infinite impulse response (FIR/IIR) filtering, subband filtering, discrete orthogonal transforms (DT), adaptive filtering, and motion estimation for the host processor in DSP applications. Since the properties of each programmed function such as parallelism and pipelinability have been fully exploited in this design, the computational speed of this computing engine can be as fast as ASIC designs that are optimized for individual specific applications. We also show that the system can be easily configured to perform multirate FIR/IIR/DT operations at negligible hardware overhead. Since the processing elements are operated at half of the input data rate, we are able to double the processing speed on-the-fly based on the same system architecture without using high-speed/full-custom circuits. The programmable/high-speed features of the proposed design make it very suitable for cost-effective video-rate DSP applications. An-Yeu Wu, K. J. Ray Liu, Arun Raghupathy |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1998 | Algorithm-based low-power transform coding architectures: the multirate approachabstractIn most low-power VLSI designs, the supply voltage is usually reduced to lower the total power consumption. However, the device speed will be degraded as the supply voltage goes down. In this paper, we propose new algorithmic-level techniques to compensate the increased delays based on the multirate approach. We apply the technique of polyphase decomposition to design low-power transform coding architectures, in which the transform coefficients are computed through decimated low-speed input sequences. Since the operating frequency is M-times slower than the original design while the system throughput rate is still maintained, the speed penalty can be compensated at the architectural level. We start with the design of low-power multirate discrete cosine transform (DCT)/inverse discrete cosine transform (IDCT) VLSI architectures. Then the multirate low-power design is extended to the modulated lapped transform (MLT), extended lapped transform (ELT), and a unified low-power transform coding architecture. Finally, we perform finite-precision analysis for the multirate DCT architectures. The analytical results can help us to choose the optimal wordlength for each DCT channel under required signal-to-noise ratio (SNR) constraint, which can further reduce the power consumption at the circuit level. The proposed multirate architectures can also be applied to very high-speed block discrete transforms in which only low-speed operators are required. An-Yeu Wu, K. J. Ray Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1995 | Algorithm-based low-power transform coding architecturesabstractIn most low-power VLSI designs, the supply voltage is usually reduced to lower the total power consumption. However, the device speed will be degraded as the supply voltage goes down. In this paper, we propose new algorithmic-level techniques for compensating the increased delays based on the multirate approach. We will show how to compute most of the discrete sinusoidal transforms through the decimated low-speed sequences with reasonable linear hardware overhead. For the case where the decimation factor is equal to two, the overall power consumption can be reduced to about one-third of the original design. The resulting multirate low-power architectures are regular, modular, and free of global communications. Such properties are very suitable for VLSI implementations. The proposed architectures can also be applied to very high-speed block transforms where only low-speed operators are required. An-Yeu Wu, K. J. Ray Liu |
ICASSP | 1 |
| 1995 | Parallel programmable video co-processor designabstractModern video applications call for computationally intensive data processing at very high data rate. In order to meet the high-performance/low-cost constraints, the state-of-the-art video processor should be a programmable design which performs various tasks in video applications without sacrificing the computational power and the manufacturing cost in exchange for such flexibility. In this paper, we present a programmable video co-processor design that is capable of performing FIR/IIR filtering, subband filtering, and most discrete orthogonal transforms (DT), for the host processor in video applications. The computational speed of this co-processor is as fast as that of ASIC designs which are optimized for individual specific applications. We also show that the system can be easily reconfigured to perform multirate FIR/IIR/DT operations at negligible hardware overhead. Hence, we can either double the processing speed on the fly based on the same processing elements, or apply this feature to the low-power implementation of this co-processor. An-Yeu Wu, K. J. Ray Liu, Arun Raghupathy, Shang-Chieh Liu |
ICIP | 1 |
| 1994 | A Low-Power and Low-Complexity DCT/IDCT VLSI Architecture Based On Backward Chebyshev RecursionabstractA low-power parallel VLSI structure for DCT/IDCT is proposed. By treating the transformations as the evaluation of the Chebyshev series, and exploiting the Backward Chebyshev Recursion (BCR), we can reduce the total number of multipliers (N+1 for IDCT, 2N-2 for DCT). The property of BCR is also used to compute the DCT/IDCT through the down-sampled even and odd sequences. Since the operation frequency for the down-sampled sequences is two times slower, the speed penalty caused by the low-voltage design can be compensated at the architectural level. The total multipliers required for the low-power design is only 2N+1 for IDCT and 3N-3 for DCT. Extension to downsampling-by-4 is also achievable at a reasonable increase in hardware complexity.> An-Yeu Wu, K. J. Ray Liu |
ISCAS | 1 |
| 1993 | A Multi-layer 2-D Adaptive Filtering Architecture Based on McClellan Transformation
K. J. Ray Liu, An-Yeu Wu |
ISCAS | 2 |