EDBT 2026 Demo / reviewers in the wild / expert
Chi-Ying Tsui
dblp:26/1737
· DBLP profile ↗
160ranked-venue papers
14as first author
36since 2021 · last 2026
0000-0002-8024-2637ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 137 · 14 first-author · 31 since 2021Computer networks · 11 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9Applied, interdisciplinary, general and emerging computing · 9 · 1 since 2021Software engineering, systems software and programming languages · 8 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploiting the Irregular Input Sparsity in Systolic Array-based DNN Accelerators via Local Soft PoolingabstractOne promising approach to mitigating the computational complexity of deep neural networks is to leverage the sparsity of input activations that results from the application of the ReLU function. However, the irregular distribution of zero-valued inputs poses a challenge for efficient implementation in existing regular architectures, such as systolic arrays. Previous works usually depend on specialized architectures to bypass the redundant computations during runtime. In contrast to these prior strategies, we propose a local soft pooling method to efficiently exploit the irregular input sparsity in systolic array-based architectures. Through local soft pooling, adjacent input rows can be safely merged at runtime, compressing the sparse input matrix into a compact format that is only 1/3 to 1/2 of its original size. The compact matrix can then be directly fed into the systolic array for computation. A computation saving of 67.78% is achieved across various networks on both CIFAR-10 and ImageNet with negligible accuracy loss. As a result, the throughput and energy efficiency are improved by 2.72 and 2.07 times, respectively. Desheng Fu, Jingbo Jiang, Jingyang Zhu, Xizi Chen, Chi-Ying Tsui |
ASP-DAC | 6 |
| 2026 | DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI ModelsabstractStochastic computing (SC) offers hardware simplicity but suffers from low throughput, while high-throughput Digital Computing-in-Memory (DCIM) is bottlenecked by costly adder logic for matrix-vector multiplication (MVM). To address this trade-off, this paper introduces a digital stochastic CIM (DS-CIM) architecture that achieves both high accuracy and efficiency. We implement signed multiply-accumulation (MAC) in a compact, unsigned OR-based circuit by modifying the data representation. Throughput is enhanced by replicating this low-cost circuit 64 times with only a 1× area increase. Our core strategy, a shared Pseudo Random Number Generator (PRNG) with 2D partitioning, enables single-cycle mutually exclusive activation to eliminate OR-gate collisions. We also resolve the 1s saturation issue via stochastic process analysis and data remapping, significantly improving accuracy and resilience to input sparsity. Our high-accuracy DS-CIM1 variant achieves 94.45% accuracy for INT8 ResNet18 on CIFAR-10 with a root-mean-squared error (RMSE) of just 0.74%. Meanwhile, our high-efficiency DS-CIM2 variant attains an energy efficiency of 3566.1 TOPS/W and an area efficiency of 363.7 TOPS/mm2, while maintaining a low RMSE of 3.81%. The DS-CIM capability with larger models is further demonstrated through experiments with INT8 ResNet50 on ImageNet and the FP8 LLaMA-7B model. Kunming Shao, Jiangnan Yu, Zhipeng Liao, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 8 |
| 2026 | BioSeek: A Design Generation Framework of Biosignal Processors with Large-Language Models for Edge Healthcare ApplicationsabstractDeep neural network (DNN)-based methodologies have shown impressive performance and robustness in the detection of abnormalities and decoding of multi-modal biosignals. While the use of DNNs provides promising classification and decoding capabilities, it also introduces significant design and cost challenges for the implementation of biomedical System on Chips (SoC). To address the increasing demand for advanced and efficient DNN-based healthcare solutions at the edge, we propose BioSeek, an agile design generation framework enhanced by cutting-edge large-language models (LLM). BioSeek offers a comprehensive solution to the design challenges associated with biosignal processors. The effectiveness of BioSeek is evaluated through the design generation of both application-specific and versatile biosignal processors, demonstrating performance that is competitive with existing solutions. Fengshi Tian, Jiakun Zheng, Hui Wu 0010, Zilu Liu, Jinbo Chen 0002, Shiqi Zhao 0001, Jie Yang 0033, Mohamad Sawan, Chi-Ying Tsui, Kwang-Ting Cheng |
ISCAS | 9 |
| 2026 | Configurable Dataflow and Adaptive Mapping Optimization for Hybrid ReRAM and SRAM Compute-in-Memory AcceleratorabstractHybrid compute-in-memory (CIM) designs have been proposed recently to facilitate the storing of large number of weights of a neural network on-chip. Notably, ReSCIM wang2024res pairs an SRAM cell with a dedicated ReRAM crossbar, allowing ReRAM to serve as the local storage, significantly enhancing the storage capacity of the SRAM-CIM. The SRAM is custom-designed not only to serve as a storage element for CIM but also to function as a sense amplifier to retrieve the data from the ReRAM, which enables super high bandwidth of weight data loading into the CIM engine. However, existing mapping tools for CIM are inadequate for ReSCIM since they do not fully exploit the unique hardware characteristics and advantages of this novel architecture. In this work, we propose an analytical energy and latency model, which incorporates four key factors: hardware, workload, dataflow, and mapping (HWDM), for executing inference of neural network on the ReSCIM accelerator. Specifically, we first characterize the ReSCIM accelerator hardware specifications and the neural network layers. Next, we introduce three dataflows for ReSCIM, leveraging the high weight-loading bandwidth to reduce memory access for various workloads and layer types. Finally, we develop an algorithm to generate optimal mapping and dataflow strategies aimed at minimizing latency or energy consumption. Using our HWDM model, we design a tile-based ReSCIM accelerator and conduct extensive simulations to obtain the cycle-accurate latency and gate-level energy consumption metrics for inference across different neural networks. We conduct design space exploration (DSE) using the HWDM model on a comprehensive set of benchmarks to minimize inference energy or latency. Experimental results show that our optimal ReSCIM accelerator achieves a 44% reduction in EDP reduction compared to the weight-stationary and fixed mapping baseline for SEResNet50. Moreover, our design exhibits 1.74× higher energy efficiency than the state-of-the-art hybrid TL-nvSRAM wang2023tl accelerator on ResNet 18. Jingyu He, Kunming Shao, Kwang-Ting Cheng, Chi-Ying Tsui |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-Fly Aligned-Mantissa Bitwidth PredictionabstractFP8 low-precision formats have gained significant adoption in transformer inference and training. However, existing digital compute-in-memory (DCIM) architectures face challenges in supporting variable FP8 aligned-mantissa bitwidths, as unified alignment strategies and fixed-precision multiply accumulate (MAC) units struggle to handle input data with diverse distributions. This work presents a flexible FP8 DCIM accelerator with three innovations: 1) a dynamic shift-aware bitwidth prediction (DSBP) with on-the-fly input prediction that adaptively adjusts weight (2/4/6/8b) and input ($2\sim 12$b) aligned-mantissa precision; 2) a FIFO-based input alignment unit (FIAU) replacing complex barrel shifters with pointer-based control; and 3) a precision-scalable INT MAC array achieving flexible weight precision with minimal overhead. Implemented in 28-nm CMOS with a$64~\times ~96$CIM array, the design achieves 20.4 TFLOPS/W for fixed E5M7, demonstrating$2.8\times $higher FP8 efficiency than previous work while supporting all FP8 formats. Results on Llama-7b show that the DSBP achieves higher efficiency than fixed bitwidth mode at the same accuracy level on both BoolQ and Winogrande datasets, with configurable parameters enabling flexible accuracy–efficiency tradeoffs. Kunming Shao, Zhipeng Liao, Xijie Huang, Kwang-Ting Cheng, Chi-Ying Tsui, Yi Zou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit SynthesisabstractDigital Computing-in-Memory (DCIM) is an innovative technology that integrates multiply-accumulation (MAC) logic directly into memory arrays to enhance the performance of modern AI computing. However, the need for customized memory cells and logic components currently necessitates significant manual effort in DCIM design. Existing tools for facilitating DCIM macro designs struggle to optimize subcircuit synthesis to meet user-defined performance criteria, thereby limiting the potential system-level acceleration that DCIM can offer. To address these challenges and enable the agile design of DCIM macros with optimal architectures, we present SynDCIM - a performance-aware DCIM compiler that employs multi-spec-oriented subcircuit synthesis. SynDCIM features an automated performance-to-layout generation process that aligns with user-defined performance expectations. This is supported by a scalable subcircuit library and a multi-spec-oriented searching algorithm for effective subcircuit synthesis. The effectiveness of SynDCIM is demonstrated through extensive experiments and validated with a test chip fabricated in a 40nm CMOS process. Testing results reveal that designs generated by SynDCIM exhibit competitive performance when compared to state-of-the-art manually designed DCIM macros. Kunming Shao, Fengshi Tian, Jiakun Zheng, Jia Chen 0032, Jingyu He, Hui Wu 0010, Jinbo Chen 0002, Xihao Guan, Fengbin Tu, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 15 |
| 2025 | LDPC Code Optimisation for OTFS Modulation with MP DetectionabstractThe orthogonal time frequency space (OTFS) modulation is a promising technique to provide reliable communications in high-mobility scenarios. However, the performance analysis for coded OTFS systems is not available in the literature. This paper investigates the extrinsic information transfer (EXIT) behaviour of LDPC coded OTFS systems with message passing (MP) detection. Different from conventional EXIT analysis, which is normally obtained by Monte Carlo simulation, the exact distribution of the extrinsic information for MP detection is presented in this paper. Then, the EXIT function for the LDPC decoder is given with the prior information of the MP detector, which illustrates the convergence behaviour of the LDPC coded OTFS systems. Finally, an algorithm is proposed for calculating the decoding threshold for LDPC coded OTFS modulation, which can be used to design LDPC codes for OTFS systems. Numerical results verify the accuracy of the EXIT analysis, and the 5G NR LDPC code optimised by the proposed method demonstrates 1 dB to 1.5 dB performance gains over the original codes. Shenghui Song 0001, Chi-Ying Tsui, Jinhong Yuan |
GLOBECOM | 3 |
| 2025 | FAS-RIS-Aided Multi-User Systems With Linear Precoding: Random Matrix Analysis and Two-Timescale DesignabstractThe reconfigurability of fluid antenna systems (FASs) and reconfigurable intelligent surfaces (RISs) can be jointly utilized to achieve unprecedented degrees of freedom for wireless communication systems. However, adjusting fluid antennas and RISs based on instantaneous channel state information (CSI) is highly challenging. To tackle this challenge, we propose a two-timescale approach for FAS-RIS-aided multi-user systems with regularized zero-forcing (RZF)/zero-forcing (ZF) precoding, where only statistical CSI is required for FAS and RIS optimization. To achieve this goal, we first obtain the closed-form evaluation for the ergodic sum rate (ESR) of FAS-RIS aided multi-user systems with RZF/ZF precoding by exploiting random matrix theory (RMT). Then, we propose an ESR maximization algorithm by jointly optimizing the port selection for FASs, phase shifts at the RIS, and regularization factor of RZF. Numerical results validate the approximation accuracy of the derived ESR evaluation and demonstrate that the performance enhancement benefiting from the joint design of FASs and RISs becomes more prominent when the number of users becomes larger. Xin Zhang 0039, Dongfang Xu, Jingjing Wang 0001, Shenghui Song 0001, Chi-Ying Tsui, Derrick Wing Kwan Ng, Mérouane Debbah |
GLOBECOM | 5 |
| 2025 | DPE-CIM: Compute-In-Memory Accelerator using Dynamic Posit Encoding and Speculative AlignmentabstractIn this study, we propose two novel approaches to address the memory wall of AI accelerators. First, based on Posit, we introduce a new format called dynamic Posit encoding (DPE), which dynamically extends the dynamic range of its representation at run time with minimal hardware overhead. Using two exponent encoding schemes, DPE accommodates the data distribution with lower quantization error compared to regular Posit. Second, we propose a compute-in-memory (CIM) architecture to implement DPE multiply-and-accumulate (MAC) computation to reduce weight data movement. Traditional CIM proposed for floating-point-alike MAC computation uses a comparator tree (CT) to compute the maximum exponent, enabling the CIM to locus on integer MAC. However, the CT-based design has poor scalability as the number of inputs increases. To address this, we propose a speculative input alignment design that significantly reduces the delay, area, and power consumption for the max exponent computation. We show that DPE outperforms state-of-the-art quantization approaches across various neural network models through software evaluations. Hardware synthesis and simulation results further illustrate that our approach achieves significant energy efficiency and area efficiency improvement compared to the state-of-the-art posit processing element. Jingyu He, Kwang-Ting Cheng, Chi-Ying Tsui |
ISCAS | 3 |
| 2025 | A Dual-Mode One-stage R3 Rectifier with Wide Loading Range for Implantable Medical DevicesabstractA 13.56 MHz wireless power receiver with a reconfigurable resonant regulating (R3) rectifier for implantable medical devices is presented. The receiver adopts the 0X/1X dual-mode modulation and achieves voltage rectification and regulation in one stage. The PWM controller consists of a ramp generator and a Type-II compensator. Hybrid on/off-delay compensation with analog Vdd-sensing and digital multiple pulsing blocking reduces the delay of the active-diode path. Embedded level-shift drivers adaptively provide high voltage to turn off the high-side PMOS active diodes with the mode signal, and they can be decoupled from driving the power PMOS transistors, thus enhancing the overall efficiency. Measurement results show steady delivery of 5V at a load of 400Ω and stable load transients between 1mA to 20mA. The measured maximum power conversion efficiency (PCE) is 92% for a load resistance of 167Ω and the load resistance can be as large as 10kΩ. Tsz Fai Kwok, Pok Man Leung, Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 5 |
| 2025 | Analysis and Prevention of Coupling-Dependent Data Flipping in Series-Series Resonant Wireless Power Transfer SystemsabstractLoad shift keying (LSK) is commonly used in wireless power transfer (WPT) systems for backscattering information from the receiver back to the transmitter. However, when the coupling coefficient (k) between the coupling coils falls below a threshold value (kDF), the demodulated LSK data can unexpectedly flip from '1' to '0' and '0' to '1'. This phenomenon is referred to as coupling-dependent data flipping (CDDF). This research investigates the factors that lead to CDDF in series-series resonant WPT systems, by taking into account of parasitic parameters of the coupled link and validates the analysis through SPICE simulations. To mitigate CDDF, we propose a carrier-frequency auto-tuning scheme that is also verified by simulation results. Sayan Sarkar, Fengshi Tian, Wing-Hung Ki, Chi-Ying Tsui, Yang Liu 0061 |
ISCAS | 5 |
| 2025 | A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight CombinationabstractDeploying mixed-precision neural networks on edge devices is friendly to hardware resources and power consumption. To support fully mixed-precision neural network inference, it is necessary to design flexible hardware accelerators for continuous varying precision operations. However, the previous works have issues on hardware utilization and overhead of reconfigurable logic. In this paper, we propose an efficient accelerator for 2 ∼ 8-bit precision scaling with serial activation input and parallel weight preloaded. First, we set two loading modes for the weight operands and decompose the weight into the corresponding bitwidths, which extends the weight precision support efficiently. Then, to improve hardware utilization of low-precision operations, we design the architecture that performs bit-serial MAC operation with systolic dataflow, and the partial sums are combined spatially. Furthermore, we designed an efficient carry save adder tree supporting both signed and unsigned number summation across rows. The experiment result shows that the proposed accelerator, synthesized with TSMC 28nm CMOS technology, achieves peak throughput of 4.09TOPS and peak energy efficiency of 68.94TOPS/W at 2/2-bit operations. Kunming Shao, Fengshi Tian, Kwang-Ting Cheng, Chi-Ying Tsui, Yi Zou 0001 |
ISCAS | 5 |
| 2025 | NeuroEye: A 54.59mW, 12200FPS Event-Driven Near-Sensor Eye-Tracking Processor with Pipelined Spatial-Temporal Spike-StreamingabstractThis paper presents a design of an eye tracking system based on neuromorphic computing to enhance user interaction in augmented reality (AR) and virtual reality (VR) environments. Traditional methods face challenges of high computational demands and power consumption. To address these issues, we propose a fully-spike eye-tracking system that utilizes dynamic vision sensors (DVS) for asynchronous pixel-level change detection, thereby reducing data redundancy and improving temporal resolution. We proposed a pipelined processor specifically tailored for handling DVS events and Spiking Neural Network (SNN) computations. Our spatial-temporal spike-streaming architecture enables cascaded computation across all layers, achieving high energy efficiency and high frame rate in eye-tracking tasks. Implemented in a 40nm CMOS process, NeuroEye demonstrates up to 12200 frame-per-second (FPS) and 4.47uJ/frame energy efficiency with 54.59mW power consumption in post-layout evaluations. Jiakun Zheng, Fengshi Tian, Jinbo Chen 0002, Chaoming Fang, Jie Yang 0033, Mohamad Sawan, Kwang-Ting Cheng, Chi-Ying Tsui |
ISCAS | 9 |
| 2025 | DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM ComputationabstractRetrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge retrieval but faces challenges on edge devices due to high storage, energy, and latency demands. Computing-in-Memory (CIM) offers a promising solution by storing document embeddings in CIM macros and enabling in-situ parallel retrievals but is constrained by either low memory density or limited computational accuracy. To address these challenges, we present DIRC-RAG, a novel edge RAG acceleration architecture leveraging Digital In-ReRAM Computation (DIRC). DIRC integrates a high-density multi-level ReRAM subarray with an SRAM cell, utilizing SRAM and differential sensing for robust ReRAM readout and digital multiply-accumulate (MAC) operations. By storing all document embeddings within the CIM macro, DIRC achieves ultra-low-power, single-cycle data loading, substantially reducing both energy consumption and latency compared to off-chip DRAM. A query-stationary (QS) dataflow is supported for RAG tasks, minimizing on-chip data movement and reducing SRAM buffer requirements. We introduce error optimization for the DIRC ReRAM-SRAM cell by extracting the bit-wise spatial error distribution of the ReRAM subarray and applying targeted bit-wise data remapping. An error detection circuit is also implemented to enhance readout resilience against device-and circuit-level variations.Simulation results demonstrate that DIRC-RAG under TSMC 40nm process achieves an on-chip non-volatile memory density of 5.18Mb/mm2and a throughput of 131 TOPS. It delivers a 4MB retrieval latency of 5.6μs/query and an energy consumption of 0.956μJ/query, while maintaining the retrieval precision. Kunming Shao, Zhipeng Liao, Jiangnan Yu, Xijie Huang, Jingyu He, Fengshi Tian, Yi Zou 0001, Kwang-Ting Cheng, Chi-Ying Tsui |
ISLPED | 12 |
| 2025 | Exploiting the Memory-Compute-Coupling Feature for CIM Accelerator Design OptimizationabstractSRAM computing-in-memory (CIM) accelerators have evolved as a promising solution to the memory wall problem in neural network (NN) models. By integrating memory and compute resources in each macro, CIM accelerators offer massive in-situ computing parallelism and large memory capacity, enabling spatial mapping with layer fusion and potentially keeping layers stationary in CIM. However, CIM’s memory-compute coupling (MCC) feature poses challenges in designing CIM accelerators. From an architecture aspect, designers must balance CIM’s memory and compute resources by optimizing the macro’s memory-compute ratio (MCR) configuration across diverse scenarios. From a mapping aspect, conventional mappings, which allocate each macro exclusively to each layer, face two major problems: a layer-fusion dilemma (the accelerator suffers from excessive memory access due to layer replications or performance degradation due to load imbalance) and a layer-eviction issue (storing layers stationary in CIM is usually infeasible due to limited CIM capacity). To address these challenges, this paper introduces MCC-DSE, an MCC-aware Design Space Exploration framework for architecture-mapping co-optimization of CIM accelerators. We also propose a three-axis CIM division mapping, which interleaves multiple layers in each macro to concurrently optimize memory access and performance during layer fusion as well as reserves a part of CIM memory in each macro for layer pinning. Compared to baseline architecture and mapping, MCC-DSE shows a 1.4x 8.3x EDP reduction across various workloads and chip areas. Moreover, MCC-DSE provides insights into CIM accelerator optimization, such as selecting optimal MCR and configuring CIM dynamically for different scenarios. Yongkun Wu, Jia Chen 0032, Zhenhua Zhu 0002, Jingyu He, Pingcheng Dong, Yonghao Tan, Xin Zhao 0044, Liang Chang 0002, Yu Wang 0002, Fengbin Tu, Chi-Ying Tsui, Kwang-Ting Cheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 12 |
| 2025 | High Efficiency Active Rectifier Using SAR Digital-to-Time Converter for Wireless Power Transfer SystemabstractThis paper presents a 13.56MHz active rectifier used in a wireless power transfer system for implantable medical devices that employs digital-to-time converters in replacing analog comparators to generate delay-compensated gate control signals for the power transistors. The low-power digital controller employs a successive approximation register (SAR) to generate digital codes for delay compensation, achieving zero-voltage switching and eliminating reverse conduction loss. Fabricated in a standard 65nm CMOS process, the proposed rectifier has an active area of 0.02mm2. The quiescent power is$13.6\mu $W, 15 times lower than the traditional design. The power transfer efficiency is maintained above 90% from 3mW to 40mW with maximum efficiency of 95% at 16mW. Under light load condition, the proposed design achieves more than 20% efficiency enhancement compared to the rectifier without delay compensation. Yang Liu 0061, Chenchang Zhan, Chi-Ying Tsui, Wing-Hung Ki |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Hysteresis-Dependent Synchronized Load Shift Keying and Reconfigurable Class-D Power Amplifier-Based Fully Integrated Adaptive Control in Wireless Power Transfer SystemabstractA 13.56-MHz wireless power transfer (WPT) system with fully integrated transmitter (TX) and receiver (RX) chips is presented. The receiver’s output voltage is locally regulated using a linear current-sink-based regulator, while global power regulation is achieved at the transmitter through a hybrid control strategy that combines constant off-time and hysteretic control for a reconfigurable power amplifier. Synchronized load-shift keying at the receiver improves the relative change in the primary current of the transmitter by >15%. The adaptive digitally controlled active rectifier achieves a voltage conversion ratio (VCR) and power conversion efficiency (PCE) of 0.92 and 92.4%, respectively, for a 200 Ω load resistance. The end-to-end efficiency is improved by 25% at heavy load and 14% at light load by enabling TXglobal power regulation. Both TXand RXchips were fabricated in the BCDlite 180 nm process with 1.8 V/5 V devices. This system achieves a greater operating distance, higher output power, and faster load-transient response while significantly reducing circuit and system design complexity. Sayan Sarkar, Wing-Hung Ki, Chi-Ying Tsui |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | Improved Step-GRAND: Low-Latency Soft-Input Guessing Random Additive Noise DecodingabstractThe ultrareliable low-latency communication (URLLC) application scenario requires the adoption of short linear block codes to satisfy the low-latency requirements. Guessing random additive noise decoding (GRAND) is a prominent universal decoding solution for short linear block codes that lends itself to efficient hardware implementations. GRAND-based hardware implementations generally offer reduced average decoding latency but their high worst-case (W.C.) latency renders them unsuitable for deployment in mission-critical applications. This article presents an improved version of step-GRAND, a soft-input variant of GRAND that features a novel test error pattern (TEP) generating approach. A novel very large-scale integration (VLSI) architecture is developed for the execution of the improved step-GRAND algorithm with reduced W.C. decoding latency. Application specific integrated circuit (ASIC) implementation results, employing low-power (LP) TSMC 65-nm CMOS technology, demonstrate that the proposed improved step-GRAND can achieve an average decoding latency as low as 10 ns for decoding a$(128,105)$linear block code at a target frame error rate (FER) of$10^{-7}$, while the W.C. decoding latency can reach$300~\text {ns}\sim 1~\mu \text { s}$depending on the parametric settings. Compared with the previously proposed baseline soft-input ordered reliability bits GRAND (ORBGRAND) hardware implementation with similar decoding performance at target FER of$10^{-7}$, the improved step-GRAND hardware achieves$7 \times \sim 17\times $reduction in W.C. latency,$7\times $reduction in power consumption, and$37 \times \sim 66\times $higher area efficiency in the W.C. scenario. Furthermore, the proposed hardware can achieve an average throughput of up to 10.5 Gb/s and a W.C. throughput of$102\sim 350$Mb/s. Syed Mohsin Abbas, Marwan Jalaleddine, Chi-Ying Tsui, Warren J. Gross |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | RWriC: A Dynamic Writing Scheme for Variation Compensation for RRAM-based In-Memory ComputingabstractRRAM-based compute-in-memory (CIM) suffers from programming variation issues, specifically device-to-device variation (DDV) and cycle-to-cycle variation (CCV), which can have a detrimental impact on inference accuracy. To address these variation issues, we propose RWriC, a dynamic Writing scheme for variation Compensation for RRAM-based CIM. RWriC sequentially programs the weights, implemented by multiple RRAM cells, starting from the high significance cell (HSC) and moving towards the low significance cell (LSC). This approach leverages the knowledge of current cumulative errors and the programming targets (PTs) of other RRAM cells to dynamically adjust the PT of the RRAM currently under programming. By shifting the PT of HSC, RWriC enables the LSC to compensate for the programming errors of the HSC. Moreover, when the variation is substantial, RWriC allows the magnitude of LSC to be scaled up, providing an even wider compensation range. Through the combined application of the shifting and scaling techniques, experimental results show that the inference accuracy for ResNet50 on the CIFAR-10 dataset only drops by 0.9% under 18% device variation. In comparison to the conventional writing scheme, our RWriC approach achieves a 5-11x improvement in variation robustness for ResNet50 and Yolov8 across different tasks. Yucong Huang, Jingyu He, Kwang-Ting Cheng, Chi-Ying Tsui, Terry Tao Ye |
DAC | 4 |
| 2024 | AdaP-CIM: Compute-in-Memory Based Neural Network Accelerator Using Adaptive PositabstractThis study proposes two novel approaches to address memory wall issues in AI accelerator designs for large neural networks. The first approach introduces a new format called adaptive Posit (AdaP) with two exponent encoding schemes that dynamically extend the dynamic range of its representation at run time with minimal hardware overhead. The second approach proposes using compute-in-memory (CIM) with speculative input alignment (SAU) to implement the AdaP multiply-and-accumulate (MAC) computation, significantly reducing the delay, area, and power consumption for the max exponent computation. The proposed approaches outperform state-of-the-art quantization methods and achieve significant energy and area efficiency improvements. Jingyu He, Fengbin Tu, Kwang-Ting Cheng, Chi-Ying Tsui |
DATE | 4 |
| 2024 | ReSCIM: Variation-Resilient High Weight-Loading Bandwidth In-Memory Computation Based on Fine-Grained Hybrid Integration of Multi-Level ReRAM and SRAM CellsabstractSRAM-CIM is a promising approach to implement efficient accelerator architecture as it enables accurate, energy-efficient AI computing, supporting both analog and digital computation. However, it has low area efficiency. On the other hand, Resistive RAM (ReRAM) provides dense on-chip storage, especially with multi-level cells (MLC), but ReRAM-CIM may introduce inaccuracies due to device variation and only supports analog computation. To leverage the strengths of both technologies, a hybrid architecture that combines them at a fine granularity is desirable. Previous hybrid designs incorporate ReRAM resistors into SRAM to improve storage density. However, they face scalability limitations and restricted signal margins for multi-level RRAM readout, leading to degraded computation accuracy. In this work, we propose ReSCIM, a hybrid compute-in-memory (CIM) architecture that seamlessly integrates multi-level ReRAM into SRAM cells at a fine-grained level. By incorporating a compact ReRAM crossbar in each SRAM cell, a dense CIM marco using SRAM-based computation is achieved. We develop an energy-efficient differential sensing scheme that enables parallel weight loading from local ReRAM crossbars to SRAM cells. This scheme allows multi-bit ReRAM data readout using a single SRAM cell and offers resilience to device variations. Furthermore, We designed a ReSCIM accelerator architecture for efficient AI acceleration, fully utilizing the highly scalable storage and exceptional weight-loading bandwidth. We employ a folded weight-mapping approach for MLC ReRAM cells to guarantee accurate classification even under substantial ReRAM device variations. Experimental results show that ReSCIM accelerators based on both analog and digital-based CIM achieve 60% energy savings and 98% latency savings, and 59× higher area efficiency compared to state-of-the-art all-weights-on-chip AI accelerators on AlexNet. Jingyu He, Kunming Shao, Jiakun Zheng, Fengshi Tian, Kwang-Ting Cheng, Chi-Ying Tsui |
ICCAD | 7 |
| 2024 | Time Domain Analysis of Secondary Stage With Series Resonance Driving Rectifier LoadabstractThe secondary stage of a wireless power transfer system with a series-resonant circuit driving a rectifier load is analyzed in the time-domain. Depending on the load current, the inductor current may operate in continuous, boundary or discontinuous conduction modes (CCM, BCM, DCM). Analytic solutions are derived and confirmed by SPICE simulation and measurement results. Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 3 |
| 2024 | Adaptive Digitally-Controlled Active Rectifier-Based Receiver for BioimplantsabstractThis paper presents a wireless power transfer (WPT) receiver with an adaptive digitally-controlled on-off delay-compensated active rectifier for high-current biomedical implants. High efficiency is achieved by using adaptive digital techniques to compensate for turn-on and turn-off delays, reduce reverse current, and eliminate multiple pulsing. The adaptive scheme improves voltage conversion ratio (VCR) and power conversion efficiency (PCE) at different PVT corners. The generation of the optimal delay compensation current is 1.5 times faster than published works. The proposed design is implemented with 0.18 μm CMOS process. The measured maximum VCR is 0.978 for a 2 kΩ load resistance, and the maximum PCE is 92.2% for a 200 Ω load resistance. Sayan Sarkar, Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 4 |
| 2024 | BOLS: A Bionic Sensor-direct On-chip Learning System with Direct-Feedback-Through-Time for Personalized Wearable Health MonitoringabstractPrecise bio-signal classification techniques for edge healthcare have been extensively researched, yet the scalability and efficiency of existing studies remain constrained by challenges in sensing, learning, and processing. Additionally, a deficiency in cross-level integration for the development of comprehensive healthcare systems has been observed. To tackle these issues and facilitate ultra-efficient personalized edge healthcare, this paper introduces the pioneering bionic sensor-direct on-chip learning and inference system with direct-feedback-through-time for user-specific cardiac arrhythmia detection, termed BOLS. This innovative system encompasses a compact sensor-direct feature extractor and a pipelined bionic processor, enabling end-to-end on-chip learning and inference. Employing cross-level co-design, our proposed bionic on-chip learning approach attains exceptional classification performance, boasting an accuracy of 98.6%, which ranks among the highest. The entire system has been implemented using 40nm CMOS process and subsequently verified. Remarkably, the proposed BOLS system consumes a mere 1.18mW for inference and 2.57mW for learning, resulting in an impressive power saving of over ×2000 compared to existing commercial training platforms. Fengshi Tian, Jiakun Zheng, Jingyu He, Jinbo Chen 0002, Chaoming Fang, Jie Yang 0033, Mohamad Sawan, Chi-Ying Tsui, Kwang-Ting Cheng |
ISCAS | 9 |
| 2024 | How Robust is Federated Learning to Communication Error? A Comparison Study Between Uplink and Downlink ChannelsabstractBecause of its privacy-preserving capability, federated learning (FL) has attracted significant attention from both academia and industry. However, when being implemented over wireless networks, it is not clear how much communication error can be tolerated by FL. This paper investigates the robustness of FL to the uplink and downlink communication error. Our theoretical analysis reveals that the robustness depends on two critical parameters, namely the number of clients and the numerical range of model parameters. It is also shown that the uplink communication in FL can tolerate a higher bit error rate (BER) than downlink communication, and this difference is quantified by a proposed formula. The findings and theoretical analyses are further validated by extensive experiments. Linping Qu, Shenghui Song 0001, Chi-Ying Tsui, Yuyi Mao |
WCNC | 3 |
| 2024 | Analysis of Off-Resonant Flexible Tertiary Coils in Biomedical Implants for Back Telemetry DetectionabstractA systematic design technique of flexible tertiary coil in a wireless power transfer (WPT) system for back telemetry detection is presented. The detection (tertiary) coil is co-planar with the primary coil and receives the backscattered load shift keying (LSK) data from the secondary stage. Tertiary coil parameters such as the number of turns, compensation capacitance, and quality factor are determined to maximize the link gain and the end-to-end (E2E) efficiency. The analysis is validated through SPICE simulations and measurements. A novel area-saving inner-tertiary coil is further proposed and critically compared with traditional outer-tertiary coil. The RF electromagnetic radiation or Specific Absorption Rate (SAR) value of the coils was evaluated for possible health hazards using 3D head models and ANSYS/HFSS finite element software. In comparison to a heuristic outer-tertiary coil design, the primary-secondary (operating) distance is improved by 50%, from 12 mm to 18 mm. At an operating distance of 15 mm, the link gain and E2E efficiency are improved by ~207% and ~18%, respectively. Sayan Sarkar, Wing-Hung Ki, Chi-Ying Tsui |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | Energy-Efficient Channel Decoding for Wireless Federated Learning: Convergence Analysis and Adaptive DesignabstractOne of the most critical challenges for deploying distributed learning solutions, such as federated learning (FL), in wireless networks is the limited battery capacity of mobile clients. While it is a common belief that the major energy consumption of mobile clients comes from the uplink data transmission, this paper presents a novel finding, namely channel decoding also contributes significantly to the overall energy consumption of mobile clients in FL. Motivated by this new observation, we propose an energy-efficient adaptive channel decoding scheme that leverages the intrinsic robustness of FL to model errors. In particular, the robustness is exploited to reduce the energy consumption of channel decoders at mobile clients by adaptively adjusting the number of decoding iterations. We theoretically prove that wireless FL with communication errors can converge at the same rate as the case with error-free communication provided the bit error rate (BER) is properly constrained. An adaptive channel decoding scheme is then proposed to improve the energy efficiency of wireless FL systems. Experimental results demonstrate that the proposed method maintains the same learning accuracy while reducing the channel decoding energy consumption by$\sim ~20$% when compared to an existing approach. Linping Qu, Yuyi Mao, Shenghui Song 0001, Chi-Ying Tsui |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | RVComp: Analog Variation Compensation for RRAM-Based in-Memory ComputingabstractResistive Random Access Memory (RRAM) has shown great potential in accelerating memory-intensive computation in neural network applications. However, RRAM-based computing suffers from significant accuracy degradation due to the inevitable device variations. In this paper, we propose RVComp, a fine-grained analog Compensation approach to mitigate the accuracy loss of in-memory computing incurred by the Variations of the RRAM devices. Specifically, weights in the RRAM crossbar are accompanied by dedicated compensation RRAM cells to offset their programming errors with a scaling factor. A programming target shifting mechanism is further designed with the objectives of reducing the hardware overhead and minimizing the compensation errors under large device variations. Based on these two key concepts, we propose double and dynamic compensation schemes and the corresponding support architecture. Since the RRAM cells only account for a small fraction of the overall area of the computing macro due to the dominance of the peripheral circuitry, the overall area overhead of RVComp is low and manageable. Simulation results show RVComp achieves a negligible 1.80% inference accuracy drop for ResNet18 on the CIFAR-10 dataset under 30% device variation with only 7.12% area and 5.02% power overhead and no extra latency. Jingyu He, Yucong Huang, Miguel Angel Lastras-Montaño, Terry Tao Ye, Chi-Ying Tsui, Kwang-Ting Cheng |
ASP-DAC | 5 |
| 2023 | Late Breaking Results: Weight Decay is ALL You Need for Neural Network SparsificationabstractThe heuristic iterative pruning strategy has been widely used for neural network sparsification. However, it is challenging to identify the right connections to remove at each pruning iteration with only a one-shot evaluation of weight magnitude, especially at the early pruning stage. The erroneously removed connections, unfortunately, can hardly be recovered. In this work, we propose a weight decay strategy as a substitute for pruning, which let the "insignificant" weights moderately decay instead of being directly clamped to zero. At the end of the training, the vast majority of redundant weights will naturally become close to zero, making it easier to identify which connections could be removed safely. Experimental results show that the proposed weight decay method can achieve an ultra-high sparsity of 99%. Compared to the current pruning strategy, the model size is further reduced by 34%, improving the compression rate from 69× to 106× at the same accuracy. Xizi Chen, Fengshi Tian, Chi-Ying Tsui |
DAC | 5 |
| 2023 | AutoDCIM: An Automated Digital CIM CompilerabstractDigital Computing-in-Memory (DCIM) is an emerging architecture that integrates digital logic into memory for efficient AI computing. However, current DCIM designs heavily rely on manual efforts. This increases DCIM design time and limits the optimization space, making it challenging to satisfy the user specifications of diverse AI applications. This paper presents AutoDCIM, the first automated DCIM compiler. Au-toDCIM takes the user specifications as inputs and generates a DCIM macro architecture with an optimized layout. AutoDCIM’s template-based generation balances handcrafted cell design and agile macro development. AutoDCIM’s layout exploration loop analyzes diverse DCIM array partitioning schemes to satisfy user specifications. The auto-generated DCIM macros present competitive efficiency results in comparison with state-of-the-art silicon-verified DCIM macros. Jia Chen 0032, Fengbin Tu, Kunming Shao, Fengshi Tian, Xiao Huo, Chi-Ying Tsui, Kwang-Ting Cheng |
DAC | 6 |
| 2023 | Accelerating Large Kernel Convolutions with Nested Winograd TransformationabstractRecent literature has shown that convolutional neural networks (CNNs) with large kernels outperform vision transformers (ViTs) and CNNs with stacked small kernels in many computer vision tasks, such as object detection and image restoration. The Winograd transformation helps reduce the number of repetitive multiplications in convolution and is widely supported by many commercial AI processors. Researchers have proposed accelerating large kernel convolutions by linearly decomposing them into many small kernel convolutions and then sequentially accelerating each small kernel convolution with the Winograd algorithm. This work proposes a nested Winograd algorithm that iteratively decomposes a large kernel convolution into small kernel convolutions and proves it to be more effective than the linear decomposition Winograd transformation algorithm. Experiments show that compared to the linear decomposition Winograd algorithm, the proposed algorithm reduces the total number of multiplications by 1.4 to 10.5 times for computing 4×4 to 31×31 convolutions. Jingbo Jiang, Xizi Chen, Chi-Ying Tsui |
VLSI-SoC | 3 |
| 2023 | Tight Compression: Compressing CNN Through Fine-Grained Pruning and Weight Permutation for Efficient ImplementationabstractThe unstructured sparsity after pruning poses a challenge to the efficient implementation of deep learning models in existing regular architectures like systolic arrays. On the other hand, coarse-grained structured pruning is suitable for implementation in regular architectures but tends to have higher accuracy loss than unstructured pruning when the pruned models are of the same size. In this work, we propose a model compression method based on a novel weight permutation scheme to fully exploit the fine-grained weight sparsity in the hardware design. Through permutation, the optimal arrangement of the weight matrix is obtained, and the sparse weight matrix is further compressed to a small and dense format to make full use of the hardware resources. Two pruning granularities are explored. In addition to the unstructured weight pruning, we also propose a more fine-grained subword-level pruning to further improve the compression performance. Compared to the state-of-the-art works, the matrix compression rate is significantly improved from$5.88\times $to$14.13\times $. As a result, the throughput and energy efficiency are improved by 2.75 and 1.86 times, respectively. Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | FedDQ: Communication-Efficient Federated Learning with Descending QuantizationabstractFederated learning (FL) is an emerging learning paradigm without violating users' privacy. However, large model size and frequent model aggregation cause serious communication bottleneck for FL. To reduce the communication volume, techniques such as model compression and quantization have been proposed. Besides the fixed-bit quantization, existing adaptive quantization schemes use ascending-trend quantization, where the quantization level increases with the training stages. In this paper, we first investigate the impact of quantization on model convergence, and show that the optimal quantization level is directly related to the range of the model updates. Given the model is supposed to converge with the progress of the training, the range of the model updates will gradually shrink, indicating that the quantization level should decrease with the training stages. Based on the theoretical analysis, a descending quantization scheme named FedDQ is proposed. Experimental results show that the proposed descending quantization scheme can save up to 65.2% of the communicated bit volume and up to 68% of the communication rounds, when compared with existing schemes. Linping Qu, Shenghui Song 0001, Chi-Ying Tsui |
GLOBECOM | 3 |
| 2022 | A 16-bit Encrypted On-chip Embedded System for Implantable Medical DevicesabstractA 16-bit on-chip embedded encryption system built upon eFUSE, cipher, hash functions, and EDCs for optical nerve stimulation is presented. The foundry-provided eFUSE IP is modified with a one-shot block to support wireless power transfer operation by mitigating the supply voltage drop problem during sensing to avoid subsequent resetting. Novel logic gate-based auxiliary circuit facilitates different sensing and programming modes in eFUSE. A 128-bit cipher is reduced to 16 bits with cascade structure using the proposed divide-and-conquer algorithm, keeping the cipher strength constant. The developed resource sharing technique reduces the area and the power consumption of the ciper circuits by 2.7 times and 5.1 times, respectively. The whole system with the optical nerve simulator is fabricated with 0.18$\mu$m BCDlite process, and measurement results show the correct encryption operation when powered by wirelessly transferred power. Sayan Sarkar, Jingbo Jiang, Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 4 |
| 2022 | Design Strategy of Off-Resonant Tertiary Coils for Uplink Detection in Biomedical ImplantsabstractA detection (tertiary) coil co-planar with the transmitter coil receives and decodes backscattered data from the receiver using load shift keying (LSK). Tertiary coil parameters such as the number of turns, compensation capacitance and quality factor are studied for maximizing voltage gain and end-to-end efficiency through novel analytical modelling, SPICE simulations and measurement results. A novel “inside-tertiary coil structure” is suggested to save the overall coil area, and two variations are discussed. Voltage gain and efficiency are improved by ~197% and ~15%, respectively, at a transmitter-receiver distance of15 mm. Design strategy improves the operating distance of implant by 50%, from 12 mm to 18 mm. Sayan Sarkar, Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 4 |
| 2022 | Soft-Error-Aware Read-Stability-Enhanced Low-Power 12T SRAM With Multi-Node Upset Recoverability for Aerospace ApplicationsabstractWith the advancement of technology, the size of transistors and the distance between them are reducing rapidly. Therefore, the critical charge of sensitive nodes is reducing, making SRAM cells, used for aerospace applications, more vulnerable to soft-error. If a radiation particle strikes a sensitive node of the standard 6T SRAM cell, the stored data in the cell are flipped, causing a single-event upset (SEU). Therefore, in this paper, a Soft-Error-Aware Read-Stability-Enhanced Low-Power 12T (SARP12T) SRAM cell is proposed to mitigate SEUs. To analyze the relative performance of SARP12T, it is compared with other recently published soft-error-aware SRAM cells, QUCCE12T, QUATRO12T, RHD12T, RHPD12T and RSP14T. All the sensitive nodes of SARP12T can regain their data even if the node values are flipped due to a radiation strike. Furthermore, SARP12T can recover from the effect of single-event multi-node upsets (SEMNUs) induced at its storage node-pair. Along with these advantages, the proposed cell exhibits the highest read stability, as the ‘0’-storing storage node, which is directly accessed by the bitline during read operation, can recover from any upset. Furthermore, SARP12T consumes the least hold power. SARP12T also exhibits higher write ability and shorter write delay than most of the comparison cells. All these improvements in the proposed cell are obtained by exhibiting only a slightly longer read delay and consuming slightly higher read and write energy. Soumitra Pal 0002, Wing-Hung Ki, Chi-Ying Tsui |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | Design of a Single-Stage Wireless Charger with 92.3%-Peak-Efficiency for Portable Devices ApplicationsabstractThis summary presents a fully-integrated wireless charger to achieve high efficiency with low cost and volume. The charger realizes power rectification, voltage regulation and CCCV charging in one power stage only. A bootstrapping technique is also designed for on-chip integration of the bootstrap capacitors. A chip prototype was fabricated in a standard 0.35μm CMOS process with a die area of 8mm2. The charger achieves peak efficiency of 92.3% and 91.4% when the charging currents are 1A and 1.5A, respectively. Lin Cheng 0001, Xinyuan Ge, Wai Chiu Ng, Wing-Hung Ki, Tsz Fai Kwok, Chi-Ying Tsui, Ming Liu 0022 |
ASP-DAC | 7 |
| 2020 | Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based PermutationabstractThe unstructured sparsity after pruning poses a challenge to the efficient implementation of deep learning models in existing regular architectures like systolic arrays. The coarse-grained structured pruning, on the other hand, tends to have higher accuracy loss than unstructured pruning when the pruned models are of the same size. In this work, we propose a compression method based on the unstructured pruning and a novel weight permutation scheme. Through permutation, the sparse weight matrix is further compressed to a small and dense format to make full use of the hardware resources. Compared to the state-of-the-art works, the matrix compression rate is effectively improved from 5.88x to 10.28x. As a result, the throughput and energy efficiency are improved by 2.12 and 1.57 times, respectively. Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui |
DAC | 4 |
| 2020 | A Low-Power Motion Estimation Architecture for HEVC Based on a New Sum of Absolute Difference ComputationabstractHigh-efficiency video coding (HEVC) poses a considerable challenge to hardware implementation due to its complexity. Mobile devices are powered by batteries that are limited in capacity. Therefore, reducing the power consumption arising from the implementation of sophisticated coding tools in HEVC is an especially important issue for mobile devices. In particular, motion estimation (ME) is the major contributor to the power consumption of the encoder and the calculation of the sum of absolute difference (SAD) for ME consumes more than 50% of the total ME power. In this paper, a low-power motion estimation VLSI architecture is proposed based on a novel method of calculating the SAD. By reusing the calculation, the computation complexity and, hence, the power consumption are reduced. A low-power systolic processing elements array and a novel memory hierarchy are developed, which enable real-time processing of 8K resolution video with only half of the power consumption when compared with the state-of-the-art design. Luheng Jia, Chi-Ying Tsui, Oscar C. Au, Kebin Jia |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | CompRRAE: RRAM-based convolutional neural network accelerator with reduced computations through a runtime activation estimationabstractRecently Resistive-RAM (RRAM) crossbar has been used in the design of the accelerator of convolutional neural networks (CNNs) to solve the memory wall issue. However, the intensive multiply-accumulate computations (MACs) executed at the crossbars during the inference phase are still the bottleneck for the further improvement of energy efficiency and throughput. In this work, we explore several methods to reduce the computations for the RRAM-based CNN accelerators. First, the output sparsity resulting from the widely employed Rectified Linear Unit is exploited, and a significant portion of computations are bypassed through an early detection of the negative output activations. Second, an adaptive approximation is proposed to terminate the MAC early when the sum of the partial results of the remaining computations is considered to be within a certain range of the intermediate accumulated result and thus has an insignificant contribution to the inference. In order to determine these redundant computations, a novel runtime estimation on the maximum and minimum values of each output activation is developed and used during the MAC operation. Experimental results show that around 70% of the computations can be reduced during the inference with a negligible accuracy loss smaller than 0.2%. As a result, the energy efficiency and the throughput are improved by over 2.9 and 2.8 times, respectively, compared with the state-of-the-art RRAM-based accelerators. Xizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying Tsui |
ASP-DAC | 4 |
| 2019 | A Two-Staged Adaptive Successive Cancellation List Decoding for Polar CodesabstractPolar codes achieve outstanding error correction performance when using successive cancellation list (SCL) decoding with cyclic redundancy check. A larger list size brings better decoding performance and is essential for practical applications such as 5G communication networks. However, the decoding speed of SCL decreases with increased list size. Adaptive SCL (A-SCL) decoding can greatly enhance the decoding speed, but the decoding latency for each codeword is different so A-SCL is not a good choice for hardware-based applications. In this paper, a hardware-friendly two-staged adaptive SCL (TA-SCL) decoding algorithm is proposed such that a constant input data rate is supported even if the list size for each codeword is different. A mathematical model based on Markov chain is derived to explore the bounds of its decoding performance. Simulation results show that the throughput of TA-SCL is tripled for good channel conditions with negligible performance degradation and hardware overhead. ChenYang Xia, YouZhe Fan, Chi-Ying Tsui |
ISCAS | 3 |
| 2019 | SubMac: Exploiting the subword-based computation in RRAM-based CNN accelerator for energy saving and speedup
Xizi Chen, Jingbo Jiang, Jingyang Zhu, Chi-Ying Tsui |
Integr. | 4 |
| 2019 | Microshift: An Efficient Image Compression Algorithm for HardwareabstractIn this paper, we propose a lossy image compression algorithm called microshift. We employ an algorithm-hardware co-design methodology, yielding a hardware-friendly compression approach with low power consumption. In our method, the image is first micro-shifted, and then the sub-quantized values are further compressed. Two methods, FAST and MRF models, are proposed to recover the bitdepth by exploiting the spatial correlation of natural images. Both methods can decompress images progressively. On an average, our compression algorithm can compress images to 1.25-bits per pixel with a resulting quality that outperforms the state-of-the-art on-chip compression algorithms in both peak signal-to-noise ratio and structual similarity. Then, we propose a hardware architecture and implement the algorithm on an FPGA. The results on the ASIC design further validate the low-hardware complexity and high-power efficiency, showing that our method is promising, particularly for low-power wireless vision sensor networks. Bo Zhang 0025, Pedro V. Sander, Chi-Ying Tsui, Amine Bermak |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | A New Rate-Complexity-Distortion Model for Fast Motion Estimation Algorithm in HEVCabstractIn the high efficiency video coding (HEVC) standard, motion estimation (ME) adopts a quadtree coding structure and a larger search range to improve the coding performance. These advanced coding tools, however, dramatically increase the computational complexity. To accelerate ME, fast methods have been proposed that reduce ME complexity at the expense of rate-distortion (R-D) performance. However, none of these methods can claim their tradeoff to be optimal. In this paper, we propose an optimal fast motion estimation (FME) algorithm based on an analytical model of rate-complexity-distortion (R-C-D). We extend the traditional R-D model by introducing the ME complexity into it, which enables us to explicitly express the R-D performance under different complexity budgets. Based on the R-C-D model, the proposed FME finds the R-C-D optimized search ranges for some representative prediction units (PUs), which are then extended or refined dynamically according to motion characteristics for neighboring PUs. The proposed FME enables an R-D performance close to that of full search. When compared with the default FME method in the reference software of HEVC, the proposed fast algorithm can reduce the complexity by over 80% while improving the R-D performance. Furthermore, our proposed FME algorithm is hardware friendly, as regular data flow enables high data reuse efficiency. Luheng Jia, Chi-Ying Tsui, Oscar C. Au, Kebin Jia |
IEEE Trans. Multim. | 2 |
| 2018 | A high-throughput and energy-efficient RRAM-based convolutional neural network using data encoding and dynamic quantizationabstractTo solve the scaling, memory wall and high power density issues, recently RRAM-based accelerators, which show a better energy and area efficiency compared with the CMOS-based counterparts, have been proposed for convolutional neural networks. However, the RRAM-based architectures still face several design challenges, including the high energy and timing overhead at the analog/digital (A/D) conversion and interfacing circuits. To address these issues, we propose several novel optimization schemes in this work. First an encoding scheme for the synaptic weights and the input feature maps is proposed to reduce the energy of the in-situ computation and the bit-resolution of the A/D conversion. Then the resolution of the A/D conversion is further optimized for a lower energy consumption. Moreover, a dynamic quantization scheme for the multiply-accumulate operations (MACs) is proposed to improve the throughput and the energy efficiency by reducing the number of partial products. Experimental results show that the throughput, the energy efficiency and the area efficiency are improved by 2 to 4 times when compared with the state-of-the-art RRAM-based accelerators. Xizi Chen, Jingbo Jiang, Jingyang Zhu, Chi-Ying Tsui |
ASP-DAC | 4 |
| 2018 | SparseNN: An energy-efficient neural network accelerator exploiting input and output sparsityabstractContemporary Deep Neural Network (DNN) contains millions of synaptic connections with tens to hundreds of layers. The large computational complexity poses a challenge to the hardware design. In this work, we leverage the intrinsic activation sparsity of DNN to substantially reduce the execution cycles and the energy consumption. An end-to-end training algorithm is proposed to develop a lightweight (less than 5% overhead) run-time predictor for the output activation sparsity on the fly. Furthermore, an energy-efficient hardware architecture, SparseNN, is proposed to exploit both the input and output sparsity. SparseNN is a scalable architecture with distributed memories and processing elements connected through a dedicated on-chip network. Compared with the state-of-the-art accelerators which only exploit the input sparsity, SparseNN can achieve a 10%-70% improvement in throughput and a power reduction of around 50%. Jingyang Zhu, Jingbo Jiang, Xizi Chen, Chi-Ying Tsui |
DATE | 4 |
| 2018 | An Indoor Solar Energy Harvester with Ultra-Low-Power Reconfigurable Power-On-Reset-Styled Voltage DetectorabstractAn ultra-lower-power reconfigurable voltage detector for indoor solar energy harvester is presented. The voltage detector monitors the solar cell voltage and sends out a flag signal if the solar cell voltage surpasses the triggering threshold of the detector. Instead of using a traditional dynamic comparator, this design is based on a power-on-reset (POR) circuit. A POR circuit has ultra-low quiescent loss and a fixed triggering threshold, which is determined by its topology and process. Our improvement is to use a feedback loop that allows the triggering threshold to be reconfigurable. The average quiescent loss of this POR voltage detector circuit is 2.774nW. Process and temperature variations can also be compensated by the feedback loop. The energy harvesting system is designed with a 0.18 p,m CMOS process. Equipped with the proposed voltage detector, the whole system achieves a 93.21% peak efficiency at 200μW input power. Xiaodong Meng, Xing Li 0004, Chi-Ying Tsui, Wing-Hung Ki |
ISCAS | 4 |
| 2018 | On Path Memory in List Successive Cancellation Decoder of Polar CodesabstractPolar code is a breakthrough in coding theory. Using list successive cancellation decoding with large list size L, polar codes can achieve excellent error correction performance. The L partial decoded vectors are stored in the path memory and updated according to the results of list management. In the state-of-the-art designs, the memories are implemented with registers and a large crossbar is used for copying the partial decoded vectors from one block of memory to another during the update. The architectures are quite area-costly when the code length and list size are large. To solve this problem, we propose two optimization schemes for the path memory in this work. First, a folded path memory architecture is presented to reduce the area cost. Second, we show a scheme that the path memory can be totally removed from the architecture. Experimental results show that these schemes effectively reduce the area of path memory. ChenYang Xia, YouZhe Fan, Chi-Ying Tsui |
ISCAS | 4 |
| 2018 | Room-Temperature Dual-mode CMOS Gas-FET Sensor for Diabetes DetectionabstractA CMOS gas-sensitive field-effect transistor (Gas-FET) is proposed for noninvasive diabetes detection. The Gas-FET was fabricated in the GlobalFoundries 0.18μm 1P6M process with a lateral control gate and a floating gate to set operating point and gas sensing sensitivity, respectively. ZnO nanorods were used as the sensing material and deposited on top of the chip using a hydrothermal process at 80°C. Room-temperature acetone sensing down to sub-ppm level is demonstrated to enable noninvasive diagnosis of diabetes in exhaled breath. A dual-mode integrated readout circuit is also proposed to improve the sensor gas discrimination ability through the acquisition of 2-dimensional information by every single Gas-FET sensor. Farid Boussaïd, Amine Bermak, Chi-Ying Tsui |
ISCAS | 4 |
| 2017 | A wireless power receiver with a 3-level reconfigurable resonant regulating rectifier for mobile-charging applicationsabstractA wireless power receiver using a 3-level reconfigurable resonant regulating (R3) rectifier is presented in this summary. The receiver improves power conversion efficiency and reduces die area and off-chip components by achieving power conversion and voltage regulation in one stage, using only 4 on-chip power switches and 1 off-chip capacitor. The receiver regulates the output voltage at 5 V and delivers a maximum power of 6 W. It was fabricated in a standard 0.35 μm CMOS process with a die area of 4.77 mm2, and the measured peak efficiency reaches 92.2%. Lin Cheng 0001, Wing-Hung Ki, Chi-Ying Tsui |
ASP-DAC | 3 |
| 2017 | BHNN: A memory-efficient accelerator for compressing deep neural networks with blocked hashing techniquesabstractIn this paper, we propose a novel algorithm for compressing neural networks to reduce the memory requirements by using blocked hashing techniques. By adding blocked constraints on top of the conventional hashing technique, the test error rate is maintained while the spatial locality for the computations is preserved. Using this scheme, the synaptic connections are compressed by at least an order (10×) compared with the plain neural network with virtually no prediction accuracy loss. Compared with other compression techniques, the proposed algorithm achieves the best performance in the heavy compression regions. The blocked hashing techniques are also hardware friendly, of which the memory hierarchy of the hardware architecture can be efficiently implemented. To demonstrate the hardware efficiency, we implement the hardware architecture of the deep neural networks using the proposed blocked hashing techniques on a Xilinx Virtex-7 FPGA board. With a hardware parallelism of 32, the accelerator achieves a speed-up of 22× over the CPU, and 3~5× over the GPU in the inference phase. Jingyang Zhu, Zhiliang Qian, Chi-Ying Tsui |
ASP-DAC | 3 |
| 2017 | An implementation of list successive cancellation decoder with large list size for polar codesabstractPolar codes are the first class of forward error correction (FEC) codes with a provably capacity-achieving capability. Using list successive cancellation decoding (LSCD) with a large list size, the error correction performance of polar codes exceeds other well-known FEC codes. However, the hardware complexity of LSCD rapidly increases with the list size, which incurs high usage of the resources on the field programmable gate array (FPGA) and significantly impedes the practical deployment of polar codes. To alleviate the high complexity, in this paper, two low-complexity decoding schemes and the corresponding architectures for LSCD targeting FPGA implementation are proposed. The architecture is implemented in an Altera Stratix V FPGA. Measurement results show that, even with a list size of 32, the architecture is able to decode a codeword of 4096-bit polar code within 150 μs, achieving a throughput of 27Mbps. ChenYang Xia, YouZhe Fan, Chi-Ying Tsui, ChongYang Zeng, Bin Li 0013 |
FPL | 4 |
| 2017 | Concatenated LDPC-polar codes decoding through belief propagationabstractOwing to their capacity-achieving performance and low encoding and decoding complexity, polar codes have drawn much research interests recently. Successive cancellation decoding (SCD) and belief propagation decoding (BPD) are two common approaches for decoding polar codes. SCD is sequential in nature while BPD can run in parallel. Thus BPD is more attractive for low latency applications. However BPD has some performance degradation at higher SNR when compared with SCD. Concatenating LDPC with Polar codes is one popular approach to enhance the performance of BPD, where a short LDPC code is used as an outer code and Polar code is used as an inner code. In this work we propose a new way to construct concatenated LDPC-Polar code, which not only outperforms conventional BPD and existing concatenated LDPC-Polar code but also shows a performance improvement of 0.5 dB at higher SNR regime when compared with SCD. Syed Mohsin Abbas, YouZhe Fan, Chi-Ying Tsui |
ISCAS | 4 |
| 2017 | Dual transduction Gas sensor based on a surface acoustic wave resonatorabstractThis paper presents a novel dual transduction gas sensor providing both resistance and mass modalities for single sensor gas identification. The proposed sensor relies on a configurable dual mode frequency/resistance readout circuit, which enables the use of a single conventional surface acoustic wave (SAW) device. Unlike prior works which all rely on custom-made gas sensors, the proposed sensor is based on off-the-shelf SAW devices, making it low cost and easy to implement. Reported results validate the functionality of the proposed dual transduction gas sensor. The introduction of control switches for the dual mode readout is shown to only deteriorate the phase noise performance of the SAW oscillator by 4 dBc/Hz at 10 MHz offset and not affect the low offset part. This demonstrates that the mass sensing resolution of the SAW device is not reduced while including the resistive sensing feature. Amine Bermak, Chi-Ying Tsui, Farid Boussaïd |
ISCAS | 3 |
| 2017 | A low-offset dynamic comparator with area-efficient and low-power offset cancellationabstractA low-offset two-stage dynamic comparator has been proposed for parallel multi-channel processing. Low offset is achieved from two aspects: 1st-stage offset cancellation and 2nd-stage offset suppression. A fully dynamic offset cancellation scheme based on current auto-zeroing is adopted to effectively cancel out the 1st-stage offset. It features small area overhead and low energy consumption. For the 2nd-stage offset suppression, a high gain is designed for the 1st-stage dynamic amplifier by optimizing the overdrive voltage of input transistors. To maintain low offset performance across a wide range of input commonmode voltages, the overdrive voltage of the input pair is required to stay low. Therefore, a tail current source is employed for the 1st stage to ensure constant common-mode discharging current. As a result, the overdrive voltage can be stably kept low under various operation conditions. The proposed comparator has been designed in a standard CMOS 0.18 μm process. It operates under a supply voltage of 1.2 V at 10 MHz. Simulation results have verified the low-offset property of the comparator. The input-referred offset (1 σ) is reduced from 19.25 mV to 1.296 mV after cancellation and it remains constant with the input commonmode voltage changing from 0 V to 0.8 V. The offset is further reduced to 771 μV when the 2nd-stage input pair are enlarged by 4 times. At the same time, the energy consumption is increased from 147 fJ/Conv to 168 fJ/Conv. Xiaopeng Zhong, Amine Bermak, Chi-Ying Tsui |
VLSI-SoC | 3 |
| 2017 | High-Throughput and Energy-Efficient Belief Propagation Polar Code DecoderabstractOwing to their capacity-achieving performance and low encoding and decoding complexity, polar codes have received significant attention recently. Successive cancellation decoding (SCD) and belief propagation decoding (BPD) are two popular approaches for decoding polar codes. SCD, despite having less computational complexity when compared with BPD, suffers from long latency due to the serial nature of the SC algorithm. BPD, on the other hand, is parallel in nature and is more attractive for low-latency applications. However, due to the iterative nature of BPD, the required latency and energy dissipation increase linearly with the number of iterations. In this paper, we propose a novel scheme based on subfactor-graph freezing to reduce the average number of computations as well as the average number of iterations required by BPD, which directly translates into lower latency and energy dissipation. Simulation results show that the proposed scheme has no performance degradation and achieves significant reduction in computation complexity over the existing methods. Moreover, the hardware architecture for the proposed scheme is developed and compared with the state-of-the-art BPD implementations for (1024, 512) polar codes. A decoding throughput of 13.9 Gb/s is achieved along with a 60%-73% improvement in energy reduction and two times increase in hardware efficiency when compared with the existing BPD implementations. Syed Mohsin Abbas, YouZhe Fan, Chi-Ying Tsui |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | LRADNN: High-throughput and energy-efficient Deep Neural Network accelerator using Low Rank ApproximationabstractIn this work, we propose an energy-efficient hardware accelerator for Deep Neural Network (DNN) using Low Rank Approximation (LRADNN). Using this scheme, inactive neurons in each layer of the DNN are dynamically identified and the corresponding computations are then bypassed. Accordingly, both the memory accesses and the arithmetic operations associated with these inactive neurons can be saved. Therefore, compared to the architectures using the direct feed-forward algorithm, LRADNN can achieve a higher throughput as well as a lower energy consumption with negligible prediction accuracy loss (within 0.1%). We implement and synthesize the proposed accelerator using TSMC 65nm technology. From the experimental results, a 31% to 53% energy reduction together with a 22% to 43% throughput increase can be achieved. Jingyang Zhu, Zhiliang Qian, Chi-Ying Tsui |
ASP-DAC | 3 |
| 2016 | Low-Complexity List Successive-Cancellation Decoding of Polar Codes Using List PruningabstractThe performance of List Successive-Cancellation Decoding (LSCD) of Polar Codes with large list size have exceeded that of Turbo codes and Low-Density Parity-Check codes. However, large list size results in huge computation complexity and this limits the applicability of LSCD in high-throughput and power- sensitive applications. In this work, a low complexity design for LSCD with large list size based on list pruning is proposed. In particular, the property of the relative path metric (RPM) of each list candidate with respect to that of the most-likely candidate is investigated. It is found that the correct candidate has a low possibility of having a large value of RPM and based on this property, a list pruning method and the corresponding low-complexity LSCDs are proposed. From the simulation results, as compared to the conventional LSCD, the proposed LSCDs have negligible performance loss while the computation complexity is reduced by more than 80%. In addition, the proposed design is hardware-friendly and easily adaptable to the existing LSCDs hardware architectures. YouZhe Fan, ChenYang Xia, Chi-Ying Tsui, Kai Chen 0013, Bin Li 0013 |
GLOBECOM | 4 |
| 2016 | Hardware decoders for polar codes: An overviewabstractPolar codes are an exciting new class of error correcting codes that achieve the symmetric capacity of memoryless channels. Many decoding algorithms were developed and implemented, addressing various application requirements: from error-correction performance rivaling that of LDPC codes to very high throughput or low-complexity decoders. In this work, we review the state of the art in polar decoders implementing the successive-cancellation, belief propagation, and list decoding algorithms, illustrating their advantages. Pascal Giard, Gabi Sarkis, Alexios Balatsoukas-Stimming, YouZhe Fan, Chi-Ying Tsui, Andreas Peter Burg, Claude Thibeault, Warren J. Gross |
ISCAS | 5 |
| 2016 | An indoor solar energy harvesting system using dual mode SIDO converter with fully digital time-based MPPTabstractA dual-mode Single-Input-Dual-Output converter with fully digital time-based Maximum-Power-Point-Tracking is presented to efficiently harvest indoor solar energy. The converter uses only three power switches to minimize the chip area and reduce the switching power loss, and dual-mode to accommodate the power imbalance between the Photovoltaic (PV) cell supply and the load requirement. Under sufficient illumination condition such that the PV cell provides more power than needed by the load, the system will operate in light-load mode, supplying sufficient energy to the load and storing the surplus energy to battery. If the environment is under-illuminated, the system will switch to heavy-load mode, and the stored energy is extracted from the battery to supply load together with the PV cell. The power converter is designed with 0.18μm CMOS process. From the simulation results, a peak power conversion efficiency of 91.23% is achieved. Xiaodong Meng, Xing Li 0004, Chi-Ying Tsui, Wing-Hung Ki |
ISCAS | 3 |
| 2016 | A WLAN 2.4-GHz RF energy harvesting system with reconfigurable rectifier for wireless sensor networkabstractIn RF energy harvesting for wireless sensor network, due to the variation of the available RF power for individual node sensor, the conventional Dickson rectifier could not achieve the optimal harvesting efficiency in a wide available power range. In this work, a novel reconfigurable rectifier with adjustable conversion ratio is proposed without changing matching network. Based on this, a WLAN (2.4 GHz) RF energy harvesting system is proposed with the maximum power point tracking (MPPT). The entire system is implemented in an 180nm CMOS process. Post-layout HSPICE simulation results show that with MPPT, the reconfigurable rectifier achieves a high energy harvesting efficiency for a wide available power range and is configured with 3.3x response time reduction when compared with the existing reconfigurable rectifier to cater for the frequent changing RF power. Zizhen Zeng, Xing Li 0004, Amine Bermak, Chi-Ying Tsui, Wing-Hung Ki |
ISCAS | 4 |
| 2016 | A low-power chopper bandpass amplifier for biopotential sensorsabstractA low-power low-noise chopper operational amplifier for biosensor applications is proposed. It employs a simple ripple suppression method using a band-pass amplifier. The input referred noise is only 43nV/rtHz. Fabricated in a 130nm standard CMOS process, it occupies a chip area of 0.28mm2and consumes 33ßA from a 1.2V supply. Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 3 |
| 2016 | Low-latency approximate matrix inversion for high-throughput linear pre-coders in massive MIMOabstractThis work presents a high-throughput and low-latency matrix inversion design for a linear pre-coder for massive MIMO systems. Because of the large number of Base Station (BS) antennas as well as the multiple User Terminals (UTs) served in a massive MIMO system, the channel matrix dimensions become larger. Inversions of such large matrices using direct inversion methods, such as used in linear pre-coders like Zero Forcing (ZF), would entail prohibitive complexity. For avoiding such complexity, Neumann series based approximate inversion has been suggested for linear pre-coders in massive MIMO systems. However the performance, complexity and convergence speed of the Neumann series approach depends very much on the initial approximation of the inverse used as a starting point. In this work, we present a novel initial approximation for the Neumann series which facilitates the parallel computation of the inverse and hence results in lower latency for inversion as well as better accuracy when compared to the previous approaches. A VLSI architecture of the proposed method is implemented for the inversion of a 16 × 16 matrix, in TSMC 65nm technology. A throughput of 0.54M to 15M matrix inversion per sec is achieved at a clock frequency of 460MHz with a 117K gate count. Syed Mohsin Abbas, Chi-Ying Tsui |
VLSI-SoC | 2 |
| 2016 | BiLink: A high performance NoC router architecture using bi-directional link with double data rate
Jingyang Zhu, Zhiliang Qian, Chi-Ying Tsui |
Integr. | 3 |
| 2016 | A Low-Latency List Successive-Cancellation Decoding Implementation for Polar CodesabstractDue to their provably capacity-achieving performance, polar codes have attracted a lot of research interest recently. For a good error-correcting performance, list successive-cancellation decoding (LSCD) with large list size is used to decode polar codes. However, as the complexity and delay of the list management operation rapidly increase with the list size, the overall latency of LSCD becomes large and limits the applicability of polar codes in high-throughput and latency-sensitive applications. Therefore, in this work, the low-latency implementation for LSCD with large list size is studied. Specifically, at the system level, a selective expansion method is proposed such that some of the reliable bits are not expanded to reduce the computation and latency. At the algorithmic level, a double thresholding scheme is proposed as a fast approximate-sorting method for the list management operation to reduce the LSCD latency for large list size. A VLSI architecture of the LSCD implementing the selective expansion and double thresholding scheme is then developed, and implemented using a UMC 90 nm CMOS technology. Experimental results show that, even for a large list size of 16, the proposed LSCD achieves a decoding throughput of 460 Mbps at a clock frequency of 658 MHz. YouZhe Fan, ChenYang Xia, Chi-Ying Tsui, Hui Shen 0006, Bin Li 0013 |
IEEE J. Sel. Areas Commun. | 4 |
| 2016 | A Support Vector Regression (SVR)-Based Latency Model for Network-on-Chip (NoC) ArchitecturesabstractIn this paper, we propose SVR-NoC, a network-on-chip (NoC) latency model using support vector regression (SVR). More specifically, based on the application communication information and the NoC routing algorithm, the channel and source queue waiting times are first estimated using an analytical queuing model with two equivalent queues. To improve the prediction accuracy, the queuing theory-based delay estimations are included as features in the learning process. We then propose a learning framework that relies on SVR to collect training data and predict the traffic flow latency. The proposed learning methods can be used to analyze various traffic scenarios for the target NoC platform. Experimental results on both synthetic and real-application traffic demonstrate on average less than 12% prediction error in network saturation load, as well as more than 100× speedup compared to cycle-accurate simulations can be achieved. Zhiliang Qian, Da-Cheng Juan, Paul Bogdan, Chi-Ying Tsui, Diana Marculescu, Radu Marculescu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2016 | Performance Evaluation of NoC-Based Multicore Systems: From Traffic Analysis to NoC Latency ModelingabstractIn this survey, we review several approaches for predicting performance of Network-on-Chip (NoC)-based multicore systems, starting from the traffic models to the complex NoC models for latency evaluation. We first review typical traffic models to represent the application workloads in NoC. Specifically, we review Markovian and non-Markovian (e.g., self-similar or long-range memory-dependent) traffic models and discuss their applications on multicore platform design. Then, we review the analytical techniques to predict NoC performance under given input traffic. We investigate analytical models for average as well as maximum delay evaluation. We also review the developments and design challenges of NoC simulators. One interesting research direction in NoC performance evaluation consists of combining simulation and analytical models in order to exploit their advantages together. Toward this end, we discuss several newly proposed approaches that use hardware-based or learning-based techniques. Finally, we summarize several open problems and our perspective to address these challenges. Zhiliang Qian, Paul Bogdan, Chi-Ying Tsui, Radu Marculescu |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2015 | Low-latency list decoding of polar codes with double thresholdingabstractFor polar codes with short-to-medium code length, list successive cancellation decoding is used to achieve a good error-correcting performance. However, list pruning in the current list decoding is based on the sorting strategy and its timing complexity is high. This results in a long decoding latency for large list size. In this work, aiming at a low-latency list decoding implementation, a double thresholding algorithm is proposed for a fast list pruning. As a result, with a negligible performance degradation, the list pruning delay is greatly reduced. Based on the double thresholding, a low-latency list decoding architecture is proposed and implemented using a UMC 90nm CMOS technology. Synthesis results show that, even for a large list size of 16, the proposed low-latency architecture achieves a decoding throughput of 220 Mbps at a frequency of 641 MHz. YouZhe Fan, ChenYang Xia, Chi-Ying Tsui, Hui Shen 0006, Bin Li 0013 |
ICASSP | 4 |
| 2015 | A fast variable block size motion estimation algorithm with refined search range for a two-layer data reuse schemeabstractMotion estimation (ME) serves as a key tool in a variety of video coding standards. With the increasing need for higher resolution video format, the limited memory bandwidth becomes a bottleneck for ME implementation. The huge data loading from external memory to the on-chip memory and the frequent data fetching from the on-chip memory to the ME engine are two major problems. To reduce both off-chip and on-chip memory bandwidth, we propose a two-layer data reuse scheme. On the macroblock (MB) layer, an advanced Level C data reuse scheme is presented. It employs two cooperating on-chip caches which load data in a novel local-snake scanning manner. On the block layer, we propose a fast variable block size motion estimation with a refined search window (RSW-VBSME). A new approach for hardware implementation of VBSME is then employed based on the fast algorithm. Instead of obtain the SADs of all the modes at the same time, the ME of different block sizes are performed separately. This enables higher data reusability within an MB. The two-layer data reuse scheme archives a more than 90% reduction of off-chip memory bandwidth with a slight increase of on-chip memory size. Moreover, the on-chip memory bandwidth is also greatly reduced compared with other reuse methods with different VBSME implementations. Luheng Jia, Chi-Ying Tsui, Oscar C. Au, Amin Zheng |
ISCAS | 2 |
| 2015 | UHF energy harvesting system using reconfigurable rectifier for wireless sensor networkabstractIn RF energy harvesting, due to the variation of the available RF power, the conventional Dickson rectifier could not achieve the optimal harvesting efficiency in a wide available power range. To tackle this issue, in this work, a reconfigurable rectifier with adjustable conversion ratio is proposed without affecting the matching with the antenna under different configurations. Based on this, a UHF (900 MHz) energy harvesting system with maximum power point tracking (MPPT) is proposed. The system is implemented using TSMC 65 nm CMOS process. Post-layout HSPICE simulation is performed for verification. With the MPPT, the reconfigurable rectifier is always configured to achieve the maximum efficiency. Xing Li 0004, Chi-Ying Tsui, Wing-Hung Ki |
ISCAS | 2 |
| 2015 | Message from the technical program chairsabstractOn behalf of the Technical Program Committee of the 23rd IFIP/IEEE International Conference on Very Large Scale Integration (VLSI-SoC 2015), we welcome you to Daejeon, Korea and thank you for joining us at this important event. Youngsoo Shin, Chi-Ying Tsui |
VLSI-SoC | 2 |
| 2015 | FSNoC: A Flit-Level Speedup Scheme for Network on-Chips Using Self-Reconfigurable Bidirectional ChannelsabstractIn this paper, we explore optimizing the bandwidth utilization of the network-on-chips (NoCs). We propose a flit-level speedup scheme to improve the NoC performance using self-reconfigurable bidirectional channels. For the NoC intrarouter bandwidth, in addition to allowing flits from different packets to use the idle internal bandwidth of the crossbar, our proposed flit-level speedup scheme also allows flits within the same packet to be transmitted simultaneously. For interrouter channels, a distributed channel configuration scheme is developed to dynamically change the link directions. In this way, the effective bandwidth between two routers can change adaptively depending on the run time network traffic. We present the implementation of the proposed flit-level speedup NoC on a 2-D mesh. An input buffer architecture, which supports reading and writing two flits from the same virtual channel at the same time, is proposed. The switch allocator is also designed to support flit-level parallel arbitration. Extensive simulations on both the synthetic traffic and real applications show performance improvement in throughput and latency over the existing architectures using bidirectional channels. Zhiliang Qian, Syed Mohsin Abbas, Chi-Ying Tsui |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | A comprehensive and accurate latency model for Network-on-Chip performance analysisabstractIn this work, we propose a new, accurate, and comprehensive analytical model for Network-on-Chip (NoC) performance analysis. Given the application communication graph, the NoC architecture, and the routing algorithm, the proposed framework analyzes the links dependency and then determines the ordering of queuing analysis for performance modeling. The channel waiting times in the links are estimated using a generalized G/G/1/K queuing model, which can tackle bursty traffic and dependent arrival times with general service time distributions. The proposed model is general and can be used to analyze various traffic scenarios for NoC platforms with arbitrary buffer and packet lengths. Experimental results on both synthetic and real applications demonstrate the accuracy and scalability of the newly proposed model. Zhiliang Qian, Da-Cheng Juan, Paul Bogdan, Chi-Ying Tsui, Diana Marculescu, Radu Marculescu |
ASP-DAC | 4 |
| 2014 | Disease Diagnosis-on-a-Chip: Large Scale Networks-on-Chip based Multicore Platform for Protein Folding AnalysisabstractProtein folding is critical for many biological processes. In this work, we propose an NoC-based multi-core platform for protein folding computation. We first identify the speedup bottleneck for applying conventional genetic algorithm on a mesh-based multi-core platform. Then, we address this computation- and communication- intensive problem while taking into account both hardware and software aspects. Specifically, we group the processing cores into islands and propose an NoC-based multicore architecture for intra- and inter-island communication. The high scalability of the proposed platform allows us to integrate from 100 to 1200 cores for the folding computation. We then propose a genetic migration algorithm to take advantage of the massive parallel platform. Our simulation results show that the proposed platform offers near-linear speedup as the number of cores increases. We also report the hardware cost in area and power based on a 100-core FPGA prototype. Yuankun Xue, Zhiliang Qian, Paul Bogdan, Fan Ye 0001, Chi-Ying Tsui |
DAC | 5 |
| 2014 | Low-latency MAP demapper architecture for coded modulation with iterative decodingabstractBit-interleaved coded modulation with iterative decoding has been widely adopted in modern wireless communication systems because of its spectral efficiency and low detection complexity. Because of the iterative decoding structure, the overall decoding latency depends on the latency of both the demapper and the channel decoder. In this work, a parallel demapper architecture is proposed for a low latency implementation. Two look-ahead techniques are proposed to further reduce the latency of the parallel architecture. Techniques exploiting the symmetry property of the labeling and the common terms of the look-ahead pre-calculation are also presented to reduce the computation complexity overhead of the proposed architecture. Implementation results show that the latency is reduced by 40% when comparing with traditional sequential demapper architecture. YouZhe Fan, Chi-Ying Tsui |
ISCAS | 2 |
| 2014 | An adaptive wireless powering and data telemetry system for optic nerve stimulationabstractTo treat retinal degenerative diseases, a transcorneal electrical stimulation-based system is proposed which consists of an in vivo eye implant and an in vitro control device. The eye implant is wirelessly powered and controlled by the in vitro device to generate the required bi-phase current pattern for the transcorneal stimulation with an amplitude range of 5μA~320μA, a frequency range of 10Hz~160Hz and a duty ratio range of 2.5%~20%. Only one pair of coils is used for both the power link and the bi-directional data link. Except the secondary coil, the eye implant is fully integrated on chip and fabricated using UMC 0.13μm CMOS process with a size of 1.5mm by 1.5mm. The secondary coil is tiny and fabricated on a PCB with an outer diameter of only 4.4mm. After packaging with biocompatible silicone, the whole implant has a dimension of 6mm diameter with a thickness of less than 1mm. The whole device can be put on the sclera and beneath the eye's conjunctiva. Measurement results verify the system's functionality and demonstrate the performance. Xing Li 0004, Yan Lu 0002, Chi-Ying Tsui, Wing-Hung Ki |
ISCAS | 3 |
| 2014 | A fast intermode decision algorithm based on analysis of inter prediction residualabstractRate-distortion-optimized (RDO) intermode decision is one of the most effective tools that greatly improves the coding performance of modern coding standards, for example, H.264/AVC and HEVC. However, RDO intermode decision also leads to extremely intense computation. To reduce the complexity, a fast intermode decision algorithm is presented in this paper. Mathematical analysis of inter prediction residual is performed which explicitly shows the impact of edge information and motion characteristics to the prediction accuracy. Moreover, It is shown that with fixed quantization step that minimizing of R-D costs over different partition types is equivalent to minimizing the variance of transformed residual which can be expressed by motion vector and edge gradient components. In consequence, the complex calculation of rate and distortion is replaced by simple pre-analysis of video content. The repetition of motion estimation (ME) and entropy coding process are avoided. Inspired by the theoretical analysis, a fast inter mode decision algorithm is proposed. Experimental results show that the fast method achieves considerable complexity reduction with negligible coding performance degradation. Luheng Jia, Oscar C. Au, Chi-Ying Tsui, Wei Dai 0002, Pengfei Wan 0001 |
MMSP | 3 |
| 2014 | An efficient Network-on-Chip (NoC) based multicore platform for hierarchical parallel genetic algorithmsabstractIn this work, we propose a new Network-on-Chip (NoC) architecture for implementing the hierarchical parallel genetic algorithm (HPGA) on a multi-core System-on-Chip (SoC) platform. We first derive the speedup metric of an NoC architecture which directly maps the HPGA onto NoC in order to identify the main sources of performance bottlenecks. Specifically, it is observed that the speedup is mostly affected by the fixed bandwidth that a master processor can use and the low utilization of slave processor cores. Motivated by the theoretical analysis, we propose a new architecture with two multiplexing schemes, namely dynamic injection bandwidth multiplexing (DIBM) and time-division based island multiplexing (TDIM), to improve the speedup and reduce the hardware requirements. Moreover, a task-aware adaptive routing algorithm is designed for the proposed architecture, which can take advantage of the proposed multiplexing schemes to further reduce the hardware overhead. We demonstrate the benefits of our approach using the problem of protein folding prediction, which is a process of importance in biology. Our experimental results show that the proposed NoC architecture achieves up to 240X speedup compared to a single island design. The hardware cost is also reduced by 50% compared to a direct NoC-based HPGA implementation. Yuankun Xue, Zhiliang Qian, Guopeng Wei, Paul Bogdan, Chi-Ying Tsui, Radu Marculescu |
NOCS | 5 |
| 2014 | A Novel Single-Inductor Dual-Input Dual-Output DC-DC Converter With PWM Control for Solar Energy Harvesting SystemabstractA novel single-inductor dual-input dual-output dc-dc converter with pulse width modulation control is proposed for a solar energy harvesting system. The first input of the converter is from photovoltaic (PV) cells and the second input is a rechargeable battery. Apart from the conventional role of providing a regulated output voltage to power the loading circuits, the converter also clamps the PV cells' voltage to the maximum power point value to maximize efficiency. When the PV cells harvest more power than the load, the surplus energy is used to charge the rechargeable battery. When the PV cells cannot harvest sufficient power, the converter schedules the PV cells and the battery to power the load together. A test chip was fabricated using a 0.35-μm CMOS process and measured to verify the operation of the proposed dc-dc converter and to demonstrate the power transfer efficiency of the solar power management system. Xing Li 0004, Chi-Ying Tsui, Wing-Hung Ki |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | SVR-NoC: a performance analysis tool for network-on-chips using learning-based support vector regression modelabstractIn this work, we propose SVR-NoC, a learning-based support vector regression (SVR) model for evaluating Network-on-Chip (NoC) latency performance. Different from the state-of-the-art NoC analytical model, which uses classical queuing theory to directly compute the average channel waiting time, the proposed SVR-NoC model performs NoC latency analysis based on learning the typical training data. More specifically, we develop a systematic machine-learning framework that uses the kernel-based support vector regression method to predict the channel average waiting time and the traffic flow latency. Experimental results show that SVR-NoC can predict the average packet latency accurately while achieving about 120X speed-up over simulation-based evaluation methods. Zhiliang Qian, Da-Cheng Juan, Paul Bogdan, Chi-Ying Tsui, Diana Marculescu, Radu Marculescu |
DATE | 4 |
| 2013 | Performance evaluation of multicore systems: from traffic analysis to latency predictions (embedded tutorial)abstractAs technology scaling down allows multiple processing components to be integrated on a single chip, the modern computing systems led to the advent of Multiprocessor System-on-Chip (MPSoC) and Chip Multiprocessor (CMP) design. Network-on-Chips (NoCs) have been proposed as a promising solution to tackle the complex on-chip communication problems on these multicore platforms. In order to optimize the NoC-based multicore system design, it is essential to evaluate the NoC performance with respect to numerous configurations in a large design space. Taking the traffic characteristics into account and using an appropriate latency model become crucially important to provide an accurate and fast evaluation. In this tutorial, we survey the current progresses in these aspects. We first review the NoC workload modeling and traffic analysis techniques. Then, we discuss the mathematical formalisms of evaluating the performance under a given traffic model, for both the average and worst-case latency predictions. Finally, the advantages of combining the analytical and simulation-based techniques are discussed and new attempts for bridging these two approaches are reviewed. Zhiliang Qian, Paul Bogdan, Chi-Ying Tsui, Radu Marculescu |
ICCAD | 3 |
| 2012 | A flit-level speedup scheme for network-on-chips using self-reconfigurable bi-directional channelsabstractIn this work, we propose a flit-level speedup scheme to enhance the network-on-chip(NoC) performance utilizing bidirectional channels. In addition to the traditional efforts on allowing flits of different packets using the idling internal and external bandwidth of the bi-directional channel, our proposed flit-level speedup scheme also allows flits within the same packet to be transmitted simultaneously on the bi-directional channel. For inter-router transmission, a novel distributed channel configuration protocol is developed to dynamically control the link directions. For the intra-router transmission, an input buffer architecture which supports reading and writing two flits from the same virtual channel at the same time is proposed. The switch allocator is also designed to support flit-level parallel arbitration. Simulation results on both synthetic traffic and real benchmarks show performance improvement in throughput and latency over the existing architectures using bi-directional channels. Zhiliang Qian, Ying Fei Teh, Chi-Ying Tsui |
DATE | 3 |
| 2012 | Low-Complexity Rotated QAM Demapper for the Iterative Receiver Targeting DVB-T2 StandardabstractGray mapping and signal space diversity (SSD) are adopted in DVB-T2 to achieve better performance and system robustness. However, the traditional maximum a posteriori demapping for Gray mapped SSD signal is complicated for higher order modulation and it is not practical to be used in the iterative receiver structure. In this work, simplified demappers are proposed by approximating the 2-dimensional detection with 1-dimensional detection and compensating the loss due to the correlation between the I and Q components. Simulation results show that the proposed simplified demappers can approach the optimal demapper performance with a much lower complexity. YouZhe Fan, Chi-Ying Tsui |
VTC Fall | 2 |
| 2011 | A thermal-aware application specific routing algorithm for Network-on-Chip designabstractIn this work, we propose an application specific routing algorithm to reduce the hot-spot temperature for Network-on-chip (NoC). Using the traffic information of applications, we develop a routing scheme which can achieve a higher adaptivity than the generic ones and at the same time distribute the traffic more uniformly. A set of deadlock-free admissible paths for all the communications is first obtained. To reduce the hot-spot temperature, we find the optimal distribution ratio of the communication traffic among the set of candidate paths. The problem of finding this optimal distribution ratio is formulated as a linear programming (LP) problem and is solved offline. A router microarchitecture which supports our ratio-based selection policy is also proposed. From the simulation results, the peak energy reduction considering the energy consumption of both the processors and routers can be as high as 16.6% for synthetic traffic and real benchmarks. Zhiliang Qian, Chi-Ying Tsui |
ASP-DAC | 2 |
| 2011 | Efficient iterative receiver for LDPC coded wireless IPTV systemabstractMulti-level superposition coded modulation (SCM) is a scalable technique for wireless video broadcast/ multicast, in which iterative turbo structures provide receivers with multi-resolution demodulations subject to a high complexity. Forward error correction using low-density parity-check (LDPC) code is helpful for better received video quality but further increasing the receiver complexity. In this paper, a method is proposed to reduce the receiver complexity by using a sequential structure with a faster convergence for demodulation. In addition, the iterative demodulator and the LDPC decoder are jointly designed as a multi-loop iterative structure to reduce the decoding complexity. Experimental results show that up to 67% decoding complexity is reduced and better video quality is achieved at receivers under low signal-to-noise ratios. YouZhe Fan, James She, Chi-Ying Tsui |
ICIP | 3 |
| 2011 | A low-complexity image compression algorithm for Address-Event Representation (AER) PWM image sensorsabstractIn Pulse-Width Modulation (PWM) image sensors the incident light intensity is represented by the timing of pulses. Exceptionally high dynamic range (DR) and improved signal-to- noise-ratio (SNR) have been demonstrated for this class of image sensors. Unfortunately, their spatial resolution is limited by the need of an in-pixel memory to record the timing information. The AER protocol is an attractive method for removing this overhead, since pixel trigger events can be sent as address vectors, and in- pixel data memories are no longer required. Regrettably, the need to send address vectors can place an increased burden on the communication channel and will limit the array resolution, frame-rate, and image quality. In this paper, we present a low- complexity AER Block Compression (AERBC) algorithm which exploits the statistically ordered nature of AER pixel arrays. The address vector overhead can be dramatically reduced under this scheme. Only 0.0625 comparisons and 0.125 subtractions are performed for each pixel, and on average 30.82 dB PSNR can be achieved at 1.0 bit-per-pixel code rate. A general strategy is also developed here to optimize AERBC parameters so a balance between performance and hardware resources can be reached. Denis Guangyin Chen, Amine Bermak, Chi-Ying Tsui |
ISCAS | 3 |
| 2011 | Design and analysis of on-chip charge pumps for micro-power energy harvesting applicationsabstractCharge balance law based on conservation of charge is stated and employed to analyze charge pumps. For micro-power on-chip implementations, both the positive-plate and the negative-plate parasitic capacitors have to be considered. A first iteration approximation analysis is proposed to analyze charge pumps with parasitic capacitors. Using a 0.35µm CMOS process, 8X linear, Fibonacci and exponential charge pumps are designed and their performances are compared and confirmed by simulations. Wing-Hung Ki, Yan Lu 0002, Feng Su, Chi-Ying Tsui |
VLSI-SoC | 4 |
| 2011 | A fault-tolerant network-on-chip design using dynamic reconfiguration of partial-faulty routing resourcesabstractIn this work, we propose a fault-tolerant framework for Network on Chips (NoC) to achieve maximum performance under fault. A fine-grained fault model is first introduced. Different from the traditional link or node NoC fault models which assume the faulty resource to be totally unfunctional, we distinguish the faulty components and handle them according to their fault classes. By doing so, we can avoid unnecessary partitioning of the network and hence achieve a higher connectivity under high fault rate. Two new dynamic reconfiguration schemes at the router level, namely Dynamic Buffer Swapping (DBS) and Dynamic MUX Swapping(DMS), are proposed to deal with the buffer and cross-bar faults accordingly, which are the main sources of failure in the router. In these schemes, the healthy resources in the router are maximally utilized to mitigate the faults. Experimental results show that we can achieve higher packet acceptance rate and lower latency compared with state-of-the-art fault-tolerant routing schemes. Zhiliang Qian, Ying Fei Teh, Chi-Ying Tsui |
VLSI-SoC | 3 |
| 2011 | A fault-tolerant NoC using combined link sharing and partial fault link utilization schemeabstractWith reducing feature size of transistors and increasing number of cores on a single chip, system-on-chips (SoCs) are becoming more vulnerable to faults due to the physical level defects of VLSI fabrication. Fault tolerance and reliability have become two significant challenges for SoC designers. In this work, we propose a novel and efficient scheme to handle the faulty links of a network-on-chip (NoC) by adaptively combining two schemes, namely the link sharing scheme and partial fault link utilization scheme. With our approach, the system is able to optimize the usage of the remaining bandwidth of the links under different fault conditions. Experimental results show a significant improvement in average latency and maximum delay by using the proposed combined scheme with only 4.62% of hardware overhead cost. Our proposed scheme offers a way to increase the effective yield of large and complex NoC systems by enabling the usage of faulty chip with little compromise in the latency performance. Ying Fei Teh, Zhiliang Qian, Chi-Ying Tsui |
VLSI-SoC | 3 |
| 2011 | Vibration Energy Scavenging System With Maximum Power Tracking for Micropower ApplicationsabstractIn this work, we present a vibration-based energy scavenging system based on piezoelectric conversion for micropower applications. A novel maximum power point (MPP) tracking scheme is proposed to harvest the maximum power from the vibration system. A time-multiplexing mechanism is employed to perform energy harvesting and MPP tracking alternately. In the MPP tracking mode, a voltage reference that represents the optimal output voltage at the MPP is generated. A control unit then uses this reference to track the system operation around the MPP. The proposed system is capable of self-starting up without the help of an energy buffer. As a result, it is suitable for battery-less applications or when the energy buffer is completely drained. This tracking scheme has very small power overhead and is simple to implement in VLSI. Hence, it is especially applicable for micropower systems. The entire design was fabricated in a 0.35-μ m CMOS process. Experimental results verified the proposed MPP tracking scheme and demonstrated the system operation. Measurement results show that the power harvesting efficiency of the electrical circuitry is higher than 90%. Chao Lu 0005, Chi-Ying Tsui, Wing-Hung Ki |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Maximizing the harvested energy for micro-power applications through efficient MPPT and PMU designabstractEnergy harvesting is becoming more and more popular for micro-power applications where the environmental energy is used to power up the systems. In order to prolong the device lifetime and guarantee the system operation, the harvested power from the energy transducer to supply the system load should be maximized. This paper reviews different techniques and solutions to maximize the harvested power. Different environmental energy sources and the characteristics of the corresponding energy transducers are discussed. Algorithms to detect and track the maximum power point (MPP) of the energy transducer are summarized. Different power management unit (PMU) designs to execute MPP tracking (MPPT) algorithms are presented. Chi-Ying Tsui, Wing-Hung Ki |
ASP-DAC | 2 |
| 2010 | System level power optimizations for EPC RFID tags to improve sensitivity using load power shaping and operation schedulingabstractConventional design methodologies for EPC RFID tags focused on minimizing the power consumption of each building block to obtain the maximum achievable communication distance. In this work, we analyze different operation states of the RFID tag and their corresponding power consumption characteristics to identify the limiting factor that dictates the communication distance. Based on the analysis, we propose two approaches to extend the communication distance: (1) shaping the load power to track the unbalanced input power to the tag; (2) scheduling the load to average out the large burst power consumption. Simulation results show that the sensitivity can be improved by 1.8dB. Yunxiao Ling, Chi-Ying Tsui, Wing-Hung Ki |
ISCAS | 3 |
| 2010 | A single inductor DIDO DC-DC converter for solar energy harvesting applications using band-band controlabstractA single inductor dual input dual output (SIDIDO) DC-DC converter is proposed for solar energy harvesting applications. The converter supports hybrid power supplies from both the photovoltaic (PV) cells and the rechargeable battery. The proposed DC-DC converter provides a regulated output voltage to power the load as well as clamps the input voltage from the PV cells to the maximum power point (MPP) value so as to extract the maximum power from the PV cells. At the same time, one of the dual outputs is used to charge up a rechargeable battery when the PV cells output power is larger than the load. A simple bandband control is proposed for controlling the power flow between different sources and destination under different power operation modes. Extensive simulations were carried out to verify the operation and demonstrate the power transfer efficiency of the proposed DC-DC converter. Chi-Ying Tsui, Wing-Hung Ki |
VLSI-SoC | 2 |
| 2010 | An Energy Efficient Layered Decoding Architecture for LDPC DecoderabstractLow-density parity-check (LDPC) decoder requires large amount of memory access which leads to high energy consumption. To reduce the energy consumption of the LDPC decoder, memory-bypassing scheme has been proposed for the layered decoding architecture which reduces the amount of access to the memory storing the soft posterior reliability values. In this work, we present a scheme that achieves the optimal reduction of memory access for the memory bypassing scheme. The amount of achievable memory bypassing depends on the decoding order of the layers. We formulate the problem of finding the optimal decoding order and propose algorithm to obtain the optimal solution. We also present the corresponding architecture which combines some of memory components and results in reduction of memory area. The proposed decoder was implemented in TSMC 0.18 μm CMOS process. Experimental results show that for a LDPC decoder targeting IEEE 802.11 n specification, the amount of memory access values can be reduced by 12.9-19.3% compared with the state-of-the-art design. At the same time, 95.6%-100% hardware utilization rate is achieved. Chi-Ying Tsui |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Joint Routing and Sleep Scheduling for Lifetime Maximization of Wireless Sensor NetworksabstractThe rapid proliferation of wireless sensor networks has stimulated enormous research efforts that aim to maximize the lifetime of battery-powered sensor nodes and, by extension, the overall network lifetime. Most work in this field can be divided into two equally important threads, namely (i) energy-efficient routing that balances traffic load across the network according to energy-related metrics and (ii) sleep scheduling that reduces energy cost due to idle listening by providing periodic sleep cycles for sensor nodes. To date, these two threads are pursued separately in the literature, leading to designs that optimize one component assuming the other is pre-determined. Such designs give rise to practical difficulty in determining the appropriate routing and sleep scheduling schemes in the real deployment of sensor networks, as neither component can be optimized without pre-fixing the other one. This paper endeavors to address the lack of a joint routing-and-sleep-scheduling scheme in the literature by incorporating the design of the two components into one optimization framework. Notably, joint routing-and-sleep-scheduling by itself is a non-convex optimization problem, which is difficult to solve. We tackle the problem by transforming it into an equivalent Signomial Program (SP) through relaxing the flow conservation constraints. The SP problem is then solved by an iterative Geometric Programming (IGP) method, yielding an near optimal routing-and-sleep-scheduling scheme that maximizes network lifetime. To the best of our knowledge, this is the first attempt to obtain the optimal joint routing-and-sleep-scheduling strategy for wireless sensor networks. The near optimal solution provided by this work opens up new possibilities for designing practical and heuristic schemes targeting the same problem, for now the performance of any new heuristics can be easily evaluated by using the proposed near optimal scheme as a benchmark. Chi-Ying Tsui, Ying-Jun Angela Zhang |
IEEE Trans. Wirel. Commun. | 2 |
| 2009 | Low energy level converter design for sub-Vth logicsabstractA low energy consumption level converter (LC) is presented for logic voltage conversion from sub-Vthvoltage to nominal high voltage. By employing the multi-stage architecture and implementing a unique circuit inside each stage, the proposed LC can reduce its energy consumption by almost 3 orders and at the same time ensure the robustness of its function. The LC was fabricated and measured to verify its operation and the performance improvement. Chi-Ying Tsui |
ASP-DAC | 2 |
| 2009 | An inductor-less MPPT design for light energy harvesting systemsabstractAn inductor-less maximum power point tracker was designed for light energy harvesting systems. We target at systems under different lighting environments and sometimes the solar cell voltage may be low. A charge pump is used to convert the voltage to a higher value. At the same time, the control circuit tunes the charge pump switching frequency to track the system maximum output power point. The design was fabricated and measured to verify the system operation. Chi-Ying Tsui, Wing-Hung Ki |
ASP-DAC | 2 |
| 2009 | A Low Energy Two-step Successive Approximation Algorithm for ADC DesignabstractThis paper presents a new method for switching the capacitors in the DAC capacitor array of a successive approximation register (SAR) ADC. By separating the decoding of the most significant bits and the least significant bits, and using two different capacitor arrays with unequal size to determine their values, respectively, the average switching energy of the capacitor arrays is dramatically reduced compared to the traditional switching methods. Calibration registers are used to reduce the error of the most significant bits conversion due to the usage of a smaller capacitor array. Experiments were carried out on a 10-bit SAR-ADC designed using TSMC 0.18mum CMOS process. HSPICE simulations show that significant reduction in energy consumption is achieved using the proposed design. Ricky Yiu-kee Choi, Chi-Ying Tsui |
ISCAS | 2 |
| 2009 | Improving the Hardware Utilization Efficiency of Partially Parallel LDPC Decoder with Scheduling and Sub-matrix DecompositionabstractPartially parallel LDPC decoder is commonly used for practical applications due to its good tradeoff between the hardware cost and the throughput. In the partially parallel LDPC decoding architecture, two kinds of processor units are implemented: check node unit (CNU) and variable node unit (VNU). Because of the dependency between two kinds of processor units, the low hardware utilization efficiency (HUE) is one of the design issues for the partially parallel decoding architecture. In order to achieve the optimal hardware utilization efficiency, it is important to determine the order of the rows and columns in the LDPC parity check matrix processed by the processor units. In this paper, we model the scheduling problem as an optimization problem and use simulated annealing to find good solutions for the scheduling. In order to further increase the HUE of the partially parallel decoding architecture, sub-matrix decomposition scheme is proposed. By applying these two schemes, the HUE of some partially parallel decoding implementations can achieve 100%. Chi-Ying Tsui |
ISCAS | 2 |
| 2009 | A single inductor dual input dual output DC-DC converter with hybrid supplies for solar energy harvesting applicationsabstractA single inductor dual input dual output (SIDIDO) DC-DC converter is proposed for solar energy harvesting applications. The converter supports hybrid power supplies from both the photovoltaic (PV) cells and the rechargeable battery. Apart from the conventional role of providing a regulated output voltage to power the load, the proposed DC-DC converter also clamps the input voltage from the PV cells to the maximum power point (MPP) value so as to extract the maximum power from the PV cells. At the same time, one of the dual outputs is used to charge up a rechargeable battery. When the PV cells output power is insufficient for the load, the converter will schedule the PV cells and the battery to power the load together. When the PV cells output power is larger than the load, the surplus PV cells energy is used to charge the rechargeable battery through the converter's second output. Extensive simulations were carried out to verify the operation and demonstrate the power transfer efficiency of the proposed DC-DC converter. Copyright 2009 ACM. Chi-Ying Tsui, Wing-Hung Ki |
ISLPED | 2 |
| 2009 | The Design of a Micro Power Management System for Applications Using Photovoltaic Cells With the Maximum Output Power ControlabstractAn inductor-less on-chip micro power management system for light energy harvesting applications is presented. We target at wide variety of applications that operate at different lighting environments ranging from strong sunlight to dim indoor lighting where the output voltage from the photovoltaic cells is low. A step-up charge pump is used to directly operate the circuit or to charge a rechargeable battery. The power management system operation is discussed and the control strategy for transferring the maximum output power from the power system is presented. Low power circuit design is proposed for the implementation of the system maximum output power control. The system was implemented using a 0.35-mum CMOS process. The chip was fabricated and measurements were conducted for different lighting conditions to demonstrate the system operation and verify the control strategy. Chi-Ying Tsui, Wing-Hung Ki |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Integrated single-inductor dual-input dual-output boost converter for energy harvesting applicationsabstractAn integrated single-inductor dual-input dual-output (SI DIDO) boost converter for energy harvesting applications was designed in a 0.35μm CMOS process. It provides two regulated output voltages for the load and the charge storage device, and two sources, the energy harvesting source and the charge storage device, are multiplexed to serve as the input. The implementation has several special features. (1) The input power MUX is driven by an internal charge pump for a larger gate drive to save area. (2) The power stage is implemented with an active diode core to eliminate gate drive circuitry. (3) A 1:7 timeslot scheduling with a fixed peak inductor current is adopted to deliver energy to the two outputs with a large difference in load currents. The proposed converter could operate at 1V with up to 85% efficiency at 200mW. Ngok-Man Sze, Feng Su, Yat-Hei Lam, Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 5 |
| 2008 | An energy-adaptive MPPT power management unit for micro-power vibration energy harvestingabstractA batteryless power management unit (PMU) that manages harvested low-level vibration energy from a piezoelectric device for a wireless sensor node is presented. An energy-adaptive maximum power point tracking (EA-MPPT) scheme is proposed that allows the PMU to activate different operation modes according to the available power level. The harvested energy is processed by an ac-dc voltage doubler followed by on-chip charge pumps with variable up/down conversion ratios for higher efficiency. Interleaving technique is employed for the high-power output to reduce both current and voltage ripples. The PMU is designed using a 0.35μm CMOS process, and simulation results are presented to demonstrate its functions. Feng Su, Yat-Hei Lam, Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 5 |
| 2008 | A low power layered decoding architecture for LDPC decoder implementation for IEEE 802.11n LDPC codesabstractThis paper presents a low power LDPC decoder design based on reducing the amount of memory access. By utilizing the column overlapping of the LDPC parity check matrix, the amount of access for the memory storing the posterior values is minimized. In addition, a thresholding decoding scheme is proposed which reduces the memory access by trading off the error correcting performance. The decoder was implemented in TSMC 0.18μm CMOS process. Experimental results show that for a LDPC decoder targeting for IEEE 802.11n, the power consumption of the memory and the decoder can be reduced by 72% and 24%, respectively. Chi-Ying Tsui |
ISLPED | 2 |
| 2008 | Minimizing the dynamic and sub-threshold leakage power consumption using least leakage vector-assisted technology mapping
Chi-Ying Tsui, Robert Yi-Ching Au, Ricky Yiu-kee Choi |
Integr. | 1 |
| 2007 | Energy-Aware Synthesis of Networks-on-Chip Implemented with Voltage IslandsabstractVoltage islands provide a very good opportunity for minimizing the energy consumption of core-based Networks-on-Chip (NoC) design by utilizing a unique supply voltage for the cores on each island. This paper addresses various complex design issues for NoC implementation with voltage islands. A novel design framework based on genetic algorithm is proposed to optimize both the computation and communication energy with the creation of voltage islands concurrently for the NoC using multiple supply voltages. The algorithm automatically performs tile mapping, routing path allocation, link speed assignment, voltage island partitioning and voltage assignment simultaneously. Experiments using both real-life and artificial benchmarks were performed and results show that, by using the proposed scheme, significant energy reduction is obtained. Lap-Fai Leung, Chi-Ying Tsui |
DAC | 2 |
| 2007 | A Batteryless Vibration-based Energy Harvesting System for Ultra Low Power Ubiquitous ApplicationsabstractThis paper proposed a vibration driven energy harvesting platform based on piezoelectric material. A new maximum power point tracking (MPPT) method for this platform is also presented. This platform harvests ambient vibration energy as its power source, and is capable of self-starting, and self-powered operation without the need of a battery. The proposed MPPT methodology consumes very little power, and is especially suitable for the environments, where ambient harvested power is very low. System modeling, analysis, and implementation are developed. Simulation results show that the power harvesting efficiency with the proposed MPPT scheme is higher than 90% and the power utilization efficiency of the overall platform is higher than 50%. Chao Lu 0005, Chi-Ying Tsui, Wing-Hung Ki |
ISCAS | 2 |
| 2007 | Design and Implementation of a Low-power Baseband-system for RFID TagabstractThis article describes a low power design approach for a UHF passive RFID tag baseband system. It proposes a new RFID tag baseband architecture which is compatible with the EPC C1G2 UHF RFID protocol. Advanced low power design approaches are adopted, including separating driving clocks, applying an improved Tausworthe sequence generator, moving window PIE decoding algorithm, idle scheme and parallel operating scheme. The tag supports three commands, which are read, write and query. It consists of a 136 bits one-time programmable memory, rectifier, charge pump, clock divider, analog frontend and baseband system. SimuLink co-verification approach is applied for system functional test. The chip was designed and fabricated successfully by using 0.18 mum 6 layers CMOS technology Sau-Wing Man, Edward S. Zhang, Hin-Tat Chan, Vincent K. N. Lau, Chi-Ying Tsui, Howard C. Luong |
ISCAS | 5 |
| 2007 | An Inductor-less Micro Solar Power Management System Design for Energy Harvesting ApplicationsabstractA micro solar power management system is presented for energy harvesting applications. An inductor-less solution is proposed which facilitates the system on-chip integration. The authors target at the applications working in different lighting environments ranging from strong sunlight to dim lighting where the output voltage from the photovoltaic (PV) cells is low. A charge pump is used to step up the PV voltage to charge a battery or directly operate the circuit. The solar power system behavior is analyzed and the control strategy is derived in order to achieve maximum power transferred from the PV cells to battery or computation circuit. Simulations for the power management system were carried out to verify the proposed control strategy. Chi-Ying Tsui, Wing-Hung Ki |
ISCAS | 2 |
| 2007 | Vibration energy scavenging and management for ultra low power applicationsabstractIn this work, the design of a mechanical vibration energy scavenging and management system is presented for ultra low power applications. A new maximum power point tracking (MPPT) scheme is proposed for piezoelectric conversion. This scheme consumes very little power and is especially suitable for ultra low power energy harvesting applications. This design is capable of self-starting and self-powered, thus eliminates external battery integration and significantly reduces the system volume. System modeling, analysis, and VLSI implementation were developed. Various simulations were carried out and the simulation results show that the proposed MPPT scheme can achieve an energy harvesting efficiency higher than 90%. Chao Lu 0005, Chi-Ying Tsui, Wing-Hung Ki |
ISLPED | 2 |
| 2007 | A micro power management system and maximum output power control for solar energy harvesting applicationsabstractA micro power management system is presented for solar energy harvesting applications. An inductor-less solution is proposed which facilitates on-chip integration of the system. We target at applications working in different lighting environments ranging from strong sunlight to dim indoor lighting where the output voltage from the photovoltaic (PV) cells is low. A charge pump is used to step the PV voltage up to charge a battery or directly operate the circuit. The power management system behavior is theoretically analyzed and the control strategy is derived to transfer maximum power from the PV cells to the battery or the circuit. Circuit design and low power techniques for the maximum output power control are proposed for the system. The system was implemented using 0.35μm CMOS process. With the measured PV cells output characteristics, HSPICE simulations for the power management system were carried out to verify the control strategy and to demonstrate the system operation. Chi-Ying Tsui, Wing-Hung Ki |
ISLPED | 2 |
| 2007 | Low-Power Limited-Search Parallel State Viterbi Decoder Implementation Based on Scarce State TransitionabstractIn this paper, a low-power Viterbi decoder design based on scarce state transition (SST) is presented. A low complexity algorithm based on a limited search algorithm, which reduces the average number of the add-compare-select computation of the Viterbi algorithm, is proposed and seamlessly integrated with the SST-based decoder. The new decoding scheme has low overhead and facilitates low-power implementation for high throughput applications. We also propose an uneven-partitioned memory architecture for the trace-back survivor memory unit to reduce the overall memory access power. The new Viterbi decoder is designed and implemented in TSMC 0.18-mum CMOS process. Simulation results show that power consumption is reduced by up to 80% for high throughput wireless systems such as Multiband-OFDM Ultra-wideband applications. Chi-Ying Tsui |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Integrated direct output current control switching converter using symmetrically-matched self-biased current sensorsabstractA noninverting flyback converter using an integrated symmetrically-matched self-biased current sensor was fabricated in a 0.35/spl mu/m CMOS process. It operates in pseudo-continuous conduction mode and employs a direct output current control scheme to achieve excellent line transient response. The converter switches at 1MHz with an input of 1.2V to 2V to give an output of 1.5V and delivers 250mA. Yat-Hei Lam, Suet-Chui Koon, Wing-Hung Ki, Chi-Ying Tsui |
ASP-DAC | 4 |
| 2006 | Adaptively-biased capacitor-less CMOS low dropout regulator with direct current feedbackabstractA capacitor-less low dropout regulator (LDR) with direct current feedback is proposed. A symmetrically-matched voltage mirror in sensing the load current is employed, and gives excellent line and load regulations. The dynamic biasing results in an LDR with pole-tracking that extends the bandwidth of the loop gain at high load currents. The LDR was fabricated in a 0.35/spl mu/m CMOS process with an active area of 0.11mm/sup 2/, and measurement results corroborated well with both analysis and simulation. Yat-Hei Lam, Wing-Hung Ki, Chi-Ying Tsui |
ASP-DAC | 3 |
| 2006 | Ultra-low voltage power management circuit and computation methodology for energy harvesting applicationsabstractA power management and computation methodology is proposed for ultra-low power energy harvesting applications. An integrated exponential charge pump that accepts an input voltage of around 150mV and provides an unregulated output voltage of more than 1.5V serves as the power supply. To cater with the fluctuated energy source and unregulated power supply, a supply side charge-based computation methodology is proposed, of which the computation activity tracks with the fluctuation of the available energy. The idea is demonstrated in a test chip fabricated using a 0.35 mum technology Chi-Ying Tsui, Wing-Hung Ki, Feng Su |
ASP-DAC | 1 |
| 2006 | Optimal link scheduling on improving best-effort and guaranteed services performance in network-on-chip systemsabstractWith the advent of the multiple IP-core based design using Network on Chip (NoC), it is possible to run multiple applications concurrently. For applications with hard deadline, guaranteed services (GS) are required to satisfy the deadline requirement. GS typically under-utilizes the network resources. To increase the resources utilization efficiency, GS applications are always complement with the best-effort services (BE). To allow more resource available for BE, the resource reservation for GS applications, which depends heavily on the scheduling of the computation and communication, needs to be optimized. In this paper we propose a new approach based on optimal link scheduling to judiciously schedule the packets on each of the links such that the maximum latency of the GS application is minimized with minimum network resources utilization. To further increase the performance, we propose a novel router architecture using a shared-buffer implementation scheme. The approach is formulated using integer linear programming (ILP). We applied our algorithm on real applications and experimental results show that significant improvement on the overall execution time and link utilization can be achieved. Lap-Fai Leung, Chi-Ying Tsui |
DAC | 2 |
| 2006 | High performance single clock cycle CMOS comparatorabstractIn this paper, a novel comparison algorithm is introduced, which uses a parallel MSB checking method instead of the traditional priority-encoding based comparison algorithm. By doing so, fast dynamic NOR gates are used instead of high-fan-in NAND gates and this results in significant improvement in performance over the traditional design. A test chip has been built in AMS 0.35/spl mu/m technology and from both post-layout simulation and test chip measurement results, it is shown that the proposed design is 22% faster than the existing fastest single-cycle comparator based on priority-encoder respectively. Hing Mo Lam, Chi-Ying Tsui |
ISCAS | 2 |
| 2006 | Energy-aware optimal workload allocation among the battery-powered devices to maximize the co-operation life timeabstractTraditional low power design minimizes the total power consumption. However for systems that want to maximize co-operation life time among the battery-powered devices, the total power minimization strategy may not be the best. The workload allocation among these devices has a dramatic impact on the co-operation life time. In this paper, we address this problem by formulating it as a linear programming (LP) problem. A simplified energy-aware algorithm is proposed to provide the optimal solution according to the energy capacity ratio of the devices. This algorithm is simple enough to be embedded into real-time system and we illustrate this by an example using a configurable transmitter-receiver pair. Experimental results show that the proposed algorithm significantly outperforms the other workload allocation strategies by 400.66% in the randomly-generated systems and 23.20% in a point-to-point (P2P) communication system, respectively Chi-Ying Tsui |
ISCAS | 2 |
| 2006 | A charge based computation system and control strategy for energy harvesting applicationsabstractA novel charge based computation system is proposed for ultra low power applications based on energy harvesting. An integrated exponential charge pump that accepts an input voltage of around 150mV and provides an unregulated output voltage of more than 1.5V serves as the power supply. To track with the unregulated voltage supply, a supply side charge-based computation methodology is proposed. Charge based control strategy is further derived to efficiently use the harvested energy and to increase the number of computations during a certain time interval. The suggested system and control strategy are demonstrated in a test chip fabricated using a 0.35/spl mu/m technology. Chi-Ying Tsui, Wing-Hung Ki |
ISCAS | 2 |
| 2006 | High efficiency cross-coupled doubler with no reversion lossabstractReversion loss is analyzed in details. A 2/spl times/ cross-coupled charge pump in a 0.35/spl mu/ CMOS process is designed using a gate control strategy that eliminates reversion loss. It switches at 1MHz, delivers a current of 30mA, and reaches an efficiency of 95% for an input voltage of 1V to 1.6V. Feng Su, Wing-Hung Ki, Chi-Ying Tsui |
ISCAS | 3 |
| 2006 | A low power Viterbi decoder implementation using scarce state transition and path pruning scheme for high throughput wireless applicationsabstractThis paper presents a low power Viterbi decoder design based on Scarce State Transition (SST). We propose an approach which seamlessly integrates the path pruning techniques with the SST decoding to reduce the average add-compare-select (ACS) computation. The scheme has very low overhead and is practical for implementation. We also propose an uneven-partitioned memory architecture for the survivor memory unit to reduce the memory access power during the trace back operation. The proposed decoder is implemented in SMIC 0.18?m CMOS process. Simulation results show that significant power consumption reduction can be achieved for high throughput wireless systems such as MB-OFDM Ultra-wide-band applications. Chi-Ying Tsui |
ISLPED | 2 |
| 2006 | Low Complexity SST Viterbi DecoderabstractReducing the complexity and power consumption of the Viterbi decoder is one of the important design goals for high throughput wireless systems. Recently, a low complexity decoding algorithm was proposed to reduce the average number of ACS (Add Compare Select) operation of the Viterbi algorithm (VA) using the information of syndrome. Unfortunately, it has two limitations: the large computation overhead and the large memory requirement, which prevent it from the practical VLSI implementation. In this work, we propose an approach which facilitates the implementation of this reduced complexity algorithm by building it on the Scarce State Transition (SST) Viterbi decoding scheme. The proposed scheme achieves the similar complexity reduction while facilitates the low power implementation. Simulation results show that over 80% computation reduction can be achieved while the bit-error-rate (BER) performance of the VA is maintained. Jin Jie, Chi-Ying Tsui |
VTC Fall | 2 |
| 2005 | Exploiting Dynamic Workload Variation in Low Energy Preemptive Task SchedulingabstractA novel energy reduction strategy to maximally exploit the dynamic workload variation is proposed for the offline voltage scheduling of preemptive systems. The idea is to construct a fully preemptive schedule that leads to minimum energy consumption when the tasks take on approximately the average execution cycles yet still guarantees no deadline violation during the worst-case scenario. End-time for each sub-instance of the tasks obtained from the schedule is used for the on-line dynamic voltage scaling (DVS) of the tasks. For the tasks that normally require a small number of cycles but occasionally a large number of cycles to complete, such a schedule provides more opportunities for slack utilization and hence results in larger energy saving. The concept is realized by formulating the problem as a nonlinear programming (NLP) optimization problem. Experimental results show that, by using the proposed scheme, the total energy consumption at runtime is reduced by as much as 60% for randomly generated task sets when compared with the static scheduling approach only using worst case workload. Lap-Fai Leung, Chi-Ying Tsui, Xiaobo Sharon Hu |
DATE | 2 |
| 2004 | A dual-band switching digital controller for a buck converter
Martin Yeung-Kei Chui, Wing-Hung Ki, Chi-Ying Tsui |
ASP-DAC | 3 |
| 2004 | Minimizing energy consumption of multiple-processors-core systems with simultaneous task allocation, scheduling and voltage assignment
Lap-Fai Leung, Chi-Ying Tsui, Wing-Hung Ki |
ASP-DAC | 2 |
| 2004 | Minimizing energy consumption of hard real-time systems with simultaneous tasks scheduling and voltage assignment using statistical data
Lap-Fai Leung, Chi-Ying Tsui, Wing-Hung Ki |
ASP-DAC | 2 |
| 2004 | Fast adaptive DC-DC conversion using dual-loop one-cycle control in standard digital CMOS process
Wing-Hung Ki, Chi-Ying Tsui |
ASP-DAC | 3 |
| 2004 | Power control of CDMA systems with successive interference cancellation using the knowledge of battery power capacity
Chi-Ying Tsui, Roger S. Cheng, Wai Ho Mow |
ASP-DAC | 2 |
| 2004 | Re-Configurable Bus Encoding Scheme for Reducing Power Consumption of the Cross Coupling Capacitance for Deep Sub-Micron Instruction BusabstractIn very deep sub-micron designs, cross coupling capacitances become the dominant factor of the total bus loading and have a significant impact on the power consumption. In this paper, we propose two reconfigurable bus encoding schemes, which are based on the correlation among the bit lines, to reduce the power consumption at the cross coupling capacitances of the instruction buses. The instruction is encoded by flipping and reordering the bit lines during compilation time to reduce the total switching capacitances. A crossbar is used to map back the data to the original instruction code before sending to the instruction decoder. The reordering can be re-configured during run-time by using different configurations in the crossbar. We propose two types of re-configuration, static and dynamic. Static coding uses a fix flipping and re-configuring pattern after the corresponding program is compiled. Dynamic coding allows different re-configuring patterns during program execution. Experimental results show that by using the proposed schemes, significant energy reduction, 17-23%, can be achieved. Comparisons with existing bit lines reordering encoding scheme have also been made and on average more than 15% reduction can be obtained using our method. Siu-Kei Wong, Chi-Ying Tsui |
DATE | 2 |
| 2002 | Performance study of OFDM receiver using FFT based on log number systemabstractIn this paper, we study the performance of the OFDM receiver using FFT based on the log number system (Log-FFT) in a wireless environment. By using the log number system (LNS), complex multiplications of the FFT are reduced to simple additions. The effect of finite bit precision on the receiver performance is analyzed theoretically under the AWGN channel. Simulation results for both AWGN and fading channels are presented. It is shown that the bit width requirement of the receiver is quite small. In a coded OFDM system, there is no degradation in bit error rate performance when only two fractional bits are used for the LNS under a slow fading channel. The size of the look-up tables are small and the complexity of the log-FFT implementation is simple. Thus log-FFT is an attractive implementation method for the OFDM receiver design in wireless environment. Chi-Ying Tsui, Roger S. Cheng, Wai Ho Mow |
VTC Spring | 2 |
| 2001 | A single-inductor dual-output integrated DC/DC boost converter for variable voltage schedulingabstractAn integrated boost DC/DC converter that provides two different outputs with a 1.8V input using only one inductor is presented. The converter works in discontinuous conduction mode and employs time division multiplexing in switching the inductor current to the two outputs. Synchronous rectification for high efficiency is implemented. Techniques for current sensing, inductor ringing suppression and controller design are discussed. At an oscillator frequency of 1MHz, the conversion efficiency reaches 90% at 350mW. Wing-Hung Ki, Chi-Ying Tsui, Philip K. T. Mok |
ASP-DAC | 3 |
| 2001 | Reducing power consumption of turbo-code decoder using adaptive iteration with variable supply voltageabstractTurbo-code becomes popular for the next generation wireless communication systems because of its remarkable coding performance. One of the problems for decoding turbo-code in the receiver is the complexity and the high power consumption since multiple iterations of Soft Output Viterbi Algorithm (SOVA) or Maximum a posteriori (MAP) decoding have to be carried out to decode a data frame. To reduce the complexity of the turbo-code decoder, adaptive iteration based on cyclic redundancy checking (CRC) and output convergence approaches has been proposed to reduce the average number of iterations required for decoding a data frame. This results in a system that has variable workload since the amount of computation required for decoding each data frame is different. In this work, we propose a dynamic voltage scaling approach to further reduce the power consumption. Different from other variable workload systems, the workload here is not known at the time when the data is being decoded. Thus, optimum voltage assignment is not feasible. We propose several heuristic algorithms to assign supply voltage for different decoding iterations. Simulation results show that significant reduction of power consumption is achieved comparing with the system using fixed supply voltage. Oliver Yuk-Hang Leung, Chi-Ying Tsui, Roger S. Cheng |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | VLSI implementation of rake receiver for IS-95 CDMA Testbed using FPGAabstractNo abstract available. Oliver Yuk-Hang Leung, Chi-Ying Tsui, Roger S. Cheng |
ASP-DAC | 2 |
| 2000 | VLSI implementation of a switch fabric for mixed ATM and IP trafficabstractNo abstract available. Chi-Ying Tsui, Louis Chung-Yin Kwan, Chin-Tau A. Lea |
ASP-DAC | 1 |
| 2000 | Composite interference cancellation scheme for CDMA systemsabstractComplexity and detection delay are major issues for applying interference cancellation to practical CDMA systems. We introduce a new interference cancellation approach for multiuser detection. This detector combines both parallel interference cancellation (PIC) and successive interference cancellation (SIC) techniques. It has a short detection delay as in PIC and low computation complexity as in SIC. Billy Chi-Kin Poon, Chi-Ying Tsui, Roger S. Cheng |
GLOBECOM | 2 |
| 2000 | A reduced complexity implementation of the Log-Map algorithm for turbo-codes decodingabstractIn this paper, we propose a reduced complexity implementation scheme of the Log-Map algorithm for turbo-codes decoding. By re-arranging the structure of the computation and using adaptive approximation, the computation in the Log-Map algorithm such as the state metric and the log-likelihood ratio calculation could be simplified or reduced adaptively. Simulation results show that more than 45% of the computation can be reduced with almost no performance degradation. Chi-Ying Tsui, Roger S. Cheng |
ICASSP | 2 |
| 2000 | A low power VLSI architecture of SOVA-based turbo-code decoder using scarce state transition schemeabstractIn this paper, we propose a low power VLSI architecture of Soft-Output-Viterbi-Algorithm (SOVA) based turbo-code decoder using a scarce state transition scheme (SST). A register exchange survival memory unit (RE-SMU) and systolic block are used in the implementation of SOVA for high throughput and low latency. SST is used to reduce the power consumption. Simulation results show that the power consumption of RE-SMU and the systolic block is reduced significantly. The power consumption of the add-compare-select (ACS) block is also reduced by as much as 20% after 4 iterations of turbo-code decoding. Chi-Ying Tsui, Roger S. Cheng |
ISCAS | 2 |
| 2000 | Low complexity VLSI implementation of a joint successive interference cancellation with interleaving schemeabstractThe performance of wideband code division multiple access (WCDMA) can be severely degraded as the number of users increases due to the increase in interference. Multi-user detection is a scheme, which can cancel the interference generated by other users and hence can improve the system capacity. Recently a joint successive interference cancellation with interleaving (JSICI) scheme has been proposed for multi-user detection. However the complexity of the scheme is very high. In this paper, we propose a low complexity architecture which can implement this JSICI scheme efficiently in VLSI. Also low power features are proposed. Bob Ka-Man Wong, Chi-Ying Tsui, Roger S. Cheng |
ISCAS | 2 |
| 2000 | A low complexity architecture of the V-BLAST systemabstractThe V-BLAST architecture has been proposed as an extremely spectral efficient tool for wireless communications. However, owing to the intensive computation involved, it may be difficult to implement this architecture for high data rate communication system. We propose the replacement of the optimal decoding order by a suboptimal one and the utilization of the Gram-Schmitt orthogonalization (GSO) to substitute the computation of the pseudo-inverse in finding the weight vectors. These modifications significantly reduce the total number of arithmetic operations required to obtain the weight vectors and the decoding order in V-BLAST with virtually no performance degradation. In slowly time-varying fading channels, the proposed method can further reduce the amount of computation required by exploiting the time correlation between the channel gains. Wong Kwan Wai, Chi-Ying Tsui, Roger S. Cheng |
WCNC | 2 |
| 2000 | Low-power VLSI design for motion estimation using adaptive pixel truncationabstractPower consumption is very critical for portable video applications such as portable videophone and digital camcorder. Motion estimation (ME) in the video encoder requires a huge amount of computation, and hence consumes the largest portion of power. We propose a novel method of reducing power consumption of the ME by adaptively changing the pixel resolution during the computation of the motion vector. The pixel resolution is changed by masking or truncating the least significant bits of the pixel data, which is governed by the bit-rate control mechanism. Experimental results show that on average more than 4 bits ran be truncated without significantly affecting the picture quality. This results in more than 60% reduction in power consumption. Zhong-Li He, Chi-Ying Tsui, Kai-Keung Chan, Ming Lei Liou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1999 | Timing Optimization of Logic Network Using Gate DuplicationabstractWe present a timing optimization algorithm based on the concept of gate duplication on the technology-decomposed network. We first examine the relationship between gate duplication and delay reduction, and then introduce the notion of duplication gain for selecting the good candidate gates to be duplicated. The objective is to obtain the maximum delay reduction with the minimum duplications. The performance of the algorithm is demonstrated with experiments on benchmark circuits. Our approach can also be combined with other technology-independent timing optimizers (such as speed-up) to achieve further delay improvement. Chun-hong Chen, Chi-Ying Tsui |
ASP-DAC | 2 |
| 1999 | An Integrated Battery-Hardware Model for Portable ElectronicsabstractWe describe an integrated model of the hardware and the battery sub-systems in battery-powered VLSI systems. We demonstrate that, under this model and for a fixed operating voltage, the battery life decreases super-linearly as the average current dissipation increases. With the aid of analyses and empirical studies, we then show that the implications of this phenomenon are far-reaching and change our perceptions about low power design techniques targeted toward battery-powered VLSI circuits. Massoud Pedram, Chi-Ying Tsui, Qing Wu 0002 |
ASP-DAC | 2 |
| 1999 | Reducing power consumption of turbo code decoder using adaptive iteration with variable supply voltageabstractTurbo code becomes popular for the next generation wireless communication systems because of its remarkable coding performance. One of the problems for decoding turbo code in the receiver is the complexity and the high power consumption since multiple iterations of Soft Output Viterbi Algorithm (SOVA) have to be carried out to decode a data frame. In this paper, we address the issues of reducing the complexity and power consumption of the turbo code decoder. An approach using cyclic redundancy checking (CRC) to adaptively terminate the SOVA iteration of each frame is presented. This results in system that has variable workload of which the amount of computation required for each data frame is different. Dynamic voltage scaling is then used to further reduce the power consumption. However, since the workload is not yet known at the time when the data is being decoded, optimum voltage assignment is not feasible. In this work, we propose two heuristic algorithms to assign supply voltage for different decoding iterations. Simulation results show that significant reduction of power consumption is achieved comparing with system using fixed supply voltage. Oliver Yuk-Hang Leung, Chung-Wai Yue, Chi-Ying Tsui, Roger S. Cheng |
ISLPED | 3 |
| 1999 | Adaptive tracking of optimal bit and power allocation for OFDM systems in time-varying channelsabstractThis paper presents an adaptive tracking algorithm for adaptively updating the bit and power allocation for OFDM transmission in a time-varying channel. The proposed algorithm utilities the previous bit allocation information for initialization, and then iteratively refines the allocation until it reaches the optimal solution. We prove a theorem which guarantees the convergence of the algorithm to the optimal solution. Through simulation, we find that the proposed algorithm converges quickly in just a few iterations when the variation in the channel is small. When the channel changes rapidly the adaptive tracking algorithm is still computationally much more efficient than existing algorithms which solve the problem from scratch. Sai Kit Lai, Roger S. Cheng, Khaled Ben Letaief, Chi-Ying Tsui |
WCNC | 4 |
| 1998 | Towards the capability of providing power-area-delay trade-off at the register transfer levelabstractThis paper presents a new register-transfer level (RT-level) power estimation technique based on technology decomposition. Given the Boolean description of a circuit function, the power consumption of two typical circuit implementations, namely the minimum area implementation and the minimum delay implementation, are estimated, respectively. This provides a capability of obtaining a full power-delay-area trade-off curve at the RT level. Our method makes it possible to capture the structural and/or functional information of a circuit without going through actual gate-level implementation. Experimental results show that the accuracy is very reasonable. Chun-hong Chen, Chi-Ying Tsui |
ISLPED | 2 |
| 1998 | Accurate and efficient power simulation strategy by compacting the input vector set
Chi-Ying Tsui, Massoud Pedram |
Integr. | 1 |
| 1998 | Gate-level power estimation using tagged probabilistic simulationabstractIn this paper, we present a probabilistic simulation technique to estimate the power consumption of a CMOS circuit under a general delay model. This technique is based on the notion of a tagged (probability) waveform, which models the set of all possible events at the output of each circuit node. Tagged waveforms are obtained by partitioning the logic waveform space of a circuit node according to the initial and final values of each logic waveform and compacting all logic waveforms in each partition by a single tagged waveform. To improve the efficiency of tagged probabilistic simulation, only tagged waveforms at the circuit inputs are exactly computed. The tagged waveforms of the remaining nodes are computed using a compositional scheme that propagates the tagged waveforms from circuit inputs to circuit outputs. We obtain significant speed up over explicit simulation methods with an average error of only 6%. This also represents a factor of 2-3/spl times/ improvement in accuracy of power estimates over previous probabilistic simulation approaches. Chih-Shun Ding, Chi-Ying Tsui, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1998 | Low-power state assignment targeting two- and multilevel logic implementationsabstractThe problem of minimizing power consumption during the state encoding of a finite-state machine is addressed. A new power cost model for state encoding is proposed, and encoding techniques that minimize this power cost for two- and multilevel logic implementations are described. These techniques are compared with those that minimize area or the switching activity at the present state bits. Experimental results show significant improvements. Chi-Ying Tsui, Massoud Pedram, Alvin M. Despain |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1997 | A Power Estimation Framework for Designing Low Power Portable Video ApplicationsabstractThis paper presents a power evaluation framework designedfor estimating power consumption of a new video telephonecompression standard, ITU-H.263, at the system level.A hierarchical,mixed-level simulation environment is built andcycle-accurate power macro-modeling is used for the architecturalpower evaluation.Experimental results show the effectivenessof the proposed framework and models. Chi-Ying Tsui, Kai-Keung Chan, Qing Wu 0002, Chih-Shun Ding, Massoud Pedram |
DAC | 1 |
| 1997 | An efficient and reconfigurable VLSI architecture for different block matching motion estimation algorithmsabstractThis paper describes a VLSI architecture which can be reconfigured to support both Full Search Block-Matching algorithm and 3-step Hierarchical Search Block-Matching algorithm. By using a reconfigurable register-mux array and a parameterizable adder tree, the 2-D array architecture provides efficient real time motion estimation for many video applications. We also propose a memory architecture and an associated switching network to solve the simultaneous data access problem. Xiao-Dong Zhang 0001, Chi-Ying Tsui |
ICASSP | 2 |
| 1997 | Low power motion estimation design using adaptive pixel truncationabstractPower consumption is very critical for portable video applications such as portable video-phone. Motion estimation in the video encoder requires huge amount of computation and hence consumes the largest portion of power. In this paper we propose a novel method of reducing power consumption of the motion estimation by adaptively changing the pixel resolution during the computation of the motion vector. The pixel resolution is changed by mashing or truncating the LSBs of the pixel data which is governed by an adaptive mechanism. Experimental results show that on average more than 4 bits can be truncated without affecting the picture quality. This results in an average 70% reduction in power consumption. Zhong-Li He, Kai-Keung Chan, Chi-Ying Tsui, Ming Lei Liou |
ISLPED | 3 |
| 1996 | Improving the Efficiency of Power Simulators by Input Vector CompactionabstractAccurate power estimation is essential for low power digital CMOS circuit design.Power dissipation is input pattern dependent.To obtain an accurate power estimate, a large input vector set must be used which leads to very long simulation time.One solution is to generate a compact vector set that is representative of the original input vector set and can be simulated in a reasonable time.In this paper, we propose an input vector compaction technique that preserves the statistical properties of the original sequence.Experimental results show that a compaction ratio of 100X is achieved with less than 2% average error in the power estimates. Chi-Ying Tsui, Radu Marculescu, Diana Marculescu, Massoud Pedram |
DAC | 1 |
| 1996 | Correction to "Power Estimation Methods for Sequential Logic Circuits" [Correspondence]
Chi-Ying Tsui, José Monteiro 0001, Massoud Pedram, Srini Devadas, Alvin M. Despain, Bill Lin 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1995 | Power estimation methods for sequential logic circuitsabstractRecently developed methods for power estimation have primarily focused on combinational logic. We present a framework for the efficient and accurate estimation of average power dissipation in sequential circuits. Switching activity is the primary cause of power dissipation in CMOS circuits. Accurate switching activity estimation for sequential circuits is considerably more difficult than that for combinational circuits, because the probability of the circuit being in each of its possible states has to be calculated. The Chapman-Kolmogorov equations can be used to compute the exact state probabilities in steady state. However, this method requires the solution of a linear system of equations of size 2/sup N/ where N is the number of flip-flops in the machine. We describe a comprehensive framework for exact and approximate switching activity estimation in a sequential circuit. The basic computation step is the solution of a nonlinear system of equations which is derived directly from a logic realization of the sequential machine. Increasing the number of variables or the number of equations in the system results in increased accuracy. For a wide variety of examples, we show that the approximation scheme is within 1-3% of the exact method, but is orders of magnitude faster for large circuits. Previous sequential switching activity estimation methods can have significantly greater inaccuracies.> Chi-Ying Tsui, José Monteiro 0001, Massoud Pedram, Srini Devadas, Alvin M. Despain, Bill Lin 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 1994 | Exact and Approximate Methods for Calculating Signal and Transition Probabilities in FSMsabstractIn this paper, we consider the problem of calculating the signal and transition probabilities of the internal nodes of the combinational logic part of a finite state machine (FSM). Given the state transition graph (STG) of the FSM, we first calculate the state probabilities by iteratively solving the Chapman-Kolmogorov equations. Using these probabilities, we then calculate the exact signal and transition probabilities by an implicit state enumeration procedure. For large sequential machines where the STG cannot be explicitly built, we unroll the next state logic k times and estimate the signal probability of the state bits using an OBDD-based approach. The basic computation step consists of solving a system of non-linear equations. We then use these estimates to approximately calculate signal and transition probabilities of the internal nodes. Our experimental results indicate that the average errors of transition probabilities and power estimation(compared to the exact method) are on... Chi-Ying Tsui, Massoud Pedram, Alvin M. Despain |
DAC | 1 |
| 1994 | Low power state assignment targeting two-and multi-level logic implementations
Chi-Ying Tsui, Massoud Pedram, Chih-Ang Chen, Alvin M. Despain |
ICCAD | 1 |
| 1994 | Power efficient technology decomposition and mapping under an extended power consumption modelabstractWe propose a new power consumption model that accounts for the power consumption at the internal nodes of a CMOS gate. Next, we address the problem of minimizing the average power consumption during the technology dependent phase of logic synthesis. Our approach consists of two steps. In the first step, we generate a NAND decomposition of an optimized Boolean network such that the sum of average switching rates for all nodes in the network is minimum. In the second step, we perform a power efficient technology mapping that finds a minimal power mapping for given timing constraints (subject to the unknown load problem).> Chi-Ying Tsui, Massoud Pedram, Alvin M. Despain |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 1993 | Technology Decomposition and Mapping Targeting Low Power DissipationabstractIn his paper, we address the problem of minimizing the average power dissipation during the technology dependent phase of logic synthesis. Our approach consists of two steps. In the first step, we generate a Nand decomposition of an optimized Boolean network such that the sum of average switching rates for all nodes in the network is minimum. Our power-efficient decomposition procedure is optimal for dynamic CMOS circuits with uncorrelated input signals and produces very good results for static CMOS. In the second step, we perform a power efficient technology mapping that finds an optimal power delay trade-off value (subject to the unknown load problem) four given timing constraints. We obtain an average of 21% improvement in power at the expense of 12.6% increase in area and without any degradation in performance on a number of benchmarks. Chi-Ying Tsui, Massoud Pedram, Alvin M. Despain |
DAC | 1 |
| 1993 | Efficient estimation of dynamic power consumption under a real delay modelabstractIn CMOS circuits, glitches account for a sizeable part of the total power consumption. In this paper, we present a fast and memory efficient power estimation technique for CMOS circuits which estimates the power consumed due to the glitches. Our technique is based on the notion of tagged transition waveforms. In particular, we approximate the correlation between transition waveforms for two signal lines by the correlation between the steady state values of these lines. We obtain an order of magnitude speed up over an exact method with an average error of only 1%. Chi-Ying Tsui, Massoud Pedram, Alvin M. Despain |
ICCAD | 1 |
| 1992 | Application-Driven Design Automation for Microprocessor Design
Iksoo Pyo, Ching-Long Su, Ing-Jer Huang, Kuo-Rueih Pan, Yong-Seon Koh, Chi-Ying Tsui, Hsu-Tsun Chen, Gino Cheng, Shihming Liu, Shiqun Wu, Alvin M. Despain |
DAC | 6 |