Chen-Yi Lee

dblp:89/1481 · DBLP profile ↗
← Back
97ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-6795-0874ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 52 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1
YearPublicationVenuePosition
2026 FastDEP: A CMOS DEP Chip with Frequency and Voltage Scaling for Cell Biology Applications
Yu-Chen Chang, Lin-Hung Lai, Shao-Hua Lian, Yu-Chen Hung, Wen-Yue Lin, Bang-Yuan Xiao, Fang-Chen Lo, Chen-Yi Lee
ISCAS8
2026 A Programmable CMOS Chip for Mobile Nucleic Acid Amplification Tests
Jhan-Yi Liao, Hsi-Hao Huang, Yi-Xuan Ran, Chen-Yi Lee
ISCAS4
2025 M*: On-Chip Microfluidic Operations With A-Star for Portable Diagnostics
abstract
Precise microfluidic control is essential for biomedical diagnostics. Compared to traditional PCB-based DMFB, standard CMOS-based digital microfluidic biochips (DMFB) can integrate sensing, heating, and actuation on a single chip, eliminating the need for external equipment [1]. However, CMOS-DMFB are limited by their lower actuation voltage (usually around 90 V [2]), and their actuation stability is often not as good as that of PCB-DMFB, which can be operated at voltages as high as 300 V [3]. This is due to the fact that, according to research, the EWOD force is proportional to the square of the actuation voltage [4]. In other words, a CMOS-DMFB will only generate approximately 4% of the actuation force of a PCB-DMFB, making actuation not as easy as expected [5]. To alleviate this limitation, we would like to utilize the real-time sensing function on the chip as an actuation feedback. After each actuation, the sensing results can be used to re-adjust the path and determine whether the destination has been reached. Therefore, we designed M* algorithms for on-chip microfluidic operation using A* algorithms to improve the actuation efficiency and accuracy, which is important for high-precision biomedical applications.
Cheng-Hsuan Hsieh, Eugene Lee 0001, Peng Hsien Chi, Jiajie Diao, Chen-Yi Lee
ICASSP5
2025 Multiple Sampling and Pixel-Wise Accumulation in CMOS Capacitive Sensor Array System for Real-Time Droplet Analysis
abstract
Capacitive sensor array (CSA) is vital in precise monitoring for lab-on-chip (LOC) systems. However, shrinking electrode sizes to increase spatial resolution bring challenges like noise interference and large data volumes. This paper presents an FPGA-based system that addresses these problems with multiple sampling (MS) and pixel-wise accumulation (PWA). MS reduces Gaussian noise by sampling multiple frames and retaining only representative data points, while PWA compresses data using Block RAM and minimal combinational logic, reducing size from 118 Mb to 0.46 Mb and boosting SNR to 25.30 dB. The system enables real-time monitoring every 5 seconds instead of 17 minutes, with pipeline sensing and transmission further optimizing sensing time. Experiments demonstrate its effectiveness in distinguish between droplets and monitor evaporation in real time. MS and PWA can be easily integrated into future chip designs, offering scalable solutions for fast and precise monitoring in LOC environments.
Lin-Hung Lai, Wen-Yue Lin, Yu-Chen Hung, Yu-Hsian Wang, Hsi-Hao Huang, Chen-Yi Lee
ISCAS6
2025 Smart Pattern Generation on Programmable Dielectrophoresis Array Chip for Single Particle Manipulation
abstract
Dielectrophoresis (DEP) is a powerful tool for manipulating biological cells. However, single cell manipulation is usually time-consuming and skill-intensive. This paper presents a system that integrates AI for real-time image recognition with a programmable dielectrophoresis (DEP) array chip for automated particle manipulation. The system comprises a DEP chip, an FPGA, a computer, a microscope, and a server. The YOLO v8 model is used to detect particle positions within microscope images and generate DEP manipulation patterns. The system utilizes a Breadth-First Search (BFS) algorithm for path planning, ensuring collision-free movement of particles within a grid structure. Experimental results demonstrated the system’s effectiveness in manipulating 20 μm polystyrene particles with a success rate of over 90%. This system offers a significant advancement in automated DEP-based manipulation, providing precise control at micro scales with high computational efficiency.
Yu-Hsiang Wang, Wen-Yue Lin, Lin-Hung Lai, Chen-Yi Lee
ISCAS4
2024 A Programmable CMOS Dielectrophoresis Array Chip with 128 × 128 Electrodes for Cell Manipulation
abstract
Dielectrophoresis (DEP) is a powerful technique for manipulating biological cells. Yet, its widespread application has been limited by traditional glass-based chips with static electrode configurations that often require integrated microfluidic systems. This paper presents a novel DEP array chip fabricated in a standard CMOS process, featuring a 128 × 128 electrode matrix capable of generating dynamic, programmable electric field patterns that can be tailored to meet specific requirements for different use cases. Experiments have demonstrated the ability of the chip to manipulate fibroblast and THP-1 cells, with fibroblast movement observed at a velocity of 10µm/s with a DEP frequency of 800kHz and a peak-to-peak DEP voltage of 1.8V. The chip is designed for compatibility with standard petri dishes, obviating the requirement for microfluidics and facilitating its integration with traditional cell culture protocols. Our results indicate the chip’s potential as a versatile tool for cell biology research and applications.
Wen-Yue Lin, Lin-Hung Lai, Yi-Wei Lin, Chen-Yi Lee
ISCAS4
2024 A 2.56-µs Dynamic Range, 31.25-ps Resolution 2-D Vernier Digital-to-Time Converter (DTC) for Cell-Monitoring
abstract
Capacitive sensor array (CSA) has emerged as a prominent approach in the field of biomedical detection, particularly for the analysis of cell morphology and the construction of a Cell-on-CMOS platform. To enhance overall sensitivity, this paper presents a novel 2-D vernier digital-to-time converter (2D V-DTC) with a resolution of 31.25 ps and a dynamic range of 2.56 µs. The 2-D vernier structure attains sufficient resolution while reducing the number of delay elements required, and seamlessly integrates with a counter-based controller, extending the dynamic range to effectively cover a significantly larger sensing window. The CSA biochip, fabricated in 180-nm CMOS technology, achieves results with an overall sensitivity of 880 code/fF. This achievement translates into an sensing resolution of 1.13 aF, demonstrating its potential to further advance the development of CMOS-based cell-monitoring platform.
Heng-Yu Liu, Lin-Hung Lai, Wen-Yue Lin, Yu-Wei Lu, Yi-Wei Lin, Chen-Yi Lee
ISCAS6
2023 A Pattern-Control Digital Microfluidic Bio-Chip for Fast Thermal Cycle in Nucleic Acid Amplification Tests
abstract
A pattern-control digital biochip is proposed for fast medical tests. With integrated circuit modules in each basic element, also known as micro-electrode, this biochip can achieve digital microfluidic operations, capacitive sensing, and thermal cycle via different control patterns. As a result, bio-protocols can be derived from target biomedical tests to reach better test accuracy on the proposed chip. For the mentioned fast medical tests, samples/reagents can be identified first by capacitive sensing, followed by microfluidic and thermal cycle operations. Preliminary measurements show that heating/cooling rate of 5°C/sec can be achieved and demonstrate each thermal cycle$(95\rightarrow 55\rightarrow 72)$for polymerase chain reaction (PCR) can be completed in less than 20 seconds with power consumption of 256–444 uW per micro-electrode while dealing with nano-liter samples. This implies both test time and power consumption per sample test can be further improved, making our proposed biochip very suitable for point-of-care test (POCT) applications.
Yun-Sheng Chan, Jiajie Diao, Chen-Yi Lee
ISCAS3
2023 A Stack-Based In-Pixel Storage Circuit for SPAD Photon Counting
abstract
Single-photon avalanche diodes (SPADs) have attracted a lot of attention these days because of the ability to detect a single photon for many emerging applications. However, planar SPAD sensor arrays often suffer from serious photon loss because the readout bottleneck dominates the overall dead time. This paper presents a stack-based in-pixel storage circuit for high-throughput SPAD imaging. The proposed circuit helps a SPAD imaging chip solve the buffer saturation problem and reduces its dead time by half compared to single-bit storage. Fabricated in TSMC HV$0.18\ \mu \mathrm{m}$CMOS technology, each pixel in the SPAD array can record at most three photons in 50 ns, resulting in 40Mfps. The minimum integration time to form an 8-bit image is reduced to$6.4\ \mu \mathrm{s}$while maintaining global shutter exposure.
Tzu-Yun Huang, Hsi-Hao Huang, Chun-Hsien Liu, Sheng-Di Lin, Chen-Yi Lee
ISCAS5
2023 Self-Restoring and Low-Jitter Circuits for High Timing-Resolution SPAD Sensing Applications
abstract
Single-photon avalanche diode (SPAD) imagers have recently emerged in the fields of high-resolution imaging, such as fluorescence lifetime imaging microscopy. This paper presents a new design of passive quenching active reset (PQAR) circuit for high timing-resolution SPAD imagers. The PQAR circuit contains a logic control circuit for self-restoring and dead time reduction in the event of two adjacent incoming photons. Post-layout simulation result shows that the dead time of the sensor has been reduced to 5.46 ns for two adjacent incoming photons, and measurement result shows that the jitter of the photon-induced signal width has been reduced to 63ps. We also present the properties and measurement results of the passive quenching passive reset circuit and the passive quenching active clock-driven reset circuit. Comparison shows that of the three circuits we demonstrate, the PQAR circuit is better suited for the high timing-resolution SPAD imagers.
Hsi-Hao Huang, Chun-Hsien Liu, Tzu-Yun Huang, Sheng-Di Lin, Chen-Yi Lee
ISCAS5
2023 Cross-Resolution Flow Propagation for Foveated Video Super-Resolution
abstract
The demand of high-resolution video contents has grown over the years. However, the delivery of high-resolution video is constrained by either computational resources required for rendering or network bandwidth for remote transmission. To remedy this limitation, we leverage the eye trackers found alongside existing augmented and virtual reality headsets. We propose the application of video super-resolution (VSR) technique to fuse low-resolution context with regional high-resolution context for resource-constrained consumption of high-resolution content without perceivable drop in quality. Eye trackers provide us the gaze direction of a user, aiding us in the extraction of the regional high-resolution context. As only pixels that falls within the gaze region can be resolved by the human eye, a large amount of the delivered content is redundant as we can’t perceive the difference in quality of the region beyond the observed region. To generate a visually pleasing frame from the fusion of high-resolution region and low-resolution region, we study the capability of a deep neural network of transferring the context of the observed region to other regions (low-resolution) of the current and future frames. We label this task a Foveated Video Super-Resolution (FVSR), as we need to super-resolve the low-resolution regions of current and future frames through the fusion of pixels from the gaze region. We propose Cross-Resolution Flow Propagation (CRFP) for FVSR. We train and evaluate CRFP on REDS dataset on the task of 8× FVSR, i.e. a combination of 8× VSR and the fusion of foveated region. Departing from the conventional evaluation of per frame quality using SSIM or PSNR, we propose the evaluation of past foveated region, measuring the capability of a model to leverage the noise present in eye trackers during FVSR. Code is made available at https://github.com/eugenelet/CRFP.
Eugene Lee 0001, Lien-Feng Hsu, Chen-Yi Lee
WACV4
2021 Few-Shot and Continual Learning with Attentive Independent Mechanisms
abstract
Deep neural networks (DNNs) are known to perform well when deployed to test distributions that shares high similarity with the training distribution. Feeding DNNs with new data sequentially that were unseen in the training distribution has two major challenges — fast adaptation to new tasks and catastrophic forgetting of old tasks. Such difficulties paved way for the on-going research on few-shot learning and continual learning. To tackle these problems, we introduce Attentive Independent Mechanisms (AIM). We incorporate the idea of learning using fast and slow weights in conjunction with the decoupling of the feature extraction and higher-order conceptual learning of a DNN. AIM is designed for higher-order conceptual learning, modeled by a mixture of experts that compete to learn independent concepts to solve a new task. AIM is a modular component that can be inserted into existing deep learning frameworks. We demonstrate its capability for few-shot learning by adding it to SIB and trained on MiniImageNet and CIFAR-FS, showing significant improvement. AIM is also applied to ANML and OML trained on Omniglot, CIFAR-100 and MiniImageNet to demonstrate its capability in continual learning. Code made publicly available at https://github.com/huang50213/AIM-Fewshot-Continual.
Eugene Lee 0001, Cheng-Han Huang, Chen-Yi Lee
ICCV3
2020 NeuralScale: Efficient Scaling of Neurons for Resource-Constrained Deep Neural Networks
abstract
Deciding the amount of neurons during the design of a deep neural network to maximize performance is not intuitive. In this work, we attempt to search for the neuron (filter) configuration of a fixed network architecture that maximizes accuracy. Using iterative pruning methods as a proxy, we parametrize the change of the neuron (filter) number of each layer with respect to the change in parameters, allowing us to efficiently scale an architecture across arbitrary sizes. We also introduce architecture descent which iteratively refines the parametrized function used for model scaling. The combination of both proposed methods is coined as NeuralScale. To prove the efficiency of NeuralScale in terms of parameters, we show empirical simulations on VGG11, MobileNetV2 and ResNet18 using CIFAR10, CIFAR100 and TinyImageNet as benchmark datasets. Our results show an increase in accuracy of 3.04%, 8.56% and 3.41% for VGG11, MobileNetV2 and ResNet18 on CIFAR10, CIFAR100 and TinyImageNet respectively under a parameter-constrained setting (output neurons (filters) of default configuration with scaling factor of 0.25).
Eugene Lee 0001, Chen-Yi Lee
CVPR2
2020 Meta-rPPG: Remote Heart Rate Estimation Using a Transductive Meta-learner
Eugene Lee 0001, Chen-Yi Lee
ECCV (27)3
2020 Cross-Domain Adaptation for Biometric Identification Using Photoplethysmogram
abstract
The adoption of biomedical signals such as photoplethysmogram (PPG) and electrocardiogram (ECG) for health parameter estimation on wearable devices is growing in tandem with the increase of attention in mobile healthcare. In our work, we use PPG signals extracted from PPG sensors which are used for biometric identification. A challenge for biometric identification using PPG signal is the variation in domain (placement of sensors, wavelengths, device variation, etc.). In this work, we propose the use of both unsupervised and semi-supervised adversarial learning techniques for cross-domain adaptation. As such algorithm will be deployed on wearable devices, we propose a compact model meeting tight memory footprint limitation. All experiments will be simulated using a public dataset (TROIKA) and our in-house dataset. By introducing a cross-domain adaptation approach across sensors, we observe an accuracy gain of 4.15% on our in-house dataset. The proposed semi-supervised learning technique gives an additional accuracy boost of 2.02%.
Eugene Lee 0001, Annie Ho, Cheng-Han Huang, Chen-Yi Lee
ICASSP5
2020 An Area-Efficient High-Throughput SM4 Accelerator with SCA-Countermeasure for TV Applications
abstract
The SM4 algorithm is the first commercial cipher published by China in 2012 which is widely used in WLAN WAPI resource restricted devices. This paper proposed the single-round-iterative architecture which can operate at 500 MHZ clock frequency and reach 2Gbps throughput. In order to resist side channel attack, we changed the S-box structure and add secret sharing during the computation process. According to the CPA result, this hardware is secure under the condition of collecting 1 million power traces. The gate count of this design is about 15.19k, gaining almost 17.7% area reduction to the BMTSM4 which can reach the similar throughput [1].
Wei Chiang, Hsie-Chia Chang, Chen-Yi Lee
ISCAS3
2020 Convolutional neural networks for classification of music-listening EEG: comparing 1D convolutional kernels with 2D kernels and cerebral laterality of musical influence
Kit Hwa Cheah, Humaira Nisar, Yap Vooi Voon, Chen-Yi Lee
Neural Comput. Appl.4
2020 Multitarget Sample Preparation Using MEDA Biochips
abstract
Sample preparation, as a key procedure in many biochemical protocols, mixes various samples, and/or reagents into solutions that contain the target concentrations. Digital microfluidic biochips (DMFBs) have been adopted as a platform for sample preparation because they provide automatic procedures that require less reactant consumption and reduce human-induced errors. However, the most existing methods only consider two-reactant sample preparation, and they cannot be used for many biochemical applications that involve multiple reactants. In addition, the existing methods that can be used for multiple-reactant sample preparation were proposed on traditional DMFBs where only the (1:1) mixing model is available. In the (1:1) mixing model, only two droplets of the same volume can be mixed at a time, which results in higher completion time and the wastage of valuable reactants. To overcome this limitation, the micro-electrode-dot-array (MEDA) architecture has been introduced; it provides the flexibility of mixing multiple droplets of different volumes in a single operation. In this article, we present a generic multiple-reactant sample preparation algorithm that exploits the novel fluidic operations on MEDA biochips. We also propose an enhanced algorithm that increases the operation-sharing opportunities when multiple target concentrations are needed, and therefore the usage of reactants can be further reduced. The simulated experiments show that the proposed method outperforms existing methods in terms of saving reactant cost, minimizing the number of operations, and reducing the amount of waste.
Tung-Che Liang, Yun-Sheng Chan, Tsung-Yi Ho, Krishnendu Chakrabarty, Chen-Yi Lee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 Feature Consistency Training With JPEG Compressed Images
abstract
Deep neural networks (DNNs) are recently found to be vulnerable to JPEG compression artifacts, which distort the feature representations of DNNs leading to serious accuracy degradation. Most existing training methods which aim to address this problem add compressed images to the training data to enhance the robustness of DNNs. However, their improvements are limited since these methods usually regard the compressed images as new training samples instead of distorted samples. The feature distortions between the raw images and the compressed images are not investigated. In this work, we propose a new training method, called Feature Consistency Training, that is designed to minimize the feature distortions caused by JPEG artifacts. At each training iteration, we simultaneously input a raw image and its compressed version with a randomly sampled quality into a DNN model and extract the features from the internal layers. By adding feature consistency constraint to the objective function, the feature distortions in the representation space are minimized in order to learn robust filters. Besides, we present a residual mapping block which takes the quality factor of the compressed image as an additional information to further reduce the feature distortion. Extensive experiments demonstrate that our method outperforms several existed training methods on JPEG compressed images. Furthermore, DNN models trained by our method are found to be more robust to unseen distortions.
Sheng Wan, Tung-Yu Wu, Heng-Wei Hsu, Wing Hung Wong, Chen-Yi Lee
IEEE Trans. Circuits Syst. Video Technol.5
2020 FADE: Feature Aggregation for Depth Estimation With Multi-View Stereo
abstract
Both structural and contextual information is essential and widely used in image analysis. However, current multi-view stereo (MVS) approaches usually use a single common pre-trained model as pixel descriptor to extract features, which mix structural and contextual information together and thus increase the difficulty of matching correspondence. In this paper, we propose FADE (feature aggregation for depth estimation), which treats spatial and context information separately and focuses on aggregating features for efficient learning of the MVS problem. Spatial information includes image details such as edges and corners, whereas context information comprises object features such as shapes and traits. To aggregate these multi-level features, we use an attention mechanism to select important features for matching. We then build a plane sweep volume by using a homography backward warping method to generate match candidates. Furthermore, we propose a novel cost volume regularization network aims to minimize the noise in the matching candidates. Finally, we take advantage of 3D stacked hourglass and regression to produces high-quality depth maps. With these well-aggregated features, FADE can efficiently perform dense depth reconstruction, achieving state-of-the-art performance in terms of accuracy and requiring the least amount of model parameters.
Hsiao-Chien Yang, Po-Heng Chen, Kuan-Wen Chen, Chen-Yi Lee, Yong-Sheng Chen
IEEE Trans. Image Process.4
2019 Sample preparation for multiple-reactant bioassays on micro-electrode-dot-array biochips
abstract
Sample preparation, as a key procedure in many biochemical protocols, mixes various samples and/or reagents into solutions that contain the target concentrations. Digital microfluidic biochips (DMFBs) have been adopted as a platform for sample preparation because they provide automatic procedures that require less reactant consumption and reduce human-induced errors. However, traditional DMFBs only utilize the (1:1) mixing model, i.e., only two droplets of the same volume can be mixed at a time, which results in higher completion time and the wastage of valuable reactants. To overcome this limitation, a next-generation micro-electrode-dot-array (MEDA) architecture that provides flexibility of mixing multiple droplets of different volumes in a single operation was proposed. In this paper, we present a generic multiple-reactant sample preparation algorithm that exploits the novel fluidic operations on MEDA biochips. Simulated experiments show that the proposed method outperforms existing methods in terms of saving reactant cost, minimizing the number of operations, and reducing the amount of waste.
Tung-Che Liang, Yun-Sheng Chan, Tsung-Yi Ho, Krishnendu Chakrabarty, Chen-Yi Lee
ASP-DAC5
2019 Joint Capacitive Sensing and Frequency Selection for Fast Medical Tests
abstract
This paper presents a new approach with frequency selection for fast medical tests. Many medical tests need high speed sensor because of its short reaction time. High frequency biosensor with high sensitivity plays a vital role in fast medical tests. Each medical test has its own reaction time. How to select proper frequency for specific medical test is an important issue. In this paper, one frequency-based capacitive sensor is proposed to prove this new approach with salt water. According to the experimental results, the limit of detection (LOD) can be improved if proper working frequency is selected. The 0.1%wt salt water in 20ul can be detected on proposed system and coefficient of variation (CV) is only 0.19% at 40MHz.
Yun-Sheng Chan, Kuan-Yu Lung, Chen-Yi Lee
ISCAS4
2019 Centralized State Sensing using Sensor Array on Wearable Device
abstract
Signal acquisition from the wrist using wearable device is usually corrupted by noise, sometimes up to a level where the noise completely dominates the signal. In this paper, we use the acquisition of photoplethysmography (PPG) signal as an example to demonstrate our proposed algorithm implemented on a sensor array. To increase the chances of acquiring clean PPG signal, we develop a uniform linear sensor array located above the radial artery. We also develop a wearable device to integrate our sensor array and our proposed algorithm. We propose centralized state sensing (CSS), a centralized algorithm suited to our sensor array, to increase the efficiency in multi-sensor estimation. To support our proposed algorithm, we use heart rate estimation as an example and compare it with the estimation of a single sensor and the statistical mean of the estimation of multiple sensors. Experimental results demonstrate that our proposal achieve a reduction in mean absolute error of 26.8%, compared to the estimation result obtained using the average of the estimation of all sensors. Operating on a 80 MHz processor, our algorithm introduces a 8 ms overhead, making it highly suitable for wearable devices.
Eugene Lee 0001, Tsu-Jui Hsu, Chen-Yi Lee
ISCAS3
2019 QuatNet: Quaternion-Based Head Pose Estimation With Multiregression Loss
abstract
Head pose estimation has attracted immense research interest recently, as its inherent information significantly improves the performance of face-related applications such as face alignment and face recognition. In this paper, we conduct an in-depth study of head pose estimation and present a multiregression loss function, an L2 regression loss combined with an ordinal regression loss, to train a convolutional neural network (CNN) that is dedicated to estimating head poses from RGB images without depth information. The ordinal regression loss is utilized to address the nonstationary property observed as the facial features change with respect to different head pose angles and learn robust features. The L2 regression loss leverages these features to provide precise angle predictions for input images. To avoid the ambiguity problem in the commonly used Euler angle representation, we further formulate the head pose estimation problem in quaternions. Our quaternion-based multiregression loss method achieves state-of-the-art performance on the AFLW2000, AFLW test set, and AFW datasets and is closing the gap with methods that utilize depth information on the BIWI dataset.
Heng-Wei Hsu, Tung-Yu Wu, Sheng Wan, Wing Hung Wong, Chen-Yi Lee
IEEE Trans. Multim.5
2018 Diabetic Retinopathy Detection Based on Deep Convolutional Neural Networks
abstract
Diabetic retinopathy is the primary cause of blindness in the working-age population of the developed world. Diagnosing the disease heavily relies on imaging studies, which is a time consuming and a manual process performed by trained clinicians. Enhancing the accuracy and speed of the detection process can potentially have a significant impact on population health via early diagnosis and intervention. Motivated by this, we propose a recognition pipeline based on deep convolutional neural networks. In our pipeline, we design lightweight networks called SI2DRNet-vl along with six methods to further boost the detection performance. Without any fine-tuning, our recognition pipeline outperforms state of the art on the Messidor dataset along with 5.26x fewer in total parameters and 2.48x fewer in total floating operations.
Yi-Wei Chen, Tung-Yu Wu, Wing Hung Wong, Chen-Yi Lee
ICASSP4
2018 Correlation-Based Face Detection for Recognizing Faces in Videos
abstract
Finding the locations and identities of faces in videos is a very important task in numerous applications. In this paper, we propose a correlation-based face detection approach to improve the performance of face recognition tasks for videos. We apply correlation measures to pairs of response maps which are generated from automatically selected neurons in deep convolutional neural network (CNN) models to detect faces in each video frame. The embeddings extracted from faces cropped by our proposed approach are more consistent across each video sequence and more suitable for face recognition and clustering tasks. Experimental results from the YouTube Faces (YTF) dataset demonstrate that our proposed approach is more robust and achieves better recognition accuracy compared to state-of-the-art face detection approaches.
Heng-Wei Hsu, Tung-Yu Wu, Wing Hung Wong, Chen-Yi Lee
ICASSP4
2018 Confnet: Predict with Confidence
abstract
In this paper, we propose Confidence Network (ConfNet) which not only makes predictions on input images but also generates a confidence score that estimates the probability of correctness of each prediction. Furthermore, Confidence Loss is proposed to make ConfNet automatically learn confidence scores in the training phase. The experiments on two public datasets show that the confidence scores generated by ConfNet are highly correlated with the model accuracy and outperforms two related methods. When stacking two ConfNets in a cascade structure, 3.8x computational cost can be saved compared to the single state-of-the-art model with only 0.1 % increase of error rate.
Sheng Wan, Tung-Yu Wu, Wing Hung Wong, Chen-Yi Lee
ICASSP4
2018 Efficient and Adaptive Error Recovery in a Micro-Electrode-Dot-Array Digital Microfluidic Biochip
abstract
A digital microfluidic biochip (DMFB) is an attractive technology platform for automating laboratory procedures in biochemistry. In recent years, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have been proposed. MEDA biochips can provide advantages of better capability of droplet manipulation and real-time sensing ability. However, errors are likely to occur due to defects, chip degradation, and the lack of precision inherent in biochemical experiments. Therefore, an efficient error-recovery strategy is essential to ensure the correctness of assays executed on MEDA biochips. By exploiting MEDA-specific advances in droplet sensing, we present a novel error-recovery technique to dynamically reconfigure the biochip using real-time data provided by on-chip sensors. Local recovery strategies based on probabilistic-timed-automata are presented for various types of errors. An online synthesis technique and a control flow are also proposed to connect local-recovery procedures with global error recovery for the complete bioassay. Moreover, an integer linear programming-based method is also proposed to select the optimal local-recovery time for each operation. Laboratory experiments using a fabricated MEDA chip are used to characterize the outcomes of key droplet operations. The PRISM model checker and three benchmarks are used for an extensive set of simulations. Our results highlight the effectiveness of the proposed error-recovery strategy.
Kelvin Yi-Tse Lai, John McCrone, Po-Hsien Yu, Krishnendu Chakrabarty, Miroslav Pajic, Tsung-Yi Ho, Chen-Yi Lee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2018 Structural and Functional Test Methods for Micro-Electrode-Dot-Array Digital Microfluidic Biochips
abstract
A digital microfluidic biochip (DMFB) is an attractive platform for immunoassays, point-of-care clinical diagnostics, DNA sequencing, and other laboratory procedures in biochemistry. More recently, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have been proposed, and droplet manipulations on MEDA biochips have also been experimentally demonstrated. In order to ensure robust fluidic operations and high confidence in the outcome of biochemical experiments, MEDA biochips must be adequately tested before they can be used for bioassay execution. This paper presents the first approach for testing of MEDA biochips that include both CMOS circuits and microfluidic components. We first present structural test techniques to evaluate the pass/fail status of each microcell (droplet actuation, droplet maintenance, and droplet sensing) and identify faulty microcells. In order to ensure correct operation of functional units, e.g., mixers and diluters, we also present functional test techniques to address fundamental MEDA operations, such as droplet dispensing, transportation, mixing, and splitting. We evaluate the proposed test methods using simulations as well as experiments for fabricated MEDA biochips.
Kelvin Yi-Tse Lai, Po-Hsien Yu, Krishnendu Chakrabarty, Tsung-Yi Ho, Chen-Yi Lee
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2017 Design considerations and clinical applications of closed-loop neural disorder control SoCs
abstract
This paper presents the closed-loop neural disorder control concept and some design considerations. Two architectures of closed-loop neuromodulation for Parkinson's disease and epileptic seizure are proposed. One is a closed-loop deep brain stimulator, which meets the IEC 60601-1 standard. The other one is an implantable SoC for epileptic seizure control, which is verified by animal experiment.
Chung-Yu Wu, Cheng-Hsiang Cheng, Yi-Huan Ou-Yang, Chiung-Ghu Chen, Wei-Ming Chen, Ming-Dou Ker, Chen-Yi Lee, Sheng-Fu Liang, Fu-Zen Shaw
ASP-DAC7
2016 High-level synthesis for micro-electrode-dot-array digital microfluidic biochips
abstract
A digital microfluidic biochip (DMFB) is an attractive technology platform for automating laboratory procedures in biochemistry. However, today's DMFBs suffer from several limitations: (i) constraints on droplet size and the inability to vary droplet volume in a fine-grained manner; (ii) the lack of integrated sensors for real-time detection; (iii) the need for special fabrication processes and reliability/yield concerns. To overcome the above problems, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have recently been demonstrated. However, due to the inherent differences between today's DMFBs and MEDA, existing synthesis solutions cannot be utilized for MEDA-based biochips. We present the first biochip synthesis approach that can be used for MEDA. The proposed synthesis method targets operation scheduling, module placement, routing of droplets of various sizes, and diagonal movement of droplets in a two-dimensional array. Simulation results using benchmarks and experimental results using a fabricated MEDA biochip demonstrate the effectiveness of the proposed co-optimization technique.
Kelvin Yi-Tse Lai, Po-Hsien Yu, Tsung-Yi Ho, Krishnendu Chakrabarty, Chen-Yi Lee
DAC6
2016 Error recovery in a micro-electrode-dot-array digital microfluidic biochip?
abstract
A digital microfluidic biochip (DMFB) is an attractive technology platform for automating laboratory procedures in biochemistry. However, today's DMFBs suffer from several limitations: (i) constraints on droplet size and the inability to vary droplet volume in a fine-grained manner; (ii) the lack of integrated sensors for real-time detection; (iii) the need for special fabrication processes and the associated reliability/yield concerns. To overcome the above problems, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have been proposed recently, and droplet manipulation on these devices has been experimentally demonstrated. Errors are likely to occur due to defects, chip degradation, and the lack of precision inherent in biochemical experiments. Therefore, an efficient error-recovery strategy is essential to ensure the correctness of assays executed on MEDA biochips. By exploiting MEDA-specific advances in droplet sensing, we present a novel error-recovery technique to dynamically reconfigure the biochip using real-time data provided by on-chip sensors. Local recovery strategies based on probabilistic-timed-automata are presented for various types of errors. A control flow is also proposed to connect local recovery procedures with global error recovery for the complete bioassay. Laboratory experiments using a fabricated MEDA chip are used to characterize the outcomes of key droplet operations. The PRISM model checker and three analytical chemistry benchmarks are used for an extensive set of simulations. Our results highlight the effectiveness of the proposed error-recovery strategy.
Kelvin Yi-Tse Lai, Po-Hsien Yu, Krishnendu Chakrabarty, Miroslav Pajic, Tsung-Yi Ho, Chen-Yi Lee
ICCAD7
2016 Design of a micro-electrode cell for programmable lab-on-CMOS platform
abstract
This paper presents a programmable lab-on-CMOS (LoCMOS) with micro-electrode cell array. Array structure is suitable for programmable like CMOS VLSIs. In order to improve the utilization, each micro-electrode cell is composed of actuation and sensing circuit. In addition, a CMOS-compatible extended drain MOSFET (EDMOS) is adopted under a 3V supply. This LoCMOS platform is composed of 1,800 microelectrodes with exploiting EDMOS to enable droplet actuations. Through its field programmability, the chip can successfully perform all microfluidic operations, droplet moving/cutting/mixing on a 2-dimenional microelectrode cell array. Implemented in 0.35um standard CMOS process, the LoCMOS platform demonstrates microfluidic functions and droplet detection. Measured results show successfully for actuation and real-time droplet location sensing.
Yingchieh Ho, Gary Wang, Kelvin Yi-Tse Lai, Yi-Wen Lu, Keng-Ming Liu, Chen-Yi Lee
ISCAS7
2016 Built-in self-test for micro-electrode-dot-array digital microfluidic biochips
abstract
A digital microfluidic biochip (DMFB) is an attractive platform for immunoassays, point-of-care clinical diagnostics, DNA sequencing, and other laboratory procedures in biochemistry. However, today's DMFBs suffer from several limitations, including (i) the lack of integrated sensors for real-time detection, (ii) constraints on droplet size and the inability to vary droplet volume in a fine-grained manner, and (iii) the need for special fabrication processes and the associated reliability/yield concerns. To overcome the above limitations, DMFBs based on a micro-electrode-dot-array (MEDA) architecture have been proposed recently. Droplet manipulation on MEDA biochips has also been experimentally demonstrated. In order to ensure robust fluidic operations and high confidence in the outcome of biochemical experiments, MEDA biochips must be adequately tested before they can be used for bioassay execution. We present an efficient built-in self-test (BIST) architecture for MEDA biochips. The proposed BIST architecture can effectively detect defects in a MEDA biochip, and faulty microcells can be identified. Simulation results based on HSPICE and experiments using fabricated MEDA biochips highlight the effectiveness of the proposed BIST architecture.
Kelvin Yi-Tse Lai, Po-Hsien Yu, Krishnendu Chakrabarty, Tsung-Yi Ho, Chen-Yi Lee
ITC6
2015 A 3.46 Gb/s (9141, 8224) LDPC-based ECC scheme and on-line channel estimation for solid-state drive applications
abstract
As the reliability of NAND Flash memory keeps degrading, Low-Density Parity-Check (LDPC) codes are widely proposed to extend the endurance of Solid State Drive (SSD). However, implementing powerful decoding algorithm such as soft min-sum algorithm with high decoding speed comes along with higher hardware cost. To achieve efficient hardware cost, we propose a multi-strategy ECC scheme which consists of modified gradient descent bit-flipping (MGDBF), hard min-sum, and soft min-sum decoders. The MGDBF decoder aims to correct most of the erroneous codewords with advantages of high decoding throughput and low hardware cost, while the soft min-sum decoder is targeted to correct codewords with large number of errors under moderate decoding throughput and reasonable hardware cost. In addition, we propose a bi-sectional channel estimation technique which enables on-line estimation of distribution to generate accurate soft information for LDPC decoding with low complexity. The ECC codec and the complete Toggle DDR 1.0 NAND interface control circuits are integrated and fabricated in 90nm CMOS process. The throughput of proposed MGDBF decoder achieves 3.46 Gb/s which satisfies the throughput requirement of both toggle DDR 1.0 and 2.0 NAND interfaces.
Kin-Chu Ho, Chih-Lung Chen, Yen-Chin Liao, Hsie-Chia Chang, Chen-Yi Lee
ISCAS5
2015 Efficient Hardware Architecture of ηT Pairing Accelerator Over Characteristic Three
abstract
To support emerging pairing-based protocols related to cloud computing, an efficient algorithm/hardware codesign methodology of ηT pairing over characteristic three is presented. By mathematical manipulation and hardware scheduling, a single Miller's loop can be executed within 17 clock cycles. Furthermore, we employ torus representation and exploit the Frobenius map to lower the computation cost of final exponentiation. Pipelining and parallelization datapath are also exploited to shorten the critical path delay. Finally, by choosing suitable multiplier architecture and selecting an appropriate number of multipliers, Miller's loop and final exponentiation can be computed in a fully pipelined manner. With these schemes, a test chip for the proposed pairing accelerator has been fabricated in 90-nm CMOS 1P9M technology with a core area of 1.52 × 0.97 mm2. It performs a bilinear pairing computation over F(397) in 4.76 μs under 1.0 V supply and achieves 178% improvement to relative works in terms of area-time (AT) product. To support higher level of security, a 126-bit secure pairing accelerator that can complete a bilinear pairing computation over F(3709) in 36.2 μs is implemented and this result is at least 31% better than relative works in terms of AT product.
Szu-Chi Chung, Jing-Yu Wu, Hsing-Ping Fu, Jen-Wei Lee, Hsie-Chia Chang, Chen-Yi Lee
IEEE Trans. Very Large Scale Integr. Syst.6
2015 An MPCN-Based BCH Codec Architecture With Arbitrary Error Correcting Capability
abstract
This paper presents an area-efficient architecture of arbitrary error correction Bose-Chaudhuri-Hocquenghem codec for NAND flash memory. By factorizing the generator polynomial into several minimal polynomials and utilizing linear feedback shift registers based on minimal polynomials, our reconfigurable design cannot only support multiple error correcting capabilities at a few extra cost, but also merge the encoder and syndrome calculator for efficiently reducing hardware complexity. After being implemented in CMOS 65-nm technology, the test chip supporting t = 1-24 bits can achieve 1.33-Gb/s measured throughput with 73k gate-count while another design supporting t = 60-84 bits can provide 1.60-Gb/s synthesized throughput with 168.6k gate-count.
Chi-Heng Yang, Yi-Min Lin, Hsie-Chia Chang, Chen-Yi Lee
IEEE Trans. Very Large Scale Integr. Syst.4
2014 A 2 GOPS quad-mean shift processor with early termination for machine learning applications
abstract
This paper proposes a 2 GOPS quad-mean shift processor (Q-MSP) architecture for data clustering and machine learning applications. By exploiting the linear approximation approach and early termination mechanism, the proposed algorithm can reduce 70% and 40% computational complexity, respectively. Moreover, 4 mean shift processor cores are integrated into the proposed architecture to support parallel processing to further improve system performance. Implemented in Xilinx Virtex-7 FPGA, this architecture occupies 65k LUTs and 3.3MB block memory to achieve 2 GOPS throughput operated at 125MHz.
Chang-Hung Tsai, Hui-Hsuan Lee, Wan-Ju Yu, Chen-Yi Lee
ISCAS4
2014 Area-efficient TFM-based stochastic decoder design for non-binary LDPC codes
abstract
This paper presents a non-binary LDPC decoder based on stochastic arithmetic. Although the previous stochastic works reduce the complexity of check node by transforming the convolution of the SPA algorithm to the finite field summation, the stochastic decoder still has a implementation bottleneck due to large storage introduced by the variable node process. Considering a balance between algorithm level and implementation level, we propose a shortened TFM architecture as well as its updating criterion. A compare-and-alter counter architecture is also proposed to avoid sorting among counters which decide the decoded codeword. With these features, the proposed (136, 68) fully-parallel stochastic NB-LDPC decoder over GF(32) implemented in UMC 90-nm can achieve 120 Mb/s throughput while operating under 455 MHz with 740 k gate counts which are only 10 % of the original TFM decoder.
Chih-Wen Yang, Xin-Ru Lee, Chih-Lung Chen, Hsie-Chia Chang, Chen-Yi Lee
ISCAS5
2014 Efficient Power-Analysis-Resistant Dual-Field Elliptic Curve Cryptographic Processor Using Heterogeneous Dual-Processing-Element Architecture
abstract
Elliptic curve cryptography (ECC) for portable applications is in high demand to ensure secure information exchange over wireless channels. Because of the high computational complexity of ECC functions, dedicated hardware architecture is essential to provide sufficient ECC performance. Besides, crypto-ICs are vulnerable to side-channel information leakage because the private key can be revealed via power-analysis attacks. In this paper, a new heterogeneous dual-processing-element (dual-PE) architecture and a priority-oriented scheduling of right-to-left double-and-add-always EC scalar multiplication (ECSM) with randomized processing technique are proposed to achieve a power-analysis-resistant dual-field ECC (DF-ECC) processor. For this dual-PE design, a memory hierarchy with local memory synchronization scheme is also exploited to improve data bandwidth. Fabricated in a 90-nm CMOS technology, a 0.4- mm2160-b DF-ECC chip can achieve 0.34/0.29 ms 11.7/9.3 μJ for one GF(p)/GF(2m) ECSM. Compared to other related works, our approach is advantageous not only in hardware efficiency but also in protection against power-analysis attacks.
Jen-Wei Lee, Szu-Chi Chung, Hsie-Chia Chang, Chen-Yi Lee
IEEE Trans. Very Large Scale Integr. Syst.4
2013 Improved High Code-Rate Soft BCH Decoder Architectures With One Extra Error Compensation
abstract
Compared with traditonal hard Bose-Chaudhuri-Hochquenghem (BCH) decoders, soft BCH decoders provide better error-correcting performance but much higher hardware complexity. In this brief, an improved soft BCH decoding algorithm is presented to achieve both competitive hardware complexity and better error-correcting performance by dealing with least reliable bits and compensating one extra error outside the least reliable set. For BCH (255, 239; 2) and (255, 231; 3) codes, our proposed soft BCH decoders can achieve up to 0.75-dB coding gain with one extra error compensation and 5% less complexity than the traditional hard BCH decoders.
Yi-Min Lin, Hsie-Chia Chang, Chen-Yi Lee
IEEE Trans. Very Large Scale Integr. Syst.3
2012 An Efficient Countermeasure against Correlation Power-Analysis Attacks with Randomized Montgomery Operations for DF-ECC Processor
Jen-Wei Lee, Szu-Chi Chung, Hsie-Chia Chang, Chen-Yi Lee
CHES4
2012 A high-performance elliptic curve cryptographic processor over GF(p) with SPA resistance
abstract
In order to support high speed application such as cloud computing, we propose a new elliptic curve cryptographic (ECC) processor architecture. The proposed processor includes a 3 pipelined-stage full-word Montgomery multiplier which requires much fewer execution cycles than that of previous methods. To reach real-time requirement, the time-cost pre-computation steps of Montgomery modular multiplication are achieved by hardware as well. Moreover, our proposed processor is resistant to the simple power analysis (SPA) attack by using the Montgomery ladder-based elliptic curve scalar multiplication (ECSM). Even the Montgomery ladder method inherently has operation overhead compared with traditional binary ECSM, both of hardware sharing and parallelization techniques are exploited to improve the hardware performance. Synthesized in TSMC 90nm CMOS technology, our proposed ECC processor performs a 256-bit ECSM in 120µs over prime field with 540K gate counts. This result is at least 25% better than relative works in terms of area-time (AT) product.
Szu-Chi Chung, Jen-Wei Lee, Hsie-Chia Chang, Chen-Yi Lee
ISCAS4
2012 Stochastic decoding for LDPC convolutional codes
abstract
Among LDPC codes, LDPC convolutional codes (LDPC-CCs) seem to be more suitable for variable length applications. However, a LDPC-CC decoder is difficult to implement for its long latency and large storage usage. The stochastic computation makes the decoding of LDPC-CCs more efficient, but the boundary effect of sliding window causes poor performance. In this paper, a stochastic LDPC-CC decoder with virtual edge compensation as well as decoder architecture is presented. The simulation results based on (491, 3, 6) time-varying LDPC-CC show that under the same signal-to-noise ratio, our proposed decoder could achieve better performance, 60% less decoding latency and 40% storage reduction compared to log-BP decoder with 10 processors.
Xin-Ru Lee, Chih-Lung Chen, Hsie-Chia Chang, Chen-Yi Lee
ISCAS4
2012 Extrinsic data compression method for double-binary turbo codes
abstract
This paper presents an extrinsic data compression method for double-binary turbo codes. Frame and extrinsic data memory occupy more area in turbo decoder implementations with the frame size increasing. Besides, non-binary turbo codes have much more extrinsic memory usage than single-binary turbo codes. This proposed compression method utilizes an operation in radix-4 single-binary turbo decoder and also can simplify the integration of single and double-binary turbo decoders in hardware implementations. It can reduce one-third of extrinsic memory size, about 15% of total memory usage. Simulation results show that this method only causes 0.2dB performance loss when bit error rate is equal to 10−5with slight hardware increment.
Yi-Huan Ou-Yang, Chien-Yu Kao, Jen-Yuan Hsu, Pangan Ting, Chen-Yi Lee
ISCAS5
2012 A Low Voltage All-Digital On-Chip Oscillator Using Relative Reference Modeling
abstract
This paper presents a low voltage on-chip oscillator which can compensate process, voltage, and temperature (PVT) variation in an all-digital manner. The relative reference modeling applies a pair of ring oscillators as relative references and estimates period of the internal ring oscillator. The period estimation is parameterized by a second-order polynomial. Accordingly, the oscillator compensates frequency variations in a frequency division fashion. A 1-20 MHz adjustable oscillator is implemented in a 90-nm CMOS technology with 0.04 mm area. The fabricated chips are robust to variations of supply voltage from 0.9 to 1.1 V and temperature range from 0°C to 75°C. The low supply voltage and the small area make it suitable for low-cost and low-power systems.
Chien-Ying Yu, Jui-Yuan Yu, Chen-Yi Lee
IEEE Trans. Very Large Scale Integr. Syst.3
2011 A dual-field elliptic curve cryptographic processor with a radix-4 unified division unit
abstract
To enhance the data security in network communications, this paper presents a dual-field elliptic curve crypto- graphic processor (DECP) supporting all finite field operations and elliptic curve (EC) functions. Based on the fast radix-4 unified division algorithm, the execution time can be significantly reduced by a factor of three. By exploiting the hardware sharing and the ladder selection techniques, the proposed 160-bit and 256-bit DECP can have competitive execution cycle with only 0.29mm2and 0.45mm2silicon area in 90nm CMOS technology. In addition, the operating frequency in dual field can be increased by applying the data-path separation method and the degree checker. Our proposed DECP is over 2~6 times better in area- time product than relative works.
Yao-Lin Chen, Jen-Wei Lee, Po-Chun Liu, Hsie-Chia Chang, Chen-Yi Lee
ISCAS5
2011 An area-efficient high-accuracy prediction-based CABAC decoder architecture for H.264/AVC
abstract
This paper proposed a high-accuracy prediction scheme and area-efficient CABAC decoder architecture for H.264 video decoder. To alleviate hardware cost and keep high throughput, we propose the prediction process and optimize the memory system. In particular, simulation results show that the proposed prediction-based CABAC decoder module achieves over 90% hit rate and requires only 16K logic gates with 3,360 bits SRAM by UMC 90 nm technology. The proposed architecture operates on 150 MHz frequency (Max. 249 MHz) for realizing 1080HD video playback at 30 fps, which can achieve Level 5.0 MP in tiny gate count.
Ming-Yu Kuo, Chen-Yi Lee
ISCAS3
2011 Hardware-Assisted Reliability Enhancement for Embedded Multi-core Virtualization Design
abstract
In this paper, we propose a virtualization architecture for the multi-core embedded system to provide more system reliability and security while maintaining the same performance without introducing additional special hardware supports or having to implement complex protection mechanism in the virtualization layer. Virtualization has been widely used in embedded systems, especially in consumer electronics, albeit itself is not a new technique, because there are various needs for both GPOS (General Purpose Operating System) and RTOS (Real Time Operating System). The surge of the multi-core platform in the embedded system also helps the consolidation of the virtualization system for its better performance and lower power consumption. Embedded virtualization design usually uses two kinds of approaches. The first one is to use the traditional VMM, but it is too complicated for use in the embedded environment if there is no additional special hardware support. The other is the use of the micro kernel which imposes a modular design. The guest systems, however, would suffer from considerable amount of modifications because the micro kernel lets the guest systems to run in user space. For some RTOSes and theirs applications originally running in kernel space, it makes this approach more difficult to work because a lot of privileged instructions are used in those codes. To achieve better reliability and keep the virtualization layer design light weighted, a common hardware component adopted in the multi-core embedded processors is used in this work. In the most embedded platforms, vendors provide additional on-chip local memory for each physical core and these local memory areas are private only to their cores. By taking this memory architecture's advantage, we can mitigate above-mentioned problems at once. We choose to re-map the virtualization layer's program called SPUMONE, which it runs all its guest systems in kernel space, on the local memory. By doing so, it can provide additional reliability and security for the entire system because the SPUMONE's design in a multi-core platform has each instance being installed on a separated processor core which is different from the traditional virtualization layer design and the content of each SPUMONE is inaccessible to each others. We also achieve this goal without bringing any overhead to the overall performance.
Tsung-Han Lin, Yuki Kinebuchi, Alexandre Courbot, Hiromasa Shimada, Takushi Morita, Hitoshi Mitake, Chen-Yi Lee, Tatsuo Nakajima
ISORC7
2011 Hardware-Assisted Reliability Enhancement for Embedded Multi-core Virtualization Design
abstract
In this paper, we propose a virtualization architecture for the multi-core embedded system to provide more sys-tem reliability and security while maintaining the same performance without introducing additional special hardware supports or having to implement complex protection mechanism in the virtualization layer. Embedded virtualization design usually uses two kinds of approaches, traditional VMM and microkernel approaches, but both of them suffer from performance or engineering cost problems. To achieve better reliability and keep the virtualization layer design light weighted, a common hardware component called local memory adopted in the multi-core embedded processors is used in this work. By taking this memory ar-chitecture's advantage, we can mitigate above-mentioned problems at once. We choose to re-map the virtualizationlayer's program called SPUMONE, which it runs all its guest systems in kernel space, on the local memory. By doing so, it can provide additional reliability and security for the entire system because the SPUMONE's design in a multi-core platform has each instance being installed on a separated processor core, which is different from the traditional virtualization layer design, and therefore the content of each SPUMONE in the local memory is inaccessible to each others.
Tsung-Han Lin, Yuki Kinebuchi, Hiromasa Shimada, Hitoshi Mitake, Chen-Yi Lee, Tatsuo Nakajima
RTCSA (2)5
2011 A Low-Power and Portable Spread Spectrum Clock Generator for SoC Applications
abstract
In this paper, a novel portable and all-digital spread spectrum clock generator (ADSSCG) suitable for system-on-chip (SoC) applications with low-power consumption is presented. The proposed ADSSCG can provide flexible spreading ratios by the proposed rescheduling division triangular modulation (RDTM). Thus it can provide different EMI attenuation performance for various system applications. Furthermore, the proposed ADSSCG employs a low-power digitally controlled oscillator (DCO) to save overall power consumption significantly. Measurement results show that power consumption of the proposed ADSSCG is 1.2 mW (@54 MHz), and it provides 9.5 dB EMI reductions with 1% spreading ratio. Besides, the proposed ADSSCG has very small chip area as compared with conventional SSCGs which often required large on-chip loop filter capacitors. In addition, the proposed ADSSCG is implemented only with standard cells, making it easily portable to different processes and very suitable for SoC applications.
Duo Sheng, Ching-Che Chung, Chen-Yi Lee
IEEE Trans. Very Large Scale Integr. Syst.3
2010 An improved soft BCH decoder with one extra error compensation
abstract
In existing soft decision algorithms, a soft BCH decoders provides better error correcting performance but has much higher hardware complexity than a traditional hard BCH decoder. In this paper, a soft BCH decoder with both better error correcting performance and lower complexity is presented. The low complexity feature of the proposed architecture is achieved by dealing with least reliable bits. By compensating extra one error outside the least reliable set, the error correcting ability is improved. In addition, the proposed error locator evaluator evaluates error locations without Chien search, leading to high throughput. As compared with the traditional hard BCH decoder, the experimental result reveals that our proposed improved soft BCH decoder can achieve 0.75db coding gain for BCH (255,239) code. Implemented in standard CMOS 90nm technology, it can reach 316.3Mb/s throughput at 360MHz operation frequency with gate-count of 4.06K according to the post-layout simulations.
Yi-Min Lin, Hsie-Chia Chang, Chen-Yi Lee
ISCAS3
2009 Design of an Intra Predictor with Data Reuse for High-profile H.264 Applications
abstract
This paper presents a high-profile intra predictor for H.264 video decoder. To alleviate the starved bandwidth of intra compensation in high-definition video, we reuse the neighboring pixels and optimize the buffer size and access latency. In particular, a dedicated pixel buffer reuses neighboring pixel for realizing MB-adaptive frame-field (MBAFF) decoding in intra compensation. Moreover, a base-mode predictor is explored to optimize the area efficiency for reference sample filtering process (RSFP) in intra 8times8 modes. Simulation results show that the proposed data-reused intra prediction module requires 14 K logic gates and 688 bits SRAM, and operates on 100 MHz frequency for realizing 1080 HD video playback at 30 fps.
Yu-Fan Lai, Tsu-Ming Liu, Chen-Yi Lee
ISCAS4
2009 An eCrystal Oscillator with Self-calibration Capability
abstract
This paper reports an embedded crystal (eCrystal) oscillator which is capable of compensating process, voltage, and temperature (PVT) variation. The delay behaviors of two different delay cells are exploited for the delay estimation. Then the estimated delay is parameterized by a second order approximated function. According to the estimated delay, the eCrystal oscillator is able to calibrate the frequency error without any external reference. A 40 MHz oscillator is implemented in a 90-nm CMOS technology with area of 0.4 mm2by an all-digital approach. The fabricated chip is measured under supply voltage from 0.9 V to 1.1 V and temperature range from 0degC to 75degC with active power consumption of 237 muW.
Chien-Ying Yu, Jui-Yuan Yu, Chen-Yi Lee
ISCAS3
2008 Multi-mode message passing switch networks applied for QC-LDPC decoder
abstract
The multi-mode message passing switch networks for multi-standard QC-LDPC decoder are presented. An enhanced self-routing switch network with only one barrel shifter permutation structure and a shifter-based two-way duplicated switch network are proposed to support 19 and 3 different sub-matrices defined in IEEE 802.16e and IEEE 802.11n. These proposed switch networks can route the decoding message in parallel by different sizes without signal congestion. The enhanced self-routing switch network can switch the messages at different expansion factors with the lowest hardware complexity. Under the condition of a smaller expansion factor, the decoder throughput can be enhanced from the two-way duplicated switch network by increasing the parallelism. In the 130 nm CMOS synthesis result, the proposed enhanced self-routing and the two-way duplicated switch network gate counts are 21.9 k and 37.4 k at 384 MHz operation frequency.
Chih-Hao Liu, Chien-Ching Lin, Hsie-Chia Chang, Chen-Yi Lee, Yarsun Hsua
ISCAS4
2007 An In/Post-Loop Deblocking Filter With Hybrid Filtering Schedule
abstract
In this paper, we propose a high-throughput deblocking filter to perform the in-loop or post-loop filtering process for different standard requirements. The performance improvement is very mild if we replace a post-loop filter with an in-loop filter. To alleviate this problem, we derive an integration-oriented algorithm that can be reconfigured as the in-loop or post-loop filter. Moreover, we develop a hybrid filtering schedule to reach a lower bound of processing cycles. In particular, we reschedule the filtering order and reuse the intermediate pixels when the deblocking filter switches the filtered edges from vertical to horizontal direction. Finally, a 0.18-mum CMOS design that performs the in/post-loop filter with the hybrid filtering schedule is implemented. The synthesized gate counts are 21.1 K which is reduced to 70% of preliminary design that performs the in-loop or post-loop filter separately. Moreover, it achieves 4times105macroblock/s of throughput rate at a 100-MHz clock rate.
Tsu-Ming Liu, Wen-Ping Lee, Chen-Yi Lee
IEEE Trans. Circuits Syst. Video Technol.3
2007 A New Soft Variable Length Decoder for Wireless Video Transmission
abstract
In this paper, we propose an adaptive and scalable soft variable length code (VLC) decoder to greatly reduce overall design complexity. Generally, a soft VLC decoder needs to maintain many states for the correct decoding when the sequence length or table size grows. We propose an adaptive sorting scheme to reduce the memory accesses and design complexity. We reduce the table size by using symbol-merging and table-merging schemes. In addition, the proposed Black-Box model improves the accuracy of performance estimation by a novel measurement of "symbol-alias" and also achieves a better tradeoff between performance and complexity. Further, no side information is transmitted and the proposed soft VLC decoder is bandwidth-efficient. The proposed design is evaluated using the model of MPEG-4/UDP-Lite/UEP/AWGN, and hence it is standard-compliant. We averagely improve the peak signal-to-noise ratio by 0.4~2.9 dB as compared with traditional VLC decoding and standard-support reversible VLC decoding schemes
Tsu-Ming Liu, Sheng-Zen Wang, Bai-Jue Shieh, Chen-Yi Lee
IEEE Trans. Circuits Syst. Video Technol.4
2006 Design of a 125muW, fully-scalable MPEG-2 and H.264/AVC video decoder for mobile applications
abstract
A design of MPEG-2 and H.264/AVC video decoder is demonstrated in a 0.18?m CMOS [1]. The key design issues involved in this advanced IC are discussed, including improving area and power efficiency. Power dissipation is greatly lowered through the architectural exploration. Measurement results show that MPEG-2 and H.264/AVC real-time decoding of [email protected] are achieved at 1.15MHz with power dissipation of 108?W and 125?W respectively at 1V supply voltage.
Tsu-Ming Liu, Ching-Che Chung, Chen-Yi Lee, Ting-An Lin, Sheng-Zen Wang
DAC3
2006 A Context Adaptive Bit-Plane Coder With Maximum-Likelihood-Based Stochastic Bit-Reshuffling Technique for Scalable Video Coding
abstract
In this paper, we propose a context adaptive bit-plane coding (CABIC) with a stochastic bit reshuffling (SBR) scheme to deliver higher coding efficiency and better subjective quality for fine granular scalable (FGS) video coding. Traditional bit-plane coding in FGS algorithm suffers from poor coding efficiency and subjective quality. To improve coding efficiency, our CABIC constructs context models based on both the energy distribution in a block and the spatial correlations in the adjacent blocks. Moreover, it exploits the context across bit-planes to save side information. To improve subjective quality, our SBR reorders the coefficient bits by their estimated rate-distortion performance. Particularly, we model transform coefficients with Laplacian distributions and incorporate them into the context probability models for content-aware parameter estimation. Moreover, our SBR is implemented with a dynamic priority management that uses a low-complexity dynamic memory organization. Experimental results show that our CABIC improves the PSNR by 0.5/spl sim/1.0 dB at medium and high bit rates. While maintaining similar or even higher coding efficiency, our SBR improves the subjective quality.
Wen-Hsiao Peng, Tihao Chiang, Hsueh-Ming Hang, Chen-Yi Lee
IEEE Trans. Multim.4
2006 A low power turbo/Viterbi decoder for 3GPP2 applications
abstract
This paper presents a channel decoder that completes both turbo and Viterbi decodings, which are pervasive in many wireless communication systems, especially those that require very low signal-to-noise ratios. The trellis decoding algorithm merges them with less redundancy. However, the implementation is still challenging due to the power consumption in wearable devices. This research investigates an optimized memory scheme and rescheduled data flow to reduce power consumption and chip area. The memory access is reduced by buffering the input symbols, and the area is reduced by reducing the embedded interleaver memory. A test chip is fabricated in a 1.8 V 0.18-/spl mu/m standard CMOS technology and verified to provide 4.25-Mb/s turbo decoding and 5.26-Mb/s Viterbi decoding. The measured power dissipation is 83 mW, while decoding a 3.1 Mb/s turbo encoded data stream with six iterations for each block. The power consumption in Viterbi decoding is 25.1 mW in the 1-Mb/s data rate. The measurement shows the power dissipation is 83 mW for the turbo decoding with six iterations at 3.1 Mb/s, and 25.1 mW for the Viterbi decoding at 1 Mb/s.
Chien-Ching Lin, Yen-Hsu Shih, Hsie-Chia Chang, Chen-Yi Lee
IEEE Trans. Very Large Scale Integr. Syst.4
2005 An area-efficient and high-throughput de-blocking filter for multi-standard video applications
abstract
In this paper, we propose an area-efficient design approach to cover both in-loop and post-loop filtering processes for multiple video coding standards. In addition, we propose a hybrid filter scheduling to improve system throughput. Compared with available designs [Yu-Wen Huang et al, 2003][M. Sima et al, 2004], the proposed approach saves about one-half of processing cycles, and hence reduces power dissipation. Compared to the original loop-filter, the proposed loop/post filter only incurs 20.7% of extra cost. Simulation results show that our proposal can easily achieve real-time decoding for 1080HD when the working frequency is 100 MHz.
Tsu-Ming Liu, Wen-Ping Lee, Chen-Yi Lee
ICIP (3)3
2004 A low-complexity soft vlc decoder using performance modeling
abstract
In this paper, we propose a scalable soft VLC decoder to greatly reduce the overall design complexity. Generally, the soft VLC decoder needs to maintain many states for the correct decoding when the table size grows. We reduce the table size by using a symbol-merging scheme. We merge two symbols with the-same prefix into one. Further, to achieve the optimal trade-off between performance and complexity, we propose a black-box model. In our model, we present a novel measurement of "symbol-alias" to improve the accuracy of performance estimation. Experimental results show that our scalable soft VLC decoder using performance modeling has more than 1 dB PSNR gain and offers better subjective quality compared to traditional VLC decoding.
Tsu-Ming Liu, Chen-Yi Lee
ICIP2
2003 Error-resilient image coding (ERIC) with smart-IDCT error concealment technique for wireless multimedia transmission
abstract
The fields of multimedia and wireless communications have grown rapidly, leading to a great demand for an image coding scheme that has both compression and error-resilient capabilities. Because bandwidth is a valuable and limited resource, the compression technique is applied to wireless multimedia communication. However, strong data dependency will be created while the bit-rate reduction is achieved. Transmission errors always results in significant quality degradation. An error-resilient image coding for discrete cosine transform-based image compression is proposed. It can successfully prevent errors from propagating across image block boundaries with little overhead. Additionally, a novel post-processing error concealment scheme is presented to retain low-frequency information and discard suspicious high-frequency information. Low-resolution information, rather than total corruption, can be obtained during the image-decoding process. Because of low complexity and latency properties, it is very suitable for wireless mobile applications. Simulation results show that good image quality (PSNR=31.78dB) and a low fraction of corruptive blocks (less than 5%) can be achieved even when the bit error rate is 0.1%.
Yew-San Lee, Keng-Khai Ong, Chen-Yi Lee
IEEE Trans. Circuits Syst. Video Technol.3
2003 Two-level hierarchical Z-buffer with compression technique for 3D graphics hardware
Cheng-Hsien Chen, Chen-Yi Lee
Vis. Comput.2
2002 FPGA education and research activities in Taiwan
abstract
The role of Chip Implementation Center (CIC), founded in 1992 under the National Science Council (NSC) of Taiwan R.O.C., is to provide the services for the fabrication of multi-project chip (MPC), the procurement/integration of software CAD tools, and the promotion of IC and FPGA design/testing/CAD software technology for academia in Taiwan. To date, CIC assisted 86 universities and polytechnics to install over 6000 academic licenses of FPGA design and verification tools. Moreover, CIC promotes various training courses intensively and periodically for the FPGA design flow. In year 2002, 8 kinds of courses with over 25 classes are provided to meet the demands from academia sites. More than 1000 students are trained in these classes. The FPGA design flow, provided by CIC, is used in many of researches for implementation, verification, and prototypes. Based on the demands of academics, CIC will build a laboratory for rapid prototyping system-level design. In this laboratory, SoC (System on a Chip) and IP (Intellectual Property) designs can be downloaded into FPGA to work with a processor to verification on an SOPC (System on Programmable Chip) environment.
Yu-Tsang Chang, Yu-Te Chou, Wei-Chang Tsai, Jiann-Jenn Wang, Chen-Yi Lee
FPT5
2002 A novel DCT-based bit plane error resilient entropy coding for wireless multimedia communication
abstract
With the increasing demand for digital media, there has been significant interest in the deployment of image and video compression techniques. However, the presence of variable-length codes in most of the compression standards like the JPEG and the MPEG are highly sensitive to channel noise. Any bit of transmission error will cause the decoder to lose synchronization and the reconstructed image quality will be significantly degraded. Much effort has been invested in R&D to build error resilience into the compressed bit-stream and ensure improving the quality of image/video wireless transmission. We propose a novel DCT -based [1] bit plane error resilient entropy coding scheme, which can control and minimize the error propagation effect. The compressed rate can be similar with the JPEG standard. In addition, it uses only 18 VLC symbols. Memory requirement and power consumption can then be minimized. Base on simulation results, the proposed coding scheme can achieve high image quality (PSNR = 29.82dB) even at bit error rate of 10-3. These performances can meet various wireless multimedia system requirements.
Yew-San Lee, Cheng-Mou Yu, Chen-Yi Lee
ICASSP3
2002 Error resilient image coding and smart post-processing error concealment for wireless image transmission
abstract
In recent years, there has been a huge growth in the areas of multimedia and wireless communications. Image transmission through bandwidth limited and high bit error rate wireless channel requires both compression and error resilient capabilities. We propose an error resilient image coding and smart error concealment schemes for DCT-based [1] image compression, such as JPEG and MPEG standards. It can successfully prevent error propagation between the transmitted data blocks with little overhead. In addition, we present a novel post-processing error concealment scheme, called Smart-IDCT (SIDCT). It tries to retain error free low frequency DCT information and discarding highly suspicious high frequency information. Then, we can retrieve low-resolution information instead of totally corrupted image block during the decoding process. The required computation power is much less than conventional error detection and correction schemes. Simulation results show that the overhead of ERIC is about 5% compared to the JPEG sequential DCT-based mode without restart marker. With the SIDCT post-processing, it can achieve excellent image quality (PSNR=31.78dB).
Keng-Khai Ong, Yew-San Lee, Chen-Yi Lee
ICASSP3
2002 A high throughput low cost context-based adaptive arithmetic codec for multiple standards
abstract
For next generation image compression standard, context-based arithmetic coding is adopted for improving the compression rate. An efficient and high throughput codec design is strongly required for handling high-resolution images. We propose an efficient codec architecture for context-based adaptive arithmetic coding, which exhibits low cost, low latency, and high throughput rate. In addition, it can be programmed for supporting multiple standards such as JPEG, JPEG2000, JBIG, and JBIG2 standards. It exploits three-pipeline stages architecture. Based on parallel leading zeros detection and bit-stuffing handling, symbols can be encoded and decoded within one cycle. Therefore, the throughput rate can be increased as high as the codec operating clock rate. For 0.35 /spl mu/ 1P4M CMOS technology, both the encoding and decoding rate can run up to 185 M symbol/sec. The AC codec only costs 12 K gate count and 860 /spl mu/m/spl times/860 /spl mu/m layout area. These performances can meet high-resolution real time application requirements.
Keng-Khai Ong, Wei-Hsin Chang, Yi-Chen Tseng, Yew-San Lee, Chen-Yi Lee
ICIP (1)5
2002 A novel fixed bit plane error resilient image coding for wireless multimedia transmission
abstract
The variable length code (VLC) is the most popular technique used in DCT-based image compression standards such as JPEG, MPEG and H.26x. Unfortunately it is highly sensitive to channel noise. For wireless multimedia transmission, any error bit will cause serious error propagation and result in large image quality degradation. Moreover, it usually has no retransmission for real time applications. As a result, a high error resilient image coding scheme is important for wireless applications. We propose a novel DCT-based fixed bit plane error resilient image coding (FBP-ERIC) scheme, which can minimize the error propagation effect with low redundancy. The complexity is much less than conventional coding schemes. In addition, it has a high accurate error detection capability for performing error concealment. Even at a 0.1% bit-error-rate, 46.5% of erroneous blocks can be accurately detected. Hence, the high image quality (PSNR=32.06 dB) can be obtained by applying a simple error compensation mechanism.
Hung-Kuo Wei, Yew-San Lee, Yen-Hsu Shih, Chen-Yi Lee
ICIP (3)4
2002 A JPEG-like texture compression with adaptive quantization for 3D graphics application
Cheng-Hsien Chen, Chen-Yi Lee
Vis. Comput.2
2001 A new approach of group-based VLC codec system with full table programmability
abstract
The algorithm and architecture of a variable length code (VLC) codec system using a new group-based approach and achieving full table programmability are presented. According to the proposed codeword grouping and symbol memory mapping, both group searching and encoding/decoding procedures are completed by applying numerical properties and arithmetic operations to codewords and symbol addresses. By a novel symbol conversion, the memory requirement of the encoding process is reduced and the programmability of codewords and symbols is achieved. For MPEG applications, a 0.6-/spl mu/m CMOS design that performs concurrent VLC codec processes is shown. This VLSI implementation occupies an area of 5.0/spl times/4.5 mm/sup 2/ with 110 k transistors and satisfies a coding table up to 256-entry 12-bit symbols and 16-bit codewords. In addition, both encoding and decoding throughputs of this design achieve 100 Msymbols/s at a 100-MWz clock rate. Therefore, the proposed VLC codec system is suitable for applications which require high operation throughput, such as HDTV and simultaneous compression and decompression, such as videoconferencing.
Bai-Jue Shieh, Yew-San Lee, Chen-Yi Lee
IEEE Trans. Circuits Syst. Video Technol.3
2000 HVLC: error correctable hybrid variable length code for image coding in wireless transmission
abstract
Variable length code (VLC), also called I-iuffman code [I], is the most popular data compression technique used for solving transmission channel bandwidth bottleneck in image compression standards such as JPEG, MPEG, and M263. But, it is vulnerable to loss of synchronization if they are transmitted consecutively through a noisy wireless channel. It will result in large drops in video transmission quality. We propose a novel Hybrid VLC (HVLC) coding scheme to provide high tolerances to random and burst errors in worsening channel conditions. It exhibits high synchronization, error correction capability and low redundancy. For erroneous HVLC bitstream, it is able to self-synchronize within one codeword. The result shows that it achieves high signal-to-noise ratio (PSNK=30dB) compared to existing VLC schemes at bit error rate (BER) of environment. With efficient memory mapping, HVLC requires low memory spaces and achieves high throughput rate. It is very suitable for VLSl implementation in wireless application. 1.
Yew-San Lee, Cheng-Mou Yu, Wei-Shin Chang, Chen-Yi Lee
ICASSP4
2000 Construction of Error Resilient Synchronization Codeword for Variable-Length Code in Image Transmission
abstract
Variable-length code (VLC), also called Huffman (1952) code, is the most popular data compression technique used for image compression and transmission. But, it is vulnerable to loss of synchronization under transmission in a noisy channel. It will result in large drops in video transmission quality. For stopping error propagation, a synchronization codeword (SC) is mostly used to localize the error propagation effect. However, most of the SC cannot tolerate noise errors. We propose an iterative and efficient algorithm to construct error resilient and variable-length SC for any application defined VLC table. They can be used for synchronization of packet (ERVL-SOP). It can be appended at the end of certain number of symbols to form a packet. Thus, channel noise error can be localized. With the use of the constructed SC as the SOP, the result shows that it can achieve high signal-to-noise ratio (PSNR=26 dB) compared to the use of the normal VLC code as SOP at a bit error rate (BER) of 10/sup -3/. In addition, the incorrect ERVL-SOP detection rate can be decreased to less than 0.0013% effectively.
Yew-San Lee, Wei-Shin Chang, Hsin-Han Ho, Chen-Yi Lee
ICIP4
2000 A new approach of group-based VLC codec system
abstract
In this paper, the algorithm of a VLC codec system with new group-based approach is presented. Based on the proposed codeword grouping and symbol memory mapping, the group-searching scheme and codec processes are completed by applying numerical properties and arithmetic operations to codewords and symbol addresses. The memory requirement of encoder is reduced by a novel symbol-converting scheme. Therefore, the programmable coding table and symbol representation can be achieved. Based on MPEG-like systems, an architecture design that performs concurrent VLC codec processes with constant symbol rate is presented. Simulation results show 100 Msps with 100 MHz-clock for both encoding/decoding procedures can be achieved. As a result, it is suitable for those applications that require codec processes simultaneously, such as videoconferencing, and high throughput systems, such as HDTV.
Bai-Jue Shieh, Terng-Yin Hsu, Chen-Yi Lee
ISCAS3
2000 A high-throughput memory-based VLC decoder with codeword boundary prediction
abstract
In this paper, we present a high-throughput memory-based VLC decoder with codeword boundary prediction. The required information for prediction is added to the proposed branch models. Based on an efficient scheme, these branch models and the Huffman tree structure are mapped onto memory modules. Taking the prediction information, the decompression scheme can determine the codeword length before the decoding procedure is completed. Therefore, a parallel-processor architecture can be applied to the VLC decoder to enhance the system performance. With a clock rate of 100 MHz, a dual-processor decoding process can achieve decompression rate up to 72.5 Msymbols/s on the average. Consequently, the proposed VLC decompression scheme meets the requirements of current and advanced multimedia applications.
Bai-Jue Shieh, Yew-San Lee, Chen-Yi Lee
IEEE Trans. Circuits Syst. Video Technol.3
1999 An efficient VLC decompression scheme for user-defined coding tables
abstract
With the increase of information and data types, a high-throughput and flexible memory-based (variable length code) VLC decoder is required for user-defined coding tables to achieve a higher compression ratio. We present a memory-based VLC decoder which is quite suitable for the applications with user-defined tables. By parallel loading data into memories, the coding tables can be changed with much less time. The codeword-boundary prediction algorithm breaks the recursive dependency of the decoding procedures. As a result, the VLC decoder can be realized on a multiprocessor architecture and hence the decoding throughput is enhanced significantly. Additionally, the index-offset symbols that can recover all data with a pure VLC codeword and a smaller table size are presented. Simulation results show that the combination of the proposed VLC decoder and user-defined table can achieve a high decompression rate. As a result, it is quite suitable for high data rate applications with user-defined coding tables, such as MPEG-4.
Bai-Jue Shieh, Chen-Yi Lee
ICASSP2
1999 A Cost-Effective Lighting Processor for 3D Graphics Application
abstract
This paper presents a cost effective VLSI architecture to calculate RGB color in 3D graphics lighting stage. Phong illumination model Ell is exploited for lighting. In addition, the proposed architecture also supports point light sources, spot light sources, and infinite light sources suggested in OpenGL(R)**. These various light models are widely used in entertainment and game applications. We choose appropriate fix-point representation in this architecture, which is different from traditional floating-point geometry processor. Color accuracy is taken into account under fix-point simulation in MESA 3D graphics library. Simulation results show that the proposed architecture can achieve 1.5 M polygons/sec and be integrated with 3D graphics rasterization stage cost-effectively.
Cheng-Hsien Chen, Chen-Yi Lee
ICIP (2)2
1999 An Efficient VLSI Architecture for Separable 2-D Discrete Wavelet Transform
abstract
In this paper, we present a VLSI architecture for separable 2-D Discrete Wavelet Transform (DWT). Based on 1-D DWT recursive pyramid algorithm (RPA), a complete 2-D DWT output scheduling scheme is derived. The I/O between memory which stores the intermediate results and DWT core is simplified by “circular coefficients arrangement”. And the concept to store the “partial accumulation sum” of convolution operation in column direction is first proposed in this paper. For the computations of N×N 2-D DWT with filter length L, our architecture spends N2clock cycles and requires 2NL words in memory size, 4L multipliers, as well as 4L-2 adders. And the number of multipliers and adders can be further reduced to 2L, and 2L-1 respectively by sharing positive and negative clock edge. The architecture is suitable for VLSI implementation and various real-time video/image applications.
Wen-Shiaw Peng, Chen-Yi Lee
ICIP (2)2
1999 A New Anti-Aliasing Algorithm for Computer Graphics Images
abstract
This paper presents an area-filtering algorithm for anti-aliasing technique of computer graphics images. It can be applied to low-resolution image, which the aliasing effect is more visible. Unlike supersampling technique, which costs too much to do well, this technique is simple and suitable for VLSI implementation. Without changing the current PC-based graphics-hardware architecture, this technique costs little overhead and improves the image quality of aliasing images 5~10 dB on the average.
Yuan-Hau Yeh, Chen-Yi Lee
ICIP (2)2
1999 Cost-effective VLSI architectures and buffer size optimization for full-search block matching algorithms
abstract
This paper presents two efficient very large scale integration (VLSI) architectures and buffer size optimization for full-search block matching algorithms. Starting from an overlapped data flow of search area, both systolic- and semisystolic-array architectural solutions are derived. By means of exploiting stream memory banks, not only input/output (I/O) bandwidth can be minimized, but also processor element efficiency can be improved. In addition, the controller structure for both solutions are very straightforward, making them very suitable for VLSI implementation to meet computational requirements. Moreover, by exploring the dependency graph, we focus on the problem of reducing the internal buffer size under minimal I/O bandwidth constraint to derive guidelines on reducing redundant internal buffer as well as to achieve area-efficient VLSI architectures. Simulation results show that, for N=P=16 (N is the reference block size and P is the search range), I/O bandwidth can be reduced by 2.4 times, while buffer size increases less than 38%. Two prototype chips for N=P=16 have been designed and fabricated. Test results show that clock rate can be up to 90 MHz, implying that more than 87.9-K motion vectors per second can be achieved to meet real-time requirements specified in MPEG-2 MP@ML coding standard.
Yuan-Hau Yeh, Chen-Yi Lee
IEEE Trans. Very Large Scale Integr. Syst.2
1998 Entropy-constrained gradient-match vector quantization for image coding
abstract
Side-match VQ (SMVQ) is a well-known class of FSVQ used for low-bit rate image/video coding. It exploits the spatial correlation between the neighboring blocks to select several codewords that are very close to the encoding block from the master codebook. But if the block boundary is in the region edge area, the spatial correlations are not high and the SMVQ can't select proper codewords to encode blocks. An entropy-constrained gradient-match VQ (ECGMVQ) is proposed. Instead of exploiting the spatial correlation, the ECGMVQ uses the gradient contiguity property to select the codewords. The state function of ECGMVQ can select better codevectors than the SMVQ. In addition, the entropy-constrained rule is applied to the encoding process to reduce the bit rate. Simulation results show that the improvement of ECGMVQ over the SMVQ is up to 4-5 dB at nearly the same bit rate. Further, the perceptual image quality is better than that of SMVQ, especially in the region edge area.
Shin-Chou Juan, Chen-Yi Lee
ICASSP2
1997 Buffer size optimization for full-search block matching algorithms
abstract
This paper presents how to find optimized buffer size for VLSI architectures of full-search block matching algorithms. Starting from the DG (dependency graph) analysis, we focus in the problem of reducing the internal buffer size under minimal I/O bandwidth constraint. As a result, a systematic design procedure for buffer optimization is derived to reduce realization cost.
Yuan-Hau Yeh, Chen-Yi Lee
ASAP2
1997 A multicasting solution for ATM video applications
abstract
This paper presents a multicasting solution for shared-buffer ATM switch. The cell duplicating function is performed by a one-to-many modified demux circuit. A shared multicast server is used to translate input multicast cells to find destination ports and corresponding routing information. By using a link-list based ring structure, the multicast server can provide a high speed (5 ns to search entry in 256-entry block in 0.8-/spl mu/m CMOS process) and cascadable multicast translating function. In addition, the channel complexity of multipoint-to multipoint multicast applications can be reduced to N channels per N users. The shared-buffer controller is modified by adding controlling and scheduling functions for multicasting queues to process the multicast feature. Finally with the aid of quality of service (QoS) management, this proposed solution provides multi-QoS for multicast cells to meet the needs of various video applications.
Jer-Min Tsai, Hsin-Hsiung Fang, Chen-Yi Lee
IEEE Trans. Circuits Syst. Video Technol.3
1996 Scalable VLSI architectures for full-search block matching algorithms
abstract
This paper presents two VLSI architectures for full search block matching motion estimation (ME) algorithm based on overlapped search data flow. The proposed VLSI architectures have three specific features: (1) they contain a processor element (PE) array which provides sufficient computational power and achieves 100% hardware efficiency; (2) they contain stream memory banks which provide scheduled data flow requested by PE for computing mean absolute distortion (MAD); and (3) they both have minimum memory bandwidth to save I/O pin-count.
Yuan-Hau Yeh, Chen-Yi Lee
ICIP (2)2
1996 The outage probability in DS/CDMA for cellular mobile radio with imperfect power control
abstract
We establish a system model, including shadowing, multipath fading, antenna diversity and voice active factor, to estimate the system performance with considerations of perfect and imperfect power control by using the outage probability. Based on our calculations, it is shown that the system performance is sensitive to multipath fading, antenna diversity, and the power-control error. From our simulation results, three schemes are derived as follows: (1) the average power control is not powerful enough to maintain a minimum required number of subscribers per cell in a fading environment or in imperfect conditions; (2) antenna diversity can be used to improve the system capacity in all conditions, not only considerations with perfect power control, but also with imperfect power control; and (3) the system capacity is decreased rapidly by multipath fading and the power-control error.
Terng-Yin Hsu, Chen-Yi Lee
PIMRC2
1996 Finite state vector quantization with multipath tree search strategy for image/video coding
abstract
This paper presents a new vector quantization (VQ) algorithm exploiting the features of tree-search as well as finite state VQs for image/video coding. In the tree-search VQ, multiple candidates are identified for ongoing search to optimally determine an index of the minimum distortion. In addition, the desired codebook has been reorganized hierarchically to meet the concept of multipath search of neighboring trees so that picture quality can be improved by 4 dB on the average. In the finite state VQ, adaptation to the state codebooks is added to enhance the hit ratio of the index produced by the tree-search VQ. Thus, compressed bits can be further reduced. An identifier code is then included to indicate to which output indexes belong. Therefore, this modified algorithm not only reaches a higher compression ratio, but also achieves better quality compared to conventional finite-state and tree-search VQs. Finally, suitable VLSI architectures for real-time performance are proposed here (1) to remove the bottleneck of iteration bound in finite-state VQ and (2) to provide parallel computing structure for tree-search VQ to meet computational requirements.
Chen-Yi Lee, Shih-Chou Juan, Yen-Juan Chao
IEEE Trans. Circuits Syst. Video Technol.1
1995 Semi-systolic array based motion estimation processor design
abstract
This paper presents a new VLSI architecture for full-search block matching algorithm. The proposed architecture has two specific features: (1) it has a processor element (PE) array which provides sufficient computational power, where PEs work in a semi-systolic style and (2) it contains stream memory banks which provide scheduled data flow to reduce idle operations within PE array. By exploiting broadcasting and local data communications, hardware efficiency of the proposed architecture can be up to 100%, which outperforms those systolic-array solutions found in the literature.
Mei-Cheng Lu, Chen-Yi Lee
ICASSP2
1995 A New Multi-Path Tree-Search FSVQ Architecture for Image/Video Sequence Coding
abstract
This paper presents a new multi-path tree-search architecture to implement various FSVQs for real-time image/video coding. The proposed architecture exploits the features used in both tree-search and finite state VQs and in the mean time to adaptively update state codebooks so that PSNR is on the average 1 dB/spl sim/2 dB higher than that achieved from conventional FSVQs at the same bit rate. In particular, this architecture can be pipelined in hardware realization to overcome the iteration bounds as those encountered in conventional FSVQs, making it feasible to develop real-time cost-effective hardware solutions.
Yen-Juan Chao, Chen-Yi Lee
ISCAS2
1995 An Efficient Memory Architecture for Motion Estimation Processor Design
abstract
This paper presents a novel memory architecture for motion estimation processor design. By means of conditional selection strategy, data items which can be reused are stored in memory banks and arranged in a snake-like way. Both integer and half pixel motion vectors can be obtained by the proposed architecture and an array processor, where memory bandwidth can be minimized and hence I/O pin-count can be reduced a lot. The proposed architecture is then demonstrated by a test chip, whose hardware efficiency of processor elements is 100% when integer motion vector is demanded.
Eddie G. Tzeng, Chen-Yi Lee
ISCAS2
1994 Finite State Vector Quantization with Multi-Path Tree Search Strategy for Image/Video Coding
abstract
This paper presents a new vector quantization (VQ) algorithm exploiting the features of tree-search as well as finite state VQs for image/video coding. In the tree-search VQ, multiple candidates are identified for on-going search to optimally determine an index of the minimum distortion. In addition, the desired codebook has been reorganized hierarchically to meet the concept of multi-path search of neighboring trees so that picture quality can be improved by 4 dB on the average. In the finite state VQ, adaptation to the state codebooks is added to enhance the hit-ratio of the index produced by the tree-search VQ and hence to further reduce compressed bits. An identifier code is then included to indicate to which output indices belong. Our proposed algorithm not only reaches a higher compression ratio but also achieves better quality compared to conventional finite-state and tree-search VQs.>
Shih-Chou Juan, Yen-Jean Chao, Chen-Yi Lee
ISCAS3
1994 Design of a Fast Sequential Decoding Algorithm Based on Dynamic Searching Strategy
abstract
This paper presents a new sequential decoding algorithm based on dynamic searching strategy to improve decoding efficiency. The searching strategy is to exploit both sorting and path recording techniques. By means of sorting we can identify the correct path in a very fast way and then, by path recording, we can recover the bit sequence without degrading decoding performance. We also develop a conditional resetting scheme to overcome the buffer overflow problem encountered in conventional sequential decoding algorithms. Simulation results show that for a given code, decoding efficiency remains the same as that obtained from maximum likelihood function by appropriately selecting sorting length and decoding depth. In addition, this algorithm can be mapped onto an area-efficient VLSI architecture to implement long constraint length convolutional decoders for high-speed digital communications.>
Wen-Wei Yang, Li-Fu Jeng, Chen-Yi Lee
ISCAS3
1994 High-Throughput Data Compressor Designs Using Content Addressable Memory
abstract
This paper presents a novel VLSI architecture for high-speed data compressor designs which implement the well-known LZ77 algorithm. The architecture mainly consists of three units, namely content addressable memory, match logic, and output stage. The content address memory generates a set of hit signals which identify those positions whose symbols in a specified window are the same as input symbol. These hits signals are then passed to the match logic which determines one matched stream and its match length and location in the window to form the kernel of compressed data. These two items are then passed to the output stage for packetization before sent out. By trading off hardware complexity and compression ratio, 2KB window size and adjustable maximum match length are considered in our proto-type VLSI chip. Simulation results show that, based on a 0.8 /spl mu/m CMOS process technology, clock speed up to 50 MHz can be achieved. This implies that the developing data compressor chip can handle many real-life applications such as in video coding and high-speed data storage systems.>
Ren-Yang Yang, Chen-Yi Lee
ISCAS2
1994 High-speed median filter designs using shiftable content-addressable memory
abstract
This paper presents a very efficient VLSI architecture for real-time median filtering as requested in many image/video applications. The median is obtained by first sorting input sequences and then selecting identified order according to the number of inputs. To reach the goal of high-speed data sorting, an optimized delete-and-insert algorithm is derived and then mapped onto shiftable content-addressable memory architecture. The complete design can be decomposed into a set of processor elements, where each processor element consists of two basic cells-sort-cell and compare-cell. Thus the design becomes very regular. More specifically any specified order can be obtained within one cycle and a high-speed clock rate can be achieved. A prototype chip for 64 samples based on this architecture has been implemented and tested. Results show that a clock rate up to 50 MHz can be achieved using a 1.2 /spl mu/m CMOS double metal technology.>
Chen-Yi Lee, Po-Wen Hsieh, Jer-Min Tsai
IEEE Trans. Circuits Syst. Video Technol.1
1993 An area-efficient maximum/minimum detection circuit for digital and video signal processing
Chen-Yi Lee, Shih-Chou Juan, Wen-Wei Yang
ISCAS1
1993 VLSI implementation of an M-array image filter based on shift register array
Chen-Yi Lee, Jer-Min Tsai, Shih-Chieh Hsu
Integr.1
1991 Breaking the bottleneck of sequential decoding for high-speed digital communication
abstract
An efficient ASIC architecture for the sequential stack decoding (SSD) algorithm used for channel coding is presented. It is different from the maximal likelihood (ML) Viterbi decoder (VD), mainly in the search for the correct memory path. Due to the dedicated memory organization, the storage space and required hardware can be reduced while the decoding efficiency remains almost the same. The proposed architecture results from step by step design of the I/O interface, high-level memory management, dedicated data paths, and controller. The ordering of these steps is important in optimizing the final solution. In addition, the construction of this hardware organization can be made by using the available hardware building blocks.>
Chen-Yi Lee, Francky Catthoor, Hugo De Man
ICASSP1
1990 Efficient VLSI Architectures for a High-Performance Digital Image Communication System
abstract
A typical digital image communication system based on a number of ASIC architectural designs is currently under development. After partitioning the complete system into several stages, each selected algorithm can be implemented on a single ASIC based on an efficient architectural style. The required throughput for high-performance data compression and channel coding has been obtained due to the optimization of the critical path in the architectural design. Input/output (I/O) operations, which create the bottleneck for many image-processing algorithms, are handled by a dedicated I/O interface unit. Construction of each dedicated data path in the architecture is based on a limited parameterizable functional building block (FBB) library. The dedicated data paths have been constructed by partitioning the initial signal flow graph (SFG) into compatible graphs and by matching the graphs in each partition onto a collection of time-multiplexed FBBs. Hierarchically partitioned controllers were used to meet the high-throughput requirements. The ASIC architectures proposed are oriented to broadband integrated services digital networks (B-ISDN) as well as high-performance digital image compression systems.>
Chen-Yi Lee, Francky Catthoor, Hugo De Man
IEEE J. Sel. Areas Commun.1