VLDB 2026 Research / reviewers in the wild / expert
Ching-Te Chiu
dblp:97/4711
· DBLP profile ↗
71ranked-venue papers
7as first author
9since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 3 first-author · 2 since 2021Systems, architecture and hardware · 29 · 4 first-author · 7 since 2021Computer networks · 2Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Multi-Layer Relation Knowledge Distillation For Fingerprint RestorationabstractKnowledge distillation involves a lightweight student model learning from the high-performance teacher model. Traditional feature-based knowledge distillation has limitations as it makes the student model mimic the teacher’s features. This approach lacks flexibility, particularly when there are significant architectural differences between the teacher and student models. In this paper, we introduced a multi-layer relation knowledge distillation (MRKD). MRKD focuses on learning the similarity between input patches of the teacher model at different layers. We also utilize the Attention-based Fusion (ABF) module to learn the optimal ratios between different layers, enabling the acquisition of information from multiple layers. We designed an asymmetric one-encoder-two-decoder lightweight deep learning model to restore fingerprint quality, aiming for sub-0.1-second inference times. Compared to the teacher model, the lightweight model achieves 27x faster inference time and 10x fewer parameters. The lightweight model achieves average 50% improvement in equal error rate(EER) on the FVC2002 and FVC2004 datasets compared to state-of-the-art approaches. Yu-Min Chiu, Ching-Te Chiu, Dao-Heng Luo |
ICASSP | 2 |
| 2024 | Low DRAM Memory Access and Flexible Dataflow Convolutional Neural Network Accelerator based on RISC-V Custom InstructionabstractDeep convolutional neural networks has been widely used in several applications. However, the huge computational complexity and data access times hinder its application in edge devices. Previous works target to design a specific fixed dataflow. However, several researches point out that there are no dataflow that can be optimal across all layers or models. In this paper, first we propose an flexible dataflow accelerator which can reconfigure to weight stationary or output stationary dataflow at every layer to increase hardware utilization and data reuse. Besides, we design RISC-V custom instructions to encode the dataflow configurations. Last, we proposed a on-the-fly pooling method to compute the max pooling layer right after the convolutional layer to reduce off-chip memory access. By reconfiguring the dataflow, we improve 3.75x and 1.18x DRAM access amounts in VGG16 compared with [1], [2] respectively. Besides, we maintain a high utilization rate of 99.12%. The proposed accelerator can not only reconfigure the dataflow of each layer but also achieve high throughput, high area efficiency, and high power efficiency. The accelerator implemented in the 40nm process reaching 256 GOPS throughput with 1000 MHz, 136.9 GOPS/mm2area efficiency with 1.87 mm2area, 1.014 TOPS/W power efficiency with 252.51 mW power. Yu-Jen Chang, Ching-Te Chiu, Ming-Long Huang, Geng-Ming Liang, Chao-Lin Lee, Jenq Kuen Lee, Ping-Yu Hsieh, Wei-Chih Lai |
ISCAS | 3 |
| 2024 | Feature Points based Residual UNet with Nonlinear Decay Rate for Partial Wet Fingerprint Restoration and RecognitionabstractThis paper proposes FPN-ResUNet, which uses feature points-based restoration and nonlinear decay rate residual within a U-shape architecture to restore blurry and wet fingerprints within size constraints effectively. Our approach aims to restore fingerprints effectively while avoiding over-restoration and preserving local matching feature points, which reduces the false rejection rate (FRR). We develop novel residual blocks, which feature a nonlinear decay mechanism that optimizes weight allocation between blocks to enhance feature extraction capabilities. Furthermore, our residual feature points fusion module restores contextual information by fusing matching feature points and features from the previous level.Through comprehensive experiments on real-world data, we achieved a remarkable 9.4% reduction in FRR. Our method outperforms the basic U-Net framework [1] by 50.52% in overall performance. We improved 48.35% over FPDMNet [2] and 58.77% over DenseUNet [3] which is trained with small patches. Moreover, PGT-Net [4] is also used for small areas of wet fingerprints, but our FPNResUNet outperformed it by 72.51%. An-Ting Hsieh, Ching-Te Chiu, Tsai-Chieh Chen, Mao-Hsiu Hsu, Wenyong Long |
ISCAS | 2 |
| 2023 | RGB-D Based Pose-Invariant Face Recognition Via Attention Decomposition ModuleabstractFace recognition has recently achieved remarkable performance with the help of deep learning networks, but there is still a domain gap between frontal and profile face recognition. Generally speaking, we utilize pose-invariant face recognition methods or incorporate additional depth information to handle pose variations. However, huge backbones or multiple models are often used in RGB-D face recognition methods, which makes them hard to be applied in edge devices. In this work, we propose a RGB-D based pose-invariant face recognition model which is light enough to meet the demands of edge devices. First, we use the attention decomposition mechanism to decompose the mixed feature maps into pose- and identity-related features layer by layer. Second, a continuously indexed domain adaptation and a multi-task training framework are applied to our proposed model to decorrelate these two components. Third, we design our proposed model by embedded convolution neural network (eCNN) architecture to reduce parameters and operations. Finally, we evaluate our proposed method on the public KincetFaceDB. The Rank-1 recognition rate of our proposed method reaches 98.23% on KincetFaceDB, which is 0.13% superior to that of [1]. The parameters of our proposed eCNN model is 0.58M, which is 98.47% lower than that of [1]. Ching-Te Chiu, Kuan-Chang Shih |
ICASSP | 2 |
| 2022 | RGBD-based Hardware Friendly Head Pose Estimation System via Convolutional attention moduleabstractHead pose estimation from RGB images without depth information is a challenging task, owing to the loss of spatial information and large head pose variations in the wild. However, most studies adopt deeper convolutional neural network (CNN) models, such as ResNet50, which are limited by the enormous number of parameters to be implemented on edge devices. Owing to novel technological advancements, several edge devices have also included depth cameras and obtain high-quality images. In this study, we propose a lightweight CNN for head pose estimation. By adopting attention module and feature decoupler, we resume the performance decreasing by lower parameters. Moreover, we classify the ground-truth head pose angles of the model intermittently, and adopt the multi-loss strategy to train our model. We evaluate the proposed method on three challenging benchmark datasets, and achieved optimal results for Yaw pose and average. The obtained results indicate that although the proposed model has less parameters, it still maintains a remarkable performance. The total number of parameters is 0.19 M, including RGB and depth path, which is 50% lower than FSA-Net. Consequently, the inference speed is 0.92 ms per pair RGB-D images, which is 8% faster than FSA-Net. With fewer parameters, we achieved 3.1 MAE on yaw angle, which is 22.69% lower than that of Quatnet, including 3.5 MAE on average, which is 7.40% lower than those of other advanced methods. Yen-Yu Cheng, Ching-Te Chiu |
ISCAS | 2 |
| 2022 | Multi-View RGB-D Based 3D Point Cloud Face Model Reconstruction SystemabstractThis paper proposes a novel multi-view 3D point cloud face model reconstruction system which can well capture local asymmetry features to deal with the occlusion problem often happening in single-view-based approaches. With the help of asymmetries between facial landmarks to register 3D point clouds, this system can significantly improve the reconstruction rate even though the offsets or transformations between head poses are large. Moreover, its accuracy can be further improved via a face segmentation method to exclude non-facial elements and filter out facial outlier. Experimental results show the average distance between could points is 0.17mm which outperforms than other SoTA techniques, and 98.2 percent accuracy on face recognition is achieved. Jie-Yu Luo, Ching-Te Chiu, An-Ting Hsieh |
ISCAS | 2 |
| 2022 | Chaos LiDAR Based RGB-D Face Classification System With Embedded CNN Accelerator on FPGAsabstractFace classification is important in many applications such as surveillance, border control, and security systems. However, wide variations in environments such as insufficient light, large distances or pose angles make the task challenging. Depth sensors are added with RGB cameras for improving classification accuracy but commercial RGB-D sensors are most targeted for indoors applications. In this paper, we present and design a Chaos LiDAR depth sesnor that provides high-precision depth images through intelligent correlation processing for both indoors and outdoors applications. Our Chaos LiDAR depth sensor detects range from 2 to 40 meters with precision around 8mm at 20-meter. With the Chaos LiDAR depth as input, we design a RGB-D based face classification embedded CNN (eCNN) model for wide range applications such as dim illumination, various distances and large poses. Our Chaos LiDAR increases around 14.27% classification accuracy compared to RealSense D435i for distance from 3 to 5 meter. The eCNN face classification subsystem is implemented in Xilinx ZCU 102 and achieves 11.11 ms inference time. The eCNN engine achieves a peak throughput at 614.4 GOPS. The overall system including Chaos LiDAR, correlation and eCNN FPGA achieves face classification inference rate of 10fps. Ching-Te Chiu, Yu-Chun Ding, Wei-Jyun Chen, Shu-Yun Wu, Chao-Tsung Huang, Chun-Yeh Lin, Chia-Yu Chang, Meng-Jui Lee, Shimazu Tatsunori, Tsung Chen, Fan-Yi Lin, Yuan-Hao Huang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | High Utilization Energy-Aware Real-Time Inference Deep Convolutional Neural Network AcceleratorabstractDeep convolution Neural Network (DCNN) has been widely used in computer vision tasks. However, for edge device, even then inference has too large computational complexity and data access amount. Due to the mentioned shortcomings, the inference latency of state-of-the-art models are still impractical for real-world applications. In this paper, we proposed a high utilization energy-aware real-time inference deep convolutional neural network accelerator, which outperforms the current accelerators. First, we use 1x1 size convolution kernels as the smallest unit of the computing unit. And we design suitable computing unit for different models based on the requirement of each model. Second, we use Reuse Feature SRAM to store the output of current layer in the chip and use as the input of the next layer. Moreover, we import Output Reuse Strategy and Ring Stream Data flow not only to expand the reuse rate of data in the chip but to reduce the amount of data exchange between chips and DRAM. Finally, we present On-fly Pooling Module to let the calculation of the Pooling layer to be completed directly in the chip. With the aid of the proposed method in this paper, the implemented CNN acceleration chip has extreme high hardware utilization rate. We reduce a generous amount of data transfer on the specific module, ECNN [1]. Compared to the methods without reuse strategy, we can reduce 533 times of data access amount. At the same time, we have enough computing power to perform real-time execution of the existing image classification model, VGG16 [2] and MobileNet [3]. Compared with the design in [4], we can speed up 7.52 times and have 1.92x energy efficiency. Kuan-Ting Lin, Ching-Te Chiu, Jheng-Yi Chang, Shan-Chien Hsiao |
ISCAS | 2 |
| 2021 | Real-Time Block-Based Embedded CNN for Gesture Classification on an FPGAabstractThis paper presents a block-based embedded convolutional neural network (CNN) for gesture classification on field-programmable gate array (FPGA) in real time. Gesture recognition is an important tool to spontaneous interact with human machine interface. Many CNN architectures using RGB images have been proposed for gesture classification. RGB based gesture classification may cause incorrect results under insufficient light or similar gestures. In addition, most of the CNN architectures cannot run in real time on edge devices due to their large number of parameters and DRAM data access. In this paper, a block-based CNN using RGB-D data is proposed for gesture classification. Adding depth images to RGB images boots the classification accuracy. A CNN architecture with block-based feature maps is built for embedded FPGA implementations. The total number of parameters of the proposed RGB-D embedded CNN (eCNN) model is only 0.17M and it achieves 99.96% and 99.88% accuracy with 32-bit floating point and 8-bit fixed point implementation for America Sign Language (ASL) data set. The RTL simulation of the proposed eCNN model has the average inference speed of 0.171 milliseconds at frequency of 250MHz for a single pair RGB-D image. Implemented on a FPGA integrated with Microsoft Kinect v2 achieve an inference time in 19.42 ms which achieves high accuracy and real-time performance. Ching-Chen Wang, Yu-Chun Ding, Ching-Te Chiu, Chao-Tsung Huang, Yen-Yu Cheng, Shih-Yi Sun, Chih-Han Cheng, Hsueh-Kai Kuo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2020 | Fast Single-View 3D Object Reconstruction with Fine Details Through Dilated Downsample and Multi-Path Upsample Deep Neural NetworkabstractThree-dimensional (3D) object reconstruction is among the most important research areas in the field of computer vision. Its purpose is to reconstruct the overall shape of an object from its twodimensional (2D) image. With the development of deep learning, many methods based on convolutional neural networks (CNNs) have been applied in related research.To achieve 3D shape reconstruction with low computation time, we focus on the commonly used method: single-image reconstruction. The main issue of using a single image as an input is that the reconstruction shape often lacks structural detail. To address this issue, we proposed two methods: the dilated downsample block and the multi-path upsample block. The dilated downsample block extracts more features and the multi-path upsample block uses the features in our architecture. Thereafter, we concatenate the encoder and decoder with corresponding layers to keep the image features in reconstruction process.Finally, we perform experiments on the dataset provided by Choy et al. Results show that our method achieves 67.7% intersection over-union (IoU) accuracy, 3.6% higher than state-of-the-art method, VTN. Compared to the PSVH method, our result achieves 71.4%, an increase of 3.4%. Our average reconstruction time is 13 ms, approximately 25 times faster than PSVH. Chia-Ho Hsu, Ching-Te Chiu, Chia-Yu Kuan |
ICASSP | 2 |
| 2020 | Object Detection with Color and Depth Images with Multi-Reduced Region Proposal Network and Multi-PoolingabstractObject detection technology has received increasing research attention with recent developments in automation technology. Most studies in this field, however, use RGB images as input to deep-learning classifiers, and they rarely use depth information.So, in this paper, we use images with both RGB and depth information as input to an object detection network. We base our network on the Faster R-CNN proposed by Shih et al., and we develop a fast and accurate object detection architecture. In addition to adding depth as input, we also adjust the type of anchor boxes to improve performance on some objects. We also discuss the impact of pooling training data with multiple region proposal networks (RPN) and regions of interest (ROI).Adding depth information improved the mAP by 8.15%, from 36.86% to 45.01%, when using the SUN RGB-D dataset with 10 classes. Optimizing the anchor boxes improved the mAP from 45.01% to 45.88%. After testing various architectures with different reduced RPNs, we find that the model of 1RRPN-2ROIP performs best. The running time is 0.123 s, which is 1.8 times faster than the 3D-SSD model. Jiou-Ai Lin, Ching-Te Chiu, Yen-Yu Cheng |
ICASSP | 2 |
| 2020 | Rgb-D Based Multi-Modal Deep Learning for Face IdentificationabstractIn recent years, the rapid development of depth cameras and wide application scenarios. The depth image information becomes more influential in face identification. In the proposed architecture, we implement the networks in dual CNN paths for color and depth images separately. Moreover, we design innovative loss functions to strengthen the discrimination and the complementary features between color and depth modalities. To preserve the strengthened color and depth features, we fuse both features by concatenation before classification. The experimental results show that our multi-modal learning method achieve 4.3381% EER, 0.27 FMR1000, and 0.33 ZeroFMR on IIIT-D Kinect RGB-D Face dataset for face verification and 99.7% classification accuracy, which exceeds the most state-of-the-art methods. Moreover, the global descriptors of model output are designed to be binarized. Our method requires less memory and computation time. Tzu-Ying Lin, Ching-Te Chiu, Ching-Tung Tang |
ICASSP | 2 |
| 2020 | Depth Estimation From Single Image Through Multi-Path-Multi-Rate Diverse Feature ExtractorabstractConvolutional neural networks can effectively learn features and predict the depth by considering different scene types. However, previous studies have not accurately predicted the depth in cases wherein the objects or scenes were small and the background was complex. These studies have used the bilinear up-sampling method to enlarge the feature maps during training, or to disable the transfer of multiscale information to the end of the network. However, this has resulted in blurred regions in the depth maps and contour loss.This paper proposes a multi-path-multi-rate feature extractor, which can effectively extract multi-scale information to make accurate depth predictions. We used the U-NET [1] architecture to obtain depth maps with high resolution, and also used the proposed multi-path-multi-rate feature extractor to translate useful features from the encoder to the decoder. Dilated convolutions with different rates can provide different types of field-of-view information, which increases the precision of depth estimation and maintains the object contours. Finally, we conducted experiments using an indoor scene (NYUv2 [2]). The results show that the proposed framework achieved an improvement of 12.9% in RMSE, 9.9% in REL, and 9.3% in log10, and it requires approximately 0.048 seconds to predict a depth map from a single image. Wen-Yi Lo, Ching-Te Chiu, Jie-Yu Luo |
ICASSP | 2 |
| 2020 | Fast and Accurate Embedded DCNN for Rgb-D Based Sign Language RecognitionabstractIn this paper, fast and accurate two paths CNN architecture was designed in hardware-oriented manner. Our proposed network is composed of RGB and depth path for gesture recognition by fusing RGB and depth features, following the pre-defined constraints on dedicated hardware. The RTL simulation results indicate it only takes 0.171 milliseconds to infer a single pair of RGB image and depth maps at the operational frequency of 250MHz. Compared with running the same model at Intel i7 and GTX 1080, the speedups are 593.92x and 7.68x respectively. Besides, to increase the recognition accuracy under the diversity of the circumstance, a new RGB-D dataset, captured from Kinect, with complex background was built. Moreover, the number of parameters in our model is only 0.17M and it achieves 99.79% accuracy on the ASL Finger Spelling dataset. Compared with the Gao's CNN gesture recognition architecture, the number of parameter of our model is 2.9 times less and the accuracy is 6.49% higher. Demonstration video for sign language recognition is provided : https://youtu.be/DvO8mI7IZ5Q. Ching-Chen Wang, Ching-Te Chiu, Chao-Tsung Huang, Yu-Chun Ding, Li-Wei Wang 0013 |
ICASSP | 2 |
| 2020 | ESSA: An energy-Aware bit-Serial streaming deep convolutional neural network acceleratorabstractOver the past decade, deep convolutional neural networks (CNN) have been widely embraced in various visual recognition applications owing to their extraordinary accuracy. However, their high computational complexity and excessive data storage present two challenges when designing CNN hardware. In this paper, we propose an energy-aware bit-serial streaming deep CNN accelerator to tackle these challenges. Using ring streaming dataflow and the output reuse strategy to decrease data access, the amount of external DRAM access for the convolutional layers is reduced by 357.26x when compared with that of no output reuse case on AlexNet. We optimize the hardware utilization and avoid unnecessary computations using the loop tiling technique and by mapping the strides of the convolutional layers to unit-ones for computational performance enhancement. In addition, the bit-serial processing element (PE) is designed to use fewer bits in weights, which can reduce both the amount of computation and external memory access. We evaluate our design using the well-known roofline model. The design space is explored to find the solution with the best computational performance and communication to computation (CTC) ratio. We can reach 1.36x speed and reduce energy consumption by 41% for external memory access compared with the design in [1]. The hardware implementation for our PE Array architecture design can reach an operating frequency of 119 MHz and consumes 68 k gates with a power consumption of 10.08 mW using TSMC 90-nm technology. Compared to the 15.4 MB external memory access for Eyeriss [2] on the convolutional layers of AlexNet, our method only requires 4.36 MB of external memory access to dramatically reduce the costliest portion of power consumption. Lien-Chih Hsu, Ching-Te Chiu, Kuan-Ting Lin, Hsing-Huan Chou, Yen-Yu Pu |
J. Syst. Archit. | 2 |
| 2020 | Multi-teacher knowledge distillation for compressed video action recognition based on deep learning
Meng-Chieh Wu, Ching-Te Chiu |
J. Syst. Archit. | 2 |
| 2020 | Real-Time Object Detection With Reduced Region Proposal Network via Multi-Feature ConcatenationabstractIn recent years, object detection became more and more important following the successful results from studies in deep learning. Two types of neural network architectures are used for object detection: one-stage and two-stage. In this paper, we analyze a widely used two-stage architecture called Faster R-CNN to improve the inference time and achieve real-time object detection without compromising on accuracy. To increase the computation efficiency, pruning is first adopted to reduce the weights in convolutional and fully connected (FC) layers. However, this reduces the accuracy of detection. To address this loss in accuracy, we propose a reduced region proposal network (RRPN) with dilated convolution and concatenation of multi-scale features. In the assisted multi-feature concatenation, we propose the intra-layer concatenation and proposal refinement to efficiently integrate the feature maps from different convolutional layers; this is then provided as an input to the RRPN. Using the proposed method, the network can find object bounding boxes more accurately, thus compensating for the loss arising from compression. Finally, we test the proposed architecture using ZF-Net and VGG16 as a backbone network on the image sets in PASCAL VOC 2007 or VOC 2012. The results show that we can compress the parameters of the ZF-Net-based network by 81.2% and save 66% of computation. The parameters of VGG16-based network are compressed by 73% and save 77% of computation. Consequently, the inference speed is improved from 27 to 40 frames/s for ZF-Net and 9 to 27 frames/s for VGG16. Despite significant compression rates, the accuracy of ZF-Net is increased from 2.2% to 60.2% mean average precision (mAP) and that of VGG16 is increased from 2.6% to 69.1% mAP. Kuan-Hung Shih, Ching-Te Chiu, Jiou-Ai Lin, Yen-Yu Bu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Real-time Object Detection via Pruning and a Concatenated Multi-feature Assisted Region Proposal NetworkabstractObject detection is an important research area in the field of computer vision. Its purpose is to find all objects in an image and recognize the class of each object. Since the development of deep learning, an increasing number of studies have applied deep learning in object detection and have achieved successful results. For object detection, there are two types of network architectures: one-stage and two-stage. This study is based on the widely-used two-stage architecture, called Faster R-CNN, and our goal is to improve the inference time to achieve real-time speed without losing accuracy.First, we use pruning to reduce the number of parameters and the amount of computation, which is expected to reduce accuracy as a result. Therefore, we propose a multi-feature assisted region proposal network composed of assisted multi-feature concatenation and a reduced region proposal network to improve accuracy. Assisted multi-feature concatenation combines feature maps from different convolutional layers as inputs for a reduced region proposal network. With our proposed method, the network can find regions of interest (ROIs) more accurately. Thus, it compensates for loss of accuracy due to pruning. Finally, we use ZF-Net and VGG16 as backbones, and test the network on the PASCAL VOC 2007 dataset. Kuan-Hung Shih, Ching-Te Chiu, Yen-Yu Pu |
ICASSP | 2 |
| 2019 | Multi-teacher Knowledge Distillation for Compressed Video Action Recognition on Deep Neural NetworksabstractRecently, convolutional neural networks (CNNs) have seen great progress in classifying images. Action recognition is different from still image classification; video data contains temporal information that plays an important role in video understanding. Currently, most CNN-based approaches for action recognition have excessive computational costs, with an explosion of parameters and computation time. The currently most efficient method trains a deep network directly on compressed video containing the motion information. However, this method has a large number of parameters. We propose a multi-teacher knowledge distillation framework for compressed video action recognition to compress this model. With this framework, the model is compressed by transferring the knowledge from multiple teachers to a single small student model. With multi-teacher knowledge distillation, students learn better than with single-teacher knowledge distillation. Experiments show that we can reach a 2.4× compression rate in a number of parameters and a 1.2× computation reduction with 1.79% loss of accuracy on the UCF-101 dataset and 0.35% loss of accuracy on the HMDB51 dataset. Meng-Chieh Wu, Ching-Te Chiu, Kun-Hsuan Wu |
ICASSP | 2 |
| 2019 | An Energy-Aware Bit-Serial Streaming Deep Convolutional Neural Network AcceleratorabstractOver the past decade, deep convolutional neural networks (CNN) have been widely embraced in various visual recognition applications owing to their extraordinary accuracy. However, their high computational complexity and excessive data storage present two challenges when designing CNN hardware. We propose an energy-aware bit-serial streaming deep CNN accelerator to tackle these challenges. With several optimization methods and the proposed ring streaming dataflow, the computational performance is improved and the external memory access is reduced. We evaluate our design with the roofline model by exploring the design space to find the best performance and computation and communication (CTC) ratio solution. We can reach 1.36x speed up and reduce energy consumption by 41% for external memory access compared with the design in [1]. The hardware implementation for our PE Array architecture design can reach the operating frequency of 119 MHz and consumes 68 k gates with a power consumption of 10.08mW using TSMC 90-nm technology. Lien-Chih Hsu, Ching-Te Chiu, Kuan-Ting Lin |
ICIP | 2 |
| 2019 | Filter-based deep-compression with global average pooling for convolutional networks
Ting-Yun Hsiao, Yung-Chang Chang, Hsin-Hung Chou, Ching-Te Chiu |
J. Syst. Archit. | 4 |
| 2018 | PVDC: A Binary Descriptor Using Pore-Valley Disk Code Structure for High-Resolution Partial Fingerprint RecognitionabstractPlenty of pores can solve the lack of feature points problem on high-resolution partial fingerprints. Pore-based features are similar, so the neighbor ridge features are also taken into account. However, lots of feature points cause heavy computation required. We propose a binary descriptor, Pore-Valley Disk Code (PVDC), which encodes the local structure of a center pore and its neighbor valleys with an eight-section disk. The proposed descriptor is rotational invariant since the first section always aligns to the center pore orientation. Instead of recording pixel by pixel, we find that the direction and distance of intersected valleys in each section can efficiently represent the ridge structure with reduced computation. With the proposed fixed-length binary code, the matching time can be significantly reduced. The proposed method has 160x speedup compared with the state-of-the-art pore-based Sparse Representation based Direct Pore (SRDP) method with reasonable EER in HRF DBI database. Pei-Yin Chou, Ching-Te Chiu |
ICASSP | 2 |
| 2018 | Resource Efficient Hardware Implementation for Real-Time Traffic Sign RecognitionabstractTraffic sign recognition (TSR) is one of the Advanced Driver Assistance System (ADAS) device in modern cars. We propose a high efficiency hardware implementation for TSR, which is divided into two stages. In the detection stage, we use Normalized RGB color transform and Single-Pass Connected Component Labeling (CCL) to find the potential traffic signs. In the recognition stage, the Histogram of Oriented Gradient (HOG) is used to generate the descriptor of the signs, and we classify the signs with the Support Vector Machine (SVM). The proposed method achieves 96.61% detection rate and 90.85% recognition rate while testing with the GTSDB dataset. Our hardware implementation reduces the storage of CCL and simplifies the HOG computation. By using TSMC 90nm technology, the proposed design operates at 105 MHz clock rate and processes in 135 fps with the image size of 1360 × 800. The chip size is about lmm2and the power consumption is close to 8mW. Therefore, this work is resource efficient and achieves real-time requirement. Huai-Mao Weng, Ching-Te Chiu |
ICASSP | 2 |
| 2017 | LBP edge-mapped descriptor using MGM interest points for face recognitionabstractIn recent years, face recognition has become a popular topic in academia and industry. Current local methods such as the local binary pattern (LBP), and scale invariant feature transform (SIFT) perform better than holistic methods, but their high complexity levels limit their application. In addition, SIFT-based schemes are sensitive to illumination variation. We propose an LBP edge-mapped descriptor that uses maxima of gradient magnitude (MGM) points. It can completely illustrate facial contours and has low computational complexity. Under variable lighting, experimental results show that our proposed method has a 16.5% higher recognition rate and requires 9.06 times less execution time than SIFT in the FERET database subset fc. In addition, when applied to the Extended Yale Face Database B, our method outperformed SIFT-based approaches as well as saving about 70.9% in execution time. Furthermore, in uncontrolled conditions, our method has a 0.82% higher recognition rate than local derivative pattern histogram sequences (LDPHS) in the Unconstrained Facial Images (UFI) database. Jou Lin, Ching-Te Chiu |
ICASSP | 2 |
| 2017 | Motion clustering with hybrid-sample-based foreground segmentation for moving camerasabstractForeground segmentation/background subtraction is a vital step in many high-level video analysis applications. While many methods have been proposed for foreground segmentation, most assume the cameras to be stationary. With this assumption, they are unable to handle the movements caused by camera rotation. In this paper, we propose a robust hybrid-sample-based foreground segmentation method for moving cameras, and especially for pan-tilt-zoom cameras. First, we propose the use of motion clustering registration to reduce the impact of registration errors. Next, we propose a frame-level reinitialization scheme to solve the problem of sudden large movement between consecutive frames. Third, we adopt a hybrid-sample-based background modeling technique to easily detect camouflaged foreground objects. Lastly, in order to deal with dynamic backgrounds, we propose moving scene pixel-level feedback schemes to dynamically and locally control the sensitivity and adaptation speed of the background model. We evaluate the proposed method using the ChangeDetection.NET 2014 dataset. Experimental results show that our proposed motion clustering registration can eliminate most of the noise caused by registration errors. The proposed reinitialization scheme can handle the noises caused by sudden large movements. The proposed method performs at least 8% better than other state-of-the-art algorithms in terms of the F-score in the pan-tilt-zoom camera scenario, and it also achieves the highest F-score in camera jitter scenarios. Yi-Chan Wu, Ching-Te Chiu |
ICASSP | 2 |
| 2017 | Single image super-resolution using hybrid patch search and local self-similarityabstractIn this paper, we proposed a hybrid patch search process, which combines the gradient and low frequency (LF)-based patch search to further enhance the effects of the above mentioned methods. We use the assumption of local self-similarity to limit the search area within a small window, while obtaining similar results in most cases. In the proposed framework, two different patch search methods are applied. For edge regions, we use the gradient-based patch search, whereas in smooth regions, LF-based patch search is adopted. When the difference is close between two patches of the hybrid patch search, we further compare the gradient direction for verification. In the experimental results, compared with the SR method that only use LF-based patch search and the SR method that gradient-based patch search only, our proposed method gains higher PSNR and SSIM average values. Also, the computation for high frequency (HF) reconstruction is reduced by about half compared with the gradient-based SR method. Shen-Li Lo, Ching-Te Chiu |
ISCAS | 2 |
| 2017 | Low-complexity face recognition using contour-based binary descriptorabstractFace recognition has become a popular topic due to its applications in security, surveillance and so on. Current local methods such as the local binary pattern (LBP) or local derivative pattern (LDP) perform better than holistic methods since they are more stable on local changes such as misalignment, expression or occlusion, but their high computational complexity limit their applications. While LBP is a good feature method, the scale invariant feature transform (SIFT) is widely accepted as one of the best features to capture edge or local shape information. However, SIFT‐based schemes are sensitive to illumination variation. Thus, the authors propose an LBP edge‐mapped descriptor that uses maxima of gradient magnitude points. It accurately illustrates facial contours and has low computational complexity. Under variable lighting, experimental results show that the authors' method has a 16.5% higher recognition rate and requires 9.06 times less execution time than SIFT under FERET fc. Besides, when applied to the Extended Yale Face Database B, the authors' method outperformed SIFT‐based approaches as well as saving about 70.9% in execution time. In uncontrolled conditions, their method has a 0.82% higher recognition rate than LDP histogram sequences in the Unconstrained Facial Images database. Jou Lin, Ching-Te Chiu |
IET Image Process. | 2 |
| 2016 | A cost-effective minutiae disk code for fingerprint recognition and its implementationabstractFingerprint is one of the unique biometric features for the application of identity security. Minutiae cylinder code (MCC) constructs a cylinder for each minutia to record the contribution of the neighbor minutiae, which has great performance on fingerprint recognition. However, the computation time of the MCC is high. Therefore, we proposed a new disk structure to encode the local structure for each minutia. The proposed minutiae disk code (MDC) clearly illustrates the distribution of the neighbor minutiae and encodes the neighbor minutiae more efficiently by having 280.08× speed faster than the MCC encoding part on Matlab platform. The proposed MDC approach has 96.81% recognition rate on FVC2000 and FVC2002 datasets. The hardware implementation can achieve the operating frequency at 111MHz, which can process 1234 fingerprint images per second with the image size of 255 χ 255 and the maximum of 64 minutiae, under TSMC 90nm CMOS technology. The hardware implementation has 141.27× speed faster than the MCC method. Tsai-Te Chu, Ching-Te Chiu |
ICASSP | 2 |
| 2016 | Fingerprint recognition with ridge features and minutiae on distortionabstractIn this paper we present a new fingerprint matching method which combines different features, including minutiae and ridge features. The ridge features contain ridge count, ridge length, ridge curvature direction and ridge frequency. All ridge features are extracted in blocks around each minutia. The similarity scores of the features are summed with different weighting values as the final score of two fingerprints. Experiments are conducted on the FVC2002 database to compare the proposed method with other fingerprint methods on equal error rate (EER). The proposed method achieves better performance than other methods. The average EER value of the proposed method is 0.82 whereas the average EER value of the conventional matching method is 8.12. Chu-Chiao Liao, Ching-Te Chiu |
ICASSP | 2 |
| 2016 | Binary descriptor based SIFT and hardware implementationabstractScale-Invariant Feature Transform (SIFT) [1] has lately attracted attention in computer vision as a robust feature point detection algorithm which is invariant for scale, rotation and illumination change. However, its computational complexity is too high to apply on practical real-time applications. The iterated Gaussian blurred operations on images lead to long computational latency and high memory requirement. In addition, the gradient histogram based descriptor needs lots of calculation and costs about 60% of total calculation time on generating the descriptor. We propose a binary-based descriptor that uses the intensity difference between neighboring pixels. The binary descriptor has the advantage of lower computational complexity and memory usage. For the use of the classifying the object, the binary descriptor also have the acceptable accuracy compared with the traditional histogram based descriptor. The proposed binary descriptor has 7.07x speed up than the original SIFT descriptor. In addition, it has 4.73x speed up compared with original SIFT algorithm [1]. The hardware implementation of the SIFT algorithm with the proposed binary descriptor applies parallel computing on the stage of feature location to accelerate the computing time on detecting feature points. The final implementation uses about 493-K gate count with 90-nm CMOS technology, and offers 7600 feature points/frame for 1080p images at 30 frames/s at the clock rate of 100 MHz. Che-Yu Wu, Ching-Te Chiu, Yarsun Hsu |
ISCAS | 2 |
| 2015 | Local binary pattern orientation based face recognitionabstractScale-invariant feature transform (SIFT) is a feature point based method using the orientation descriptor for pattern recognition. It is robust under the variation of scale and rotation changes, but the computation cost increases with its feature points. Local binary pattern (LBP) is a pixel based texture extraction method that achieves high face recognition rate with low computation time. We propose a new descriptor that combines the LBP texture and SIFT orientation information to improve the recognition rate using limited number of interest points. By adding the LBP texture information, we could reduce the SIFT orientation number in the descriptor by half. Therefore, we could reduce the computation time while keeping the recognition rate. In addition, we propose a matching method to reserve the effective matching pairs and calculate the similarity between two images. By combining these two methods, we can extract different face details effectively and further reduce computational cost. We also propose an approach using the region of interest (ROI) to remove the useless interest points for saving our computation time and maintaining the recognition rate. Experimental results demonstrate that our proposed LBP orientation descriptor can reduce around 30% computation time compared with the original SIFT descriptor while maintaining the recognition rate in FERET database. Adding the ROI at our proposed LBP orientation descriptor can reduce around 58% computation time compared with the original SIFT descriptor in FERET database. For extended YaleB database, our method has 1.2% higher recognition rate than original SIFT method and reduces 28.6% computational time. The experimental results with adding ROI reduces 61.9% computation time for YaleB database. Yi-Kang Shen, Ching-Te Chiu |
ICASSP | 2 |
| 2015 | Accelerating AdaBoost algorithm using GPU for multi-object recognitionabstractTraditionally, an adaptive boosting (AdaBoost) algorithm is used for object recognition because of its prevalent usage and well-trained results. However, because the computation of AdaBoost is extremely time-consuming, it is difficult to guarantee that the computations reflect the latest information in real time. To speed-up the operation, the original AdaBoost algorithm was accelerated with a graphics processing unit (GPU). In this study, Compute Unified Device Architecture (CUDA) was used to accelerate two parts of the AdaBoost algorithm, including feature extraction and training, by applying various strategies to system components such as how the data is put in the memory, amount of CUDA streams, trunk size, and block size. In Feature Extraction of the car datasets, the most time-consuming step feature-value computation is 47.18 times faster than the CPU version. For AdaBoost Training, the total execution is accelerated by 34.23 times. Pin Yi Tsai, Yarsun Hsu, Ching-Te Chiu, Tsai-Te Chu |
ISCAS | 3 |
| 2015 | Pseudo-Multiple-Exposure-Based Tone Fusion With Local Region AdjustmentabstractNew generations of display technologies provide a significantly improved dynamic range compared to conventional display devices. Inverse tone mapping methods have been proposed to convert low dynamic range (LDR) images to HDR ones, and several of them require multiple exposure LDR images of the same scene as inputs. However, the vast majority of LDR images and videos available have only one single exposure. In this paper, we propose a region-based enhancement of the pseudo-exposures to generate an HDR image. First, we present an exposure dependent curve to convert one LDR image to the pseudo-multiple-exposures. Only certain regions of the pseudo-exposures contain noticeable detail information. We propose a region-based enhancement on the pseudo-exposures to boost details in the most distinct region. Thereby the region-enhanced pseudo-exposures are fused into an HDR image. The fused image thus enhances details in the bright region of the dark image and the dark region of the bright image. Compared with other inverse tone mapped methods, our method generates lower total contrast error measured under the dynamic range independent image quality assessment method in [1]. Tsun-Hsien Wang, Cheng-Wen Chiu, Wei-Chen Wu, Jen-Wen Wang, Ching-Te Chiu, Jing-Jia Liou |
IEEE Trans. Multim. | 6 |
| 2014 | Boosted multi-class object detection with parallel hardware implementation for real-time applicationsabstractReal-time multi-class object detection becomes popular for various applications such as vehicle vision systems, computer vision and image processing. Boosted cascades achieve fast and reliable object detection for one object class, but require parallel usage of multiple cascades for multi-class detection. The multi-class capable cascade splits the root-cascade into sub-cascades iteratively until each sub-cascade contains one class. That requires a huge number of classifiers in the generated hierarchy of interlinked cascades. In this paper, we propose a boosted multi-class object cascade that only splits one class object from the upper-level-cascade when building the sub-cascades. Since only once class object is split so we can reduce the number of classifiers in each stage. From the simulation results, the boosted multi-class object detection can reduce 46% weak classifiers compared to the multi-class capable cascade for the MIT CBCL database. The proposed method achieves high detection rate(95.54%) and low false positive rate(1.94%). We implement our proposed algorithm with a parallel architecture to accelerate the detection operation using TSMC 90nm CMOS technology. The implementation results show that the design achieves an operation frequency of 100MHz of processing images of 30 fps with size 160 × 120. Yao-Tsung Yang, Ching-Te Chiu |
ICASSP | 2 |
| 2013 | Depth-based posture recognition by radar and vision fusion for real-time applicationsabstractA radar sensor can capture the distance and angle of an object. Mapping the radar distance and angle information to the coordinates of a video frame accelerates the speed of object identification. The distance information is used to calibrate the size of an object to help the recognition. To achieve real-time performance, we use only five center of gravity points (COG) and four feature sets. Two feature sets measure the displacement of the upper and lower body COG in the vertical and horizontal directions. The other two feature sets quantize the upper and lower body angular change rate. The simulation results show that our proposed approach achieve 98.02% to 80.20% recognition rates for various postures and actions in the KTH and ISIR databases. I-Cheng Tsai, Ching-Te Chiu |
ICASSP | 2 |
| 2013 | Edge curve scaling and smoothing with cubic spline interpolationabstractImage scaling is a widely used method for many applications and numerous approaches have been proposed to this issue. Current approaches that bring promising results while the edge curves of the scaled up image still have blurring effect. This paper focuses on the edge curve of an image. We propose a simple method for edge curve scaling to avoid the disconnect and zigzag problems when using cubic spline interpolation. The scaling results of our method can avoid blurring of the edge curve and maintain the edge contour of an input image. Wei-Chen Wu, Tsun-Hsien Wang, Ching-Te Chiu |
ICIP | 3 |
| 2013 | Embedded Transition Inversion Coding With Low Switching Activity for Serial LinksabstractSerial link interconnection has been proposed for its advantages of reducing crosstalk and area. However, serializing parallel buses tends to increase bit transition and power dissipation. Several coding schemes, such as serial followed by encoding (SE) and transition inversion coding (TIC), have been proposed to reduce bit transition. TIC is capable of decreasing transitions by 15% compared to the SE scheme, but an extra indication bit is added in every data word to represent inversion occurrence. The extra bit increases the transmission overhead and the bit transitions. This paper proposes an embedded transition inversion (ETI) coding scheme that uses the phase difference between the clock and data in the transmitted serial data to tackle the problem of the extra indication bit. The ETI coding scheme reduces the transition by up to 31% compared to SE scheme. The analysis and simulation results indicate that the proposed coding scheme produces a low bit transition for different kinds of data patterns. Using the optimum degree of multiplexing, width, and spacing, the ETI coding scheme achieves 30%-60% energy reduction compared with the parallel bus without overhead. Taking circuit overhead into consideration, the power saving is up to 31.71% and 26.46% at a clock cycle of 250 ps for the 90- and 130-nm CMOS technology for m=2 where m is the number of parallel wires multiplexed into a serial link. Ching-Te Chiu, Wen-Chih Huang, Chih-Hsing Lin, Wei-Chih Lai, Ying-Fang Tsao |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Low Propagation Delay Load-Balanced 4 × 4 Switch Fabric IC in 0.13-µm CMOS TechnologyabstractA load-balanced Birkhoff-von Neumann (LB-BvN) 4 × 4 switch fabric IC is proposed for feedback-based switch systems. This is fabricated in 0.13- μm CMOS technology and the chip area is 1.380 × 1.080 mm2. The overall data rate of the LB-BvN 4 × 4 switch fabric IC is up to 32 Gb/s (8 Gb/s/channel) with only 0.8 ns propagation delay. The LB-BvN switch is highly recommended for constructing the next-generation terabit switch. In a feedback-based switch system, the long propagation delay of the switch module reduces the system throughput significantly. In this paper, we present a scalable LB-BvN 4 × 4 switch fabric IC directly in the high-speed domain. By observing the deterministic switching pattern of the N×N LB-BvN switch, we present a low-complexity pattern generator that reduces the PG complexity from O(N3) to O(1). This technique reduces the propagation delay of the switch module from 30 to 0.8 ns, and also provides 80% area saving and 85% power saving compared to serializer-deserializer interfaces. The proposed LB-BvN 4 × 4 switch fabric IC is suitable for feedback-based switch systems to solve the throughput degradation problem. Ching-Te Chiu, Yu-Hao Hsu, Wei-Chih Lai, Jen-Ming Wu, Shawn S. H. Hsu, Yang-Syu Lin, Fanta Chen, Min-Sheng Kao, Yarsun Hsu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | A 5.8 Gbps uniform mapping data center switchabstractWith the growing of cloud computing, the need of computing power no longer can be satisfied with a few powerful servers or small scale parallel computer systems. More and more servers are connected together as a data center network. Then, fault tolerance becomes an import issue when building a massive data center network. Currently, many researches focus on building fat-tree data center networks. In this paper, we propose a data center switch with uniform mapping connection patterns to provide higher fault tolerant capability for heavy traffic load fat-tree data center networks. A 4 × 4 banyan type switch IC is demonstrated as the commodity switch for building the fault tolerant fat-tree data center networks. The 4 × 4 banyan type switch IC is fabricated in 90 nm CMOS technology, and the maximum operation rate of the IC is 5.8 Gbps with only 23 ps peak-to-peak jitter. Wei-Chih Lai, Ching-Te Chiu |
ICASSP | 2 |
| 2012 | A novel low gate-count serializer topology with Multiplexer-Flip-FlopsabstractThis paper proposes Multiplexer-Flip-Flops (MUX-FFs) to be a high-throughput and low-cost solution for serial link transmitters. We also propose Multiplexer-Latches (MUX-Latches) that possess the logic function of combinational circuits and storing capacity of sequential circuits. Adopting the pipeline with MUX-FFs, which are composed of cascaded latches and MUX-Latches, many latch gates for sequencing can be removed. Analysis shows that an 8-to-1 serializer in the pipeline topology with MUX-FFs reduces 52% gate-count compared to the traditional pipeline topology. To verify the function of the proposed design, a chips is implemented with the proposed 8-to-1 serializer with MUX-FFs in 90 nm CMOS technology. The measured results show that the proposed serializer with MUX-FFs are bit-error-free (with BER-12), operating at up to 12 Gbit/s. Wei-Yu Tsai, Ching-Te Chiu, Jen-Ming Wu, Shawn S. H. Hsu, Yarsun Hsu, Ying-Fang Tsao |
ISCAS | 2 |
| 2011 | On the design and analysis of fault tolerant NoC architecture using spare routersabstractThe aggressive advent in VLSI manufacturing technology has made dramatic impacts on the dependability of devices and interconnects. In the modern manycore system, mesh based Networks-on-Chip (NoC) is widely adopted as on chip communication infrastructure. It is critical to provide an effective fault tolerance scheme on mesh based NoC. A faulty router or broken link isolates a well functional processing element (PE). Also, a set of faulty routers form faulty regions which may break down the whole design. To address these issues, we propose an innovative router-level fault tolerance scheme with spare routers which is different from the traditional microarchitecture-level approach. The spare routers not only provide redundancies but also diversify connection paths between adjacent routers. To exploit these valuable resources on fault tolerant capabilities, two configuration algorithms are demonstrated. One is shift-and-replace-allocation (SARA) and the other is defect-awareness-path-allocation (DAPA) that takes advantage of path diversity in our architecture. The proposed design is transparent to any routing algorithm since the output topology is consistent to the original mesh. Experimental results show that our scheme has remarkable improvements on fault tolerant metrics including reliability, mean time to failure (MTTF), and yield. In addition, the performance of spare router increases with the growth of NoC size but the relative connection cost decreases at the same time. This rare and valuable characteristic makes our solution suitable for large scale NoC design. Yung-Chang Chang, Ching-Te Chiu, Shih-Yin Lin, Chung-Kai Liu |
ASP-DAC | 2 |
| 2011 | A 32Gbps low propagation delay 4×4 switch IC for feedback-based system in 0.13μm CMOS technologyabstractIn this paper, a low propagation delay, low power, and area-efficient 4×4 load-balanced switch circuit for feedback-based system is presented. In this periodic and deterministic switch, only two DFFs are used to implement a pattern generator which is a O(N3) hardware complexity in traditional matching algorithm based N×N switch. For packet reordering, a feedback path is established in series of symmetric patterns. As comparing with commercial switch systems, we implement a 4×4 switch IC directly in high speed domain without the use of SERDES interfaces to achieve low propagation delay and high scalability. In CML output buffer, PMOS active load and active back-end termination are introduced. A stacked current source and symmetric topology in CML-DFF are adopted. From our results, this work efficiently deducted 28ns propagation delay, 80% area and 80% power introduced by the SERDES interface. The throughput rate is up to 32Gbps (8Gbps/Ch). Yu-Hao Hsu, Yang-Syu Lin, Ching-Te Chiu, Jen-Ming Wu, Shuo-Hung Hsu, Fanta Chen, Min-Sheng Kao, Wei-Chih Lai, Yarsun Hsu |
ASP-DAC | 3 |
| 2011 | Handover Delay Reduction and Buffer-Based Data Recovery Scheme for Inter Multicast Broadcast Service ZoneabstractMulticast broadcast service (MBS) is one of the important features supported by Mobile WiMAX to efficiently transmit data common to a group of users. As MBS services are usually delay-sensitive applications, the concept of the MBS zone is also introduced to provide better quality of service (QoS) for mobile users. However, the large inter- MBS zone handover delay and frame offsets between adjacent MBS zones cause large packet loss, but this problem is little studied. In this paper, we propose an improvement on the inter-MBS zone handover procedure to greatly reduce the handover delay. Furthermore, we also propose a data recovery scheme for inter-MBS zone handover by using an additional multicast connection as a recovery channel for each MBS service to minimize the packet loss. Simulation results show that our proposed schemes achieves almost zero packet loss when the number of recovery channels is the same with the number of MBS sessions. Sih-Kai Li, Jen-Shun Yang, Ching-Te Chiu, Po-Ting Yeh, Jenq-Neng Hwang |
GLOBECOM | 3 |
| 2011 | Curve-based and image-based JND contrast analysis for inverse tone mapping operatorsabstractRecent studies on inverse tone mapping attempt to reproduce real-world images using low dynamic range (LDR) images. Evaluation metrics to qualify the performance of inverse tone mapping operators (iTMOs) are important. Just Noticeable Difference (JND) is widely used in image analysis as visual sensitivity measure for quality assessment. However, this measure does not provide insights into how the characteristics of iTMO curves affect the quality of high dynamic range (HDR) images. Therefore, based on the probability of various contrast changes, this study proposes a curve-based and an image-based JND quality assessment metric to detect visual distortion. The results of this curve-based and image-based quality metric match those of the visible difference predictor (VDP) method. This metric does not require complex calibration and involves only simple computation. In addition, this method reveals the effect of various iTMO curve parameters. Chih-Rung Chen, Ching-Te Chiu |
ICIP | 2 |
| 2011 | Texture classification based low order local binary pattern for face recognitionabstractLocal Binary Pattern (LBP) represents a circular derivative pattern generated by the concatenation of the binary gradient directions. However, the pattern fails to extract more detailed information such as texture feature contained in the input object. In this paper, we propose a texture classification based low order LBP for face recognition. With the texture feature, we could apply this method only on eye regions rather than whole face image to reduce the computation complexity and recognition time. The experiment result shows the face recognition rate of our approach is better than the original LBP method and the average executing time ratio of proposed method to original LBP method is only 16.3%. Ching-Te Chiu, Cyuan-Jhe Wu |
ICIP | 1 |
| 2011 | Low visual difference virtual high dynamic range image synthesizer from a single legacy imageabstractNew generations of display technologies provide signicantly improved dynamic range over conventional display devices. Current inverse tone mapping schemes require multiple exposure low dynamic range (LDR) images to generate high dynamic range (HDR) images. In this work, we propose an exposure dependent S curve to convert one optimized LDR image to multiple images with different brightness, which are then fused into a virtual real scene HDR image with wide dynamic range. According to our implementation results, the dynamic range can reach about 105. The synthesizer is robust and temporally coherent, and does not require image specific parameter adjustment. This paper also presents the image quality assessment with HDR visual difference predictor (HDR-VDP) and relative entropy contrast, and our work has better performance than other inverse tone mapping operators (iTMOs) for the both image quality assessments. Tsun-Hsien Wang, Ching-Te Chiu |
ICIP | 2 |
| 2011 | A 10 to 11.5GHz rotational phase and frequency detector for clock recovery circuitabstractThis paper presents a 10.0-11.5 Gb/s full-rate phase and frequency detector integrated with the clock recovery circuit (CRC) for application in optical receivers. A rotational phase and frequency detector (RPFD) without external reference clock is proposed to train the conventional bang-bang phase detector (BBPD) to capture the clock frequency. The proposed RPFD shows 1.5 GHz capture range and works in the high speed data rate of 10 Gb/s. Only one closed loop is used to track the internal clock. The acquisition time for the clock frequency to adjust from 11.5 to 10 GHz is 160 ns. The fabricated chip occupies 0.7 mm in 90 nm CMOS process with -108.8 dBc/Hz at 1-MHz offset and consumes 52 mW power with 1.0-V supply. Fanta Chen, Min-Sheng Kao, Yu-Hao Hsu, Chih-Hsing Lin, Jen-Ming Wu, Ching-Te Chiu, Shuo-Hung Hsu |
ISCAS | 6 |
| 2011 | BiTA/SWCE: Image Enhancement With Bilateral Tone Adjustment and Saliency Weighted Contrast EnhancementabstractResearchers have proposed various image enhancement methods to make images better correlate to the human visual system. This letter proposed an innovative image enhancement framework that combines bilateral tone adjustment (BiTA) and saliency-weighted contrast enhancement (SWCE) methods. Unlike most curve-based global contrast enhancement methods, BiTA enhances the mid-tone regions that normally contain important scenes, in addition to the bright and dark regions. For local contrast enhancement, SWCE integrates the concept of image saliency into a simple filter-based contrast enhancement method. Regions with higher saliency values, which indicate that the regions have a higher extent of human interest, deserve a greater degree of enhancement. In addition, this letter presents the ratio of saliency-weighted relative entropy to noise to evaluate the enhancement quality. Simulation results show that the proposed schemes achieve high contrast enhancement with little noise and great image quality. Wei-Ming Ke, Chih-Rung Chen, Ching-Te Chiu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | A 0.64 mm 2 Real-Time Cascade Face Detection Design Based on Reduced Two-Field ExtractionabstractFace detection is widely used in portable consumer handheld devices aimed at low area, low power, and high performance applications. The boosted cascade algorithm is one of the fastest face detection algorithms in use, but its hardware implementation requires a huge amount of SRAM to store the input data, integral image, and classifiers. This paper proposes a novel cascade face detection architecture based on a reduced two-field feature extraction scheme for faster integral image calculation and feature extraction. This scheme reduces the required memory for storing integral images by 75%, and employs multiple register files instead of a single SRAM to speed up the integral image updating and feature extraction processes. The reduced integral images have only 5% of the features of original images. Although this approach requires more weak classifiers, the proposed parallel cascade detection architecture reduces the average detection time for one feature to 63% that of the original. A 0.64 mm215 mw (@390 fps) boosted cascade face detection is implemented under the UMC 90-nm CMOS technology. Experimental results show that this face detection system can achieve a high face detection rate in processing 160 × 120 grayscale images at a speed of 390 fps. Chih-Rung Chen, Wei-Su Wong, Ching-Te Chiu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2010 | Switching bilateral filter with a texture/noise detector for universal noise removalabstractIn this paper, we propose a switching bilateral filter (SBF) with a texture and noise detector for universal noise removal. Operation was carried out in two stages: detection followed by filtering. For detection, we propose the sorted quadrant median vector (SQMV) scheme, which includes important features such as edge or texture information. This information is utilized to allocate a reference median from SQMV, which is in turn compared with a current pixel to classify it as impulse noise, Gaussian noise, or noise-free. The SBF removes both Gaussian and impulse noise without adding another weighting function. The range filter inside the bilateral filter switches between the Gaussian and impulse modes depending on the noise classification result. Simulation results show that our noise detector has a high noise detection rate as well as a high classification rate for salt-and-pepper, uniform impulse noise and mixed impulse noise. Unlike most other impulse noise filters, the proposed SBF achieves high peak signal-to-noise ratio and great image quality by efficiently removing both types of mixed noise, salt-and-pepper with uniform noise and salt-and-pepper with Gaussian noise. In addition, the computational complexity of SBF is significantly less than that of other mixed noise filters. Chih-Hsing Lin, Jia Shiuan Tsai, Ching-Te Chiu |
ICASSP | 3 |
| 2010 | A 32Gbps low propagation delay 4×4 switch IC for feedback-based system in 0.13μm CMOS technologyabstractIn this paper, a low propagation delay, low power, and area-efficient 4×4 load-balanced switch circuit for feedback-based system is presented. In this periodic and deterministic switch, only two DFFs are used to implement a pattern generator which is a O(N3) hardware complexity in traditional matching algorithm based N×N switch. For packet reordering, a feedback path is established in series of symmetric patterns. As comparing with commercial switch systems, we implement a 4×4 switch IC directly in high speed domain without the use of SERDES interfaces to achieve low propagation delay and high scalability. In CML output buffer, PMOS active load and active back-end termination are introduced. A stacked current source and symmetric topology in CML-DFF are adopted. From our results, this work efficiently deducted 28ns propagation delay, 80% area and 80% power introduced by the SERDES interface. The throughput rate is up to 32Gbps (8Gbps/Ch). Yu-Hao Hsu, Yang-Syu Lin, Ching-Te Chiu, Jen-Ming Wu, Shuo-Hung Hsu, Fanta Chen, Min-Sheng Kao, Yarsun Hsu |
ISCAS | 3 |
| 2010 | Hardware-efficient image enhancement with bilateral tone adjustmentabstractVarious image enhancement methods have been proposed to make image better correlated to human visual perception. In this paper, bilateral tone adjustment (BiTA) is presented to globally enhance the regions with extremely high or low luminance. Different from most curve-based global contrast enhancement methods, proposed BiTA furthermore enhances mid-tone regions that normally contain important scenes. In fact, a general formulation of BiTA is established and two implementation approaches are presented: bilateral gamma adjustment (BiGA) and bilateral polynomial adjustment (BiPA). By the proper parameter setting, the details of the original image will be preserved. Moreover, combined with a simple filter-based local contrast enhancement method, local details can be strengthened without unacceptable noise artifacts. Our enhancement scheme is hardware-efficient since no complex operations are involved, and the simulation shows high-contrast and natural enhancement result. Wei-Ming Ke, Ching-Te Chiu |
ISCAS | 2 |
| 2010 | A packet-based emulating platform with serializer/deserializer interface for heterogeneous IP verificationabstractThis paper proposes a packet-based verification platform with serial link interface for emulating the hardware of the heterogeneous IPs before tape out. With the serial link interface Serializer/Deserializer (SerDes) added between IPs, significant amount of pin counts can be reduced in the platform. An adapter is inserted between IP and SerDes to convert parallel bus into packets and handle the handshaking. Under our proposed adapter architecture and handshaking scheme, the limitation on the number of the master adapter is eliminated compared with Bus-based Advanced High-performance Bus (AHB) architecture. Simulation results show the data transfer through our proposed architecture works correctly without the limitation on the number of masters. With the proposed adapter and SerDes architecture, the number of required signals in the interconnect is reduced from 79 to two for the AHB bus. Chih-Hsing Lin, Yung-Chang Chang, Wen-Chih Huang, Wei-Chih Lai, Ching-Te Chiu, Jen-Ming Wu, Shuo-Hung Hsu, Chun-Ming Huang, Chih-Chyau Yang, Shih-Lun Chen |
ISCAS | 5 |
| 2010 | A novel MUX-FF circuit for low power and high speed serial link interfacesabstractIn this paper, a novel multiplexer-flip-flop (MUX-FF) topology using the current mode logic (CML) is presented. A CML multiplexer-latch (MUX-latch) is proposed by combining a multiplexer and the loopback storage part of a latch into a single module so that the buffer part of a latch can be removed. A MUX-FF is implemented by cascading two stages of MUX-latches. The output of a MUX-FF is edge-triggered, so it is insensitive to input noise. All the paths from inputs to the output are symmetric. Power and area can be reduced due to the removal of DFFs. Simulation results show that a MUX-FF can achieve a similar frequency as a conventional tree-type MUX by saving 56% of area and 72% of power consumption. Wei-Yu Tsai, Ching-Te Chiu, Jen-Ming Wu, Shuo-Hung Hsu, Yarsun Hsu |
ISCAS | 2 |
| 2010 | Switching Bilateral Filter With a Texture/Noise Detector for Universal Noise RemovalabstractIn this paper, we propose a switching bilateral filter (SBF) with a texture and noise detector for universal noise removal. Operation was carried out in two stages: detection followed by filtering. For detection, we propose the sorted quadrant median vector (SQMV) scheme, which includes important features such as edge or texture information. This information is utilized to allocate a reference median from SQMV, which is in turn compared with a current pixel to classify it as impulse noise, Gaussian noise, or noise-free. The SBF removes both Gaussian and impulse noise without adding another weighting function. The range filter inside the bilateral filter switches between the Gaussian and impulse modes depending upon the noise classification result. Simulation results show that our noise detector has a high noise detection rate as well as a high classification rate for salt-and-pepper, uniform impulse noise and mixed impulse noise. Unlike most other impulse noise filters, the proposed SBF achieves high peak signal-to-noise ratio and great image quality by efficiently removing both types of mixed noise, salt-and-pepper with uniform noise and salt-and-pepper with Gaussian noise. In addition, the computational complexity of SBF is significantly less than that of other mixed noise filters. Chih-Hsing Lin, Jia Shiuan Tsai, Ching-Te Chiu |
IEEE Trans. Image Process. | 3 |
| 2009 | Hardware-efficient virtual high dynamic range image reproductionabstractHigh dynamic range (HDR) images keep dynamic range of luminance from 105to 108and preserve more details than low dynamic range (LDR) images. Conventional acquisition of HDR images requires several images with different exposure settings of one scene, so multiple cameras or static scene are necessary. Besides, in order to transform HDR images onto normal LDR display, tone-mapping algorithms are required which need intensive computations. In this paper, we propose a hardware-efficient virtual HDR image synthesizer that includes virtual photography and local contrast enhancement. Only one LDR image is enough to generate HDR-like images, with fine details and uniformly-distributed intensity. A real-time hardware display system suitable for image or video contrast enhancement is also implemented. Under UMC90nm technology, we can process video sequences with NTSC 720ç480 resolution at 60 frames per second (FPS), running at 100MHz and consume 0.3mm2silicon area. Wei-Ming Ke, Tsun-Hsien Wang, Ching-Te Chiu |
ICIP | 3 |
| 2009 | A 100MHz hardware-efficient boost cascaded face detection designabstractIn this paper, we present a novel face detection architecture based on the boosted cascade algorithm. A reduced two-field feature extraction scheme for integral image calculation is proposed. Based on this scheme, the required memory for storing integral images is reduced from 400Kbits to 2.016Kbits for a 160×120 gray scale image. The range of the feature size and location is also reduced so the learning time of the classifier decreases around 10%. In addition, input data are mapped into parallel memories to enhance processing speed in classifier evaluations. This boosted cascade face detection hardware consumes only 0.992 mm2under the UMC 90 nm technology and runs at 100 MHz. The experimental results show this face detector can achieve 91% face detection rate for processing 160×120 gray scale images at the speed of 190 fps. Wei-Su Wong, Chih-Rung Chen, Ching-Te Chiu |
ICIP | 3 |
| 2009 | A 10-Gb/s CML I/O Circuit for Backplane Interconnection in 0.18-µm CMOS TechnologyabstractA 10-Gb/s current mode logic (CML) input/output (I/O) circuit for backplane interconnect is fabricated in 0.18-mu m 1P6M CMOS process. Comparing with conventional I/O circuit, this work consists of input equalizer, limiting amplifier with active-load inductive peaking, duty cycle correction and CML output buffer. To enhance circuit bandwidth for 10-GB/s operation, several techniques include active load inductive peaking and active feedback with current buffer in Cherry-Hooper topology. With these techniques, it reduces 30%-65% of the chip area comparing with on-chip inductor peaking method. This design also passes the interoperability test with switch fabric successfully. It provides 600- mVppdifferential voltage swing in driving 50-Omega output loads, 40-dB input dynamic range, 40-dB voltage gain, and 8-mV input sensitivity. The total power consumption is only 85 mW in 1.8-V supply and the chip feature die size is 700 mum times 400 mum. Min-Sheng Kao, Jen-Ming Wu, Chih-Hsing Lin, Fanta Chen, Ching-Te Chiu, Shawn S. H. Hsu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2008 | Design optimization of a global/local tone mapping processor on arm SOC platform for real-time high dynamic range videoabstractAs the advance of high quality displays such as organic light- emitting diode (OLED) or laser TV, the importance of a real-time high dynamic range (HDR) data processing for display devices increases significantly. Many tone mapping algorithms are proposed for rendering HDR images or videos on display screens. The choice of tone mapping algorithm depends on characteristics of displays such as luminance range, contrast ratio and gamma correction. An ideal HDR tone mapping processor should include several tone mapping algorithms and be able to select an appropriate one for different kind of devices and applications. Such a HDR tone mapping processor has characteristics of robust core functionality, high flexibility, and low area consumption. An ARM core based system on chip (SOC) platform with HDR tone mapping ASIC is suitable for such applications. In this paper, we present a systematic methodology to develop an optimized architecture for tone mapping processor in the ARM SOC platform. We illustrate the approach by a HDR tone mapping processor that can handle both photographic and gradient compression. The optimization is achieved through four major steps: common module extraction, computation power enhancement, hardware/software partition and cost function analysis. Based on the proposed scheme, we develop an integrated photographic and gradient compression HDR tone mapping processor that can process 1024times768 images at 60 fps. This design runs at 100 MHz clock and consumes area of 13.8 mm2under TSMC 0.13 mum technology with 50% improvement in speed and area compared with previous results. Ching-Te Chiu, Tsun-Hsien Wang, Wei-Ming Ke, Chen-Yu Chuang, Jhih-Rong Chen, Ren-Song Tsay |
ICIP | 1 |
| 2008 | A 28Gbps 4×4 switch with low jitter SerDes using area-saving RF model in 0.13µm CMOS technologyabstractIn this paper, we present a 7 Gbps/Ch quad SerDes integrated with a 4times4 load-balanced switch fabric circuit for high speed networking applications. To achieve high-speed and low area, we propose an area-saving RF model device for the SerDes design. The area-saving RF model has almost the same speed and jitter performance with the RF model but only consumes one half of the area. In our hybrid design of the SerDes architecture, the area-saving RF model mixed with the baseband model can reduce 75% of area compared with the design using only the conventional RF model. The grounded coplanar waveguide (GCPW) type transmission line is also employed to reduce the clock tree skew for the quad SerDes to within 1 ps. The total area is 3 mm times 2.48 mm, including the switch fabric, the quad SerDes interface, and a LC-PLL. In our results, each input/output port of the 4x4 switch fabric can achieve 7 Gbps data rate, and the overall throughout is 28 Gbps. Yu-Hao Hsu, Ming-Hao Lu, Ping-Ling Yang, Fanta Chen, You-Hung Li, Min-Sheng Kao, Chih-Hsing Lin, Ching-Te Chiu, Jen-Ming Wu, Shuo-Hung Hsu, Yarsun Hsu |
ISCAS | 8 |
| 2007 | A 20 Gbps Scalable Load Balanced Birkhoff-von Neumann Symmetric TDM Switch IC with SERDES InterfacesabstractFor the first time, we implemented a reconfigurable load-balanced TDM switch IC with SERDES interface circuits for high speed networking applications. An N times N TDM switch could be constructed recursively from the TDM switch IC to achieve switching capacity of hundred gigabits per second or higher. The TDM switch IC contained a digital 8 times 8 TDM switch core with 8B10B CODECs and analog SERDES I/O interfaces. In the I/O interfaces, eight 2.56/3.2Gbps dual-mode 16/20:1 SERDES with CML buffers were developed. The 16/20:1 instead of 8/10:1 serializer and deserializer were used to reduce the required operating frequency in the switch core by half. New half-rate architectures and all static CMOS gates were used in the 16/20:1 serializer and deserializer for the low power consumption. A wide-band CML I/O buffer with our patented PMOS active load scheme was developed. All implementation were based on the 0.18 mum CMOS technology. Our implementation showed a 20 Gbps switching capacity for the 8 times 8 TDM switch IC. Yu-Hao Hsu, Min-Sheng Kao, Hou-Cheng Tzeng, Ching-Te Chiu, Jen-Ming Wu, Shuo-Hung Hsu |
ASP-DAC | 4 |
| 2007 | Block-Based Gradient Domain High Dynamic Range Compression Design for Real-Time ApplicationsabstractDue to progress in high dynamic range (HDR) capture technologies, the HDR image or video display on conventional LCD devices has become an important topic. Many tone mapping algorithms are proposed for rendering HDR images on conventional displays, but intensive computation time makes them impractical for video applications. In this paper, we present a real-time block-based gradient domain HDR compression for image or video applications. The gradient domain HDR compression is selected as our tone mapping scheme for its ability to compress and preserve details. We divide one HDR image/frame into several equal blocks and process each by the modified gradient domain HDR compression. The gradients of smaller magnitudes are attenuated less in each block to maintain local contrast and thus expose details. By solving the Poisson equation on the attenuated gradient field block by block, we are able to reconstruct a low dynamic range image. A real-time Discrete Sine Transform (DST) architecture is proposed and developed to solve the Poisson equation. Our synthesis results show that our DST Poisson solver can run at 50MHz clock and consume area of 9 mm2 under TSMC 0.18um technology. Tsun-Hsien Wang, Wei-Ming Ke, Ding-Chuang Zwao, Fang-Chu Chen, Ching-Te Chiu |
ICIP (3) | 5 |
| 2007 | Design and Implementation of a Real-Time Global Tone Mapping Processor for High Dynamic Range VideoabstractAs the development in high dynamic range (HDR) video capture technologies, the bit-depth video encoding and decoding has become an interesting topic. In this paper, we show that the real-time HDR video display is possible. A tone mapping based HDR video architecture pipelined with a video CODEC is presented. The HDR video is compressed by the tone mapping processor. The compressed HDR video can be encoded and decoded by the video standards, such as MPEG2, MPEG4 or H.264 for transmission and display. We propose and implement a modified photographic tone mapping algorithm for the tone mapping processor. The required luminance wordlength in the processor is analyzed and the quantization error is estimated. We also develop the digit-by-digit exponent and logarithm hardware architecture for the tone mapping processor. The synthesized results show that our real-time tone mapping processor can process a NTSC video with 720*480 resolution at 30 frames per second. Tsun-Hsien Wang, Wei-Su Wong, Fang-Chu Chen, Ching-Te Chiu |
ICIP (6) | 4 |
| 2007 | Low Power Design of High Performance Memory Access Architecture for HDTV DecoderabstractTo improve memory access efficiency and to reduce power consumption in HDTV video decoders, we propose a novel memory address mapping method and an efficient memory accessing architecture. The memory address mapping enables a computation-free memory address generation from the logical address of the data word in a video frame. The simple address generation is achieved by combining neighboring macroblocks into groups and stores the group of macroblocks in the same row of the external memory. By grouping suitable macroblocks, depending on interlaced or progressive scanning, we significantly reduce the cross-row memory accessing in the external memory, which is both time consuming and power consuming. In the memory accessing architecture, we rearrange the access order of luminance and chrominance data in motion compensations to further reduce the number of row changes of the external memory. Our analysis shows that the number of row changes is reduced by 87.67% and throughput of our memory-accessing scheme is improved by 30.91% compared to conventional approaches. Tsun-Hsien Wang, Ching-Te Chiu |
ICME | 2 |
| 2007 | A Scalable Load Balanced Birkhoff-von Neumann Symmetric TDM Switch IC for High-Speed Networking ApplicationsabstractFor the first time, a scalable load balanced Birkhoff-von Neumann TDM switch IC with SERDES interface circuits for high speed networking applications was implemented. AnyN×NBirkhoff-von Neumann TDM switch could be constructed recursively from the designed TDM switch IC to achieve switching capacity of hundred gigabits per second or higher. The TDM switch IC contained a digital 8×8 TDM switch core with 8B10B codecs and analog SERDES I/O interfaces. In the I/O interfaces, eight 2.56/3.2Gbps dual-mode 16/20:1 SERDES with CML buffers were developed. The 16/20:1 instead of 8/10:1 serializer and deserializer were used to reduce the required operating frequency in the switch core by half. New half-rate architectures and all static CMOS gates were used in the 16/20:1 serializer and deserializer for the low power consumption. A wide-band CML I/O buffer with our patented PMOS active load scheme was developed. All the implementations were based on the 0.18 μm CMOS technology. Test results showed a 20 Gbps switching capacity for the 8×8 TDM switch IC. Ching-Te Chiu, Yu-Hao Hsu, Min-Sheng Kao, Hou-Cheng Tzeng, Ming-Chang Du, Ping-Ling Yang, Ming-Hao Lu, Fanta Chen, Hung-Yu Lin, Jen-Ming Wu, Shuo-Hung Hsu, Yarsun Hsu |
ISCAS | 1 |
| 2007 | A 2.24GHz Wide Range Low Jitter DLL-Based Frequency Multiplier using PMOS Active Load for Communication ApplicationsabstractIn this paper, a wide-range DLL-based frequency multiplier with PMOS active load for communication applications is proposed. Adding the PMOS active load in the delay cells has the inductive-peaking effect to increase the operation frequency range. The DLL-based frequency multiplier uses simple exclusive-or (XOR) gates and phase blending technique for the frequency multiplications. The frequency multiplier can generateNtimes of frequency of the input clock when the number of delay cells (N) in the VCDL is even. The output frequency of the proposed frequency multiplier ranges from 80MHz to 2.24GHz using TSMC 0.18μm CMOS process. The locked time is 0.96ns locked time at 400MHz. The peak-to-peak jitter is 46ps at 80MHz and 95.3ps at 2.24GHz. The power consumption of proposed frequency multiplier is 25.79mW at 400MHz. Chih-Hsing Lin, Ching-Te Chiu |
ISCAS | 2 |
| 2006 | A Flexible Cross Connect LCAS for Bandwidth Maximization in 2.5G EoSabstractRecently Ethernet-over-SONET/SDH becomes an important topic due to massive growth of data traffic. In this paper, we propose a bandwidth efficient architecture that can transport 4 gigabit Ethernets and 20 fast Ethernets over OC-48 SONET connection. It includes a cross connect UTOPIA interface to aggregate traffic, a switch block capable of flexible and resilient LCAS member mapping, and an enhanced LCAS module that refines VCAT granularity into one-third of VC-3 (STS-1) so as to achieve 97% bandwidth efficiency in worst case. Moreover, we improve the performance of LCAS bandwidth management by modifying its control packet to achieve constant refresh time. In the end, we present a novel mapping scheme that eliminates the need of differential delay compensation memory for UDP/IP applications. Hou-Cheng Tzeng, Ching-Te Chiu |
NCA | 2 |
| 2005 | Design a simple and high performance switch using a two-stage architectureabstractRecently, there is tremendous interest in the research of two-stage switches. Unlike input-buffered switches, two-stage switches do not need to find matchings between inputs and outputs. However, two-stage switches usually suffer from the out-of-sequence problem. To design a simple and high performance switch using the two-stage architecture, we address three buffer design problems in this paper: re-sequencing buffers, central buffers and input buffers. We show that the size of the resequencing buffer needs to be proportional to the size of the central buffer to ensure that no packets are lost due to resequencing. Via simulations, we find that a moderate size of central buffer yields good throughput when traffic is not bursty. However, when the traffic is bursty, one needs to address the head-of-line blocking problem at the input. We also find that using the round-robin service policy for multiple virtual output queues at inputs may exhibit a catastrophic phenomenon, called a non-ergodic mode. When a switch is trapped in a non-ergodic mode, its throughput is sharply reduced. To solve such a problem in input buffers, we show that one may introduce "randomness" into a switch to jump out of a non-ergodic mode. Chih-Ying Tu, Cheng-Shang Chang, Duan-Shin Lee, Ching-Te Chiu |
GLOBECOM | 4 |
| 1994 | Optimal unified architectures for the real-time computation of time-recursive discrete sinusoidal transformsabstractAn optimal unified architecture that can efficiently compute the discrete cosine, sine, Hartley, Fourier, lapped orthogonal, and complex lapped transforms for a continuous stream of input data that arise in signal/image communications is proposed. This structure uses only half as many multipliers as the previous best known scheme (Liu and Chiu, 1993). The proposed architecture is regular, modular, and has only local interconnections in both data and control paths. There is no limitation on the transform size N and only 2N-2 multipliers are needed for the DCT. The throughput of this scheme is one input sample per clock cycle. The authors provide a theoretical justification by showing that any discrete transform whose basis functions satisfy the fundamental recurrence formula has a second-order autoregressive structure in its filter realization. They also demonstrate that dual generation transform pairs share the same autoregressive structure. They extend these time-recursive concepts to multi-dimensional transforms. The resulting d-dimensional structures are fully-pipelined and consist of only d 1D transform arrays and shift registers.> K. J. Ray Liu, Ching-Te Chiu, Ravi K. Kolagotla, Joseph F. JáJá |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1993 | Optimal unified IIR architectures for time-recursive discrete sinusoidal transforms
K. J. Ray Liu, Ching-Te Chiu, Ravi K. Kolagotla, Joseph F. JáJá |
ICASSP (3) | 2 |
| 1992 | Real-time parallel and fully pipelined two-dimensional DCT lattice structures with application to HDTV systemsabstractThe authors propose a fully pipelined architecture to compute the 2D discrete cosine transform (DCT) from a frame-recursive point of view. Based on this approach, two real-time parallel lattice structures for successive frame and block 2D DCT are developed. These structures are fully pipelined with throughput rate N clock cycles for an N*N successive input data frame. Moreover, the resulting 2D DCT architectures are modular, regular, and locally connected and require only two 1D DCT blocks that are extended directly from the 1D DCT structure without transposition. It is therefore suitable for VLSI implementation for high-speed HDTV systems. A parallel 2D DCT architecture and a scanning pattern for HDTV systems to achieve higher performance is proposed. The VLSI implementation of the 2D DCT using distributed arithmetic to increase computational efficiency and reduce round-off error is discussed.> Ching-Te Chiu, K. J. Ray Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |