Tsung-Han Tsai 0001

dblp:34/4711-1 · DBLP profile ↗
← Back
68ranked-venue papers
51as first author
18since 2021 · last 2026
0000-0001-7524-0621ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 30 first-author · 7 since 2021Systems, architecture and hardware · 23 · 17 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 2 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 A lightweight real-time karaoke system with improved singing voice separation via fundamental frequency estimation integration
Tsung-Han Tsai 0001, Ling-Wei Wu
Multim. Tools Appl.1
2024 FPGA-based control system for real-time driving of UHD Micro-LED display with color calibration
Tsung-Han Tsai 0001
Integr.1
2024 Design and implementation of deep learning-based object detection and tracking system
Tsung-Han Tsai 0001, Po-Hsien Wu
Integr.1
2024 GMM-Based Speaker Verification System with Hardware MFCC in SoC Design
Tsung-Han Tsai 0001, Chiao-Li Wang
Multim. Tools Appl.1
2024 Hardware architecture design for real-time SIFT extraction with reduced memory usage
Tsung-Han Tsai 0001, Nai-Chieh Tung
Multim. Tools Appl.1
2023 NL-DSE: Non-Local Neural Network with Decoder-Squeeze-and-Excitation for Monocular Depth Estimation
abstract
Monocular Depth Estimation is a popular and challenging problem for many years. IR CNNs (Convolutional Neural Networks)-based method with encoder-decoder architecture is proposed and shows a reasonable result. In this paper, we propose a SE-Net-based module for the decoder part in the encoder-decoder architecture to improve the result. We proposed a DSE (Decoder-Squeeze-and-Excitation) module to deal with the whole up-sampling process globally for the decoder part. We also include the Non-local Network space attention method to design the Non-Local Decoder-Squeeze-and-Excitation (NL-DSE) module. The proposed NL-DSE module is installed and evaluated on the NYU Depth V2 dataset and achieves higher accuracy. Moreover, the design is independent of the encoder-decoder architecture and can be applied in the other encoder-decoder networks to have a more accurate network.
Tsung-Han Tsai 0001, Wei-Chung Wan
ICASSP1
2023 BiSeNet V3: Bilateral segmentation network with coordinate attention for real-time semantic segmentation
abstract
For the semantic segmentation task, spatial information and the receptive field are indispensable. For semantic segmentation to be practically applicable, it must have real-time inference speed. However, most of today’s methods almost choose to compromise the spatial resolution and low-level detail information, which leads to a significant decrease in accuracy. In this paper, we propose a new architecture based on Bilateral Segmentation Network (BiSeNet) called BiSeNet V3. It introduces a new feature refinement module to optimize the feature map and a feature fusion module to combine the features efficiently. An attention mechanism is introduced to assist the model in capturing contextual information. We also use edge detection to enhance features for boundaries. Extensive experiments on the Cityscapes dataset show that our proposed approach achieves an excellent performance between segmentation accuracy and inference speed. Specifically, for a 768 × 1536 input, BiSeNet V3 achieved 79.0% mIoU on the Cityscapes test set with a speed of 93.8 FPS on an NVIDIA GTX 1080Ti. For a 720 × 960 input, BiSeNet V3 achieved 76.6% mIoU on the CamVid dataset with a speed of 147.6 FPS on an NVIDIA GTX 1080Ti. The result is significantly better than the other methods.
Tsung-Han Tsai 0001, Yu-Wei Tseng
Neurocomputing1
2023 Speech densely connected convolutional networks for small-footprint keyword spotting
Tsung-Han Tsai 0001, Xin-Hui Lin
Multim. Tools Appl.1
2022 A Skeleton-based Dynamic Hand Gesture Recognition for Home Appliance Control System
abstract
In recent years, advances in 3D sensors have dramatically promoted the development of dynamic hand gesture recognition research. On the other side, the task of hand pose estimation has seen significant progress due to the powerful feature extraction capabilities based on Convolutional Neural Networks (CNNs). In this paper, we present a lightweight CNNs method on hand gesture recognition for home appliance control system. We propose a two-stage CNN model to facilitate it. At the first stage, we utilize DetNet to detect the hand and generate 3D hand skeleton locations. At the second stage, a skeleton-based dynamic hand gesture recognition model is developed. We have 99.4% accuracy by the trained CNN model with the testing dataset. Besides, we implement this system on the Nvidia Jetson AGX Xavier to control the on/off of the fan and the light.
Tsung-Han Tsai 0001, Yi-Jhen Luo, Wei-Chung Wan
ISCAS1
2022 An Area-Efficient and High Throughput Hardware Implementation of Exponent Function
abstract
In this paper, an area-efficient and high throughput hardware implementation of the exponent function has been proposed. The proposed exponent calculation method eliminates the memory requirements leading to power and area savings. The pipelined hardware implementation results in a high-frequency design with reduced resources usage. The hardware implementation has been performed for Xilinx Virtex-4 FPGA board and TSMC 90nm process node. The throughput of 411.3 Mbps at 115.7 MHz frequency and 711.11 Mbps at 200 MHz frequency can be achieved for FPGA and ASIC design, respectively. The power consumption is 242mW and 6.1 mW for FPGA and ASIC platforms, respectively.
Muhammad Awais Hussain, Shung-Wei Lin, Tsung-Han Tsai 0001
ISCAS3
2022 A single-stage face detection and face recognition deep neural network based on feature pyramid and triplet loss
abstract
Abstract A practical deep learning face recognition system can be divided into several tasks. These tasks can be time‐consuming if each task is executed with the original image as the input data. And the feature extractors used by different tasks may duplicate its function. In this paper, a multi‐task training method based on feature pyramid and triplet loss to train a single‐stage face detection and face recognition deep neural network is proposed. As a single‐stage work, every task's data is passed through the same backbone network to avoid duplicate computation by sharing the weights and computation. The whole network is established using feature pyramid and anchor boxes to localise the face position, using triplet loss to establish the feature extractor, and finally matching the feature through a simple math function. The benefits of the approach are faster computation speed and less memory usage. On an Nvidia 2080Ti GPU accelerator, this system can achieve 212 FPS for a 640 × 640 resolution input and maintains 92.4% accuracy on the LFW data set.
Tsung-Han Tsai 0001, Po-Ting Chi
IET Image Process.1
2022 Refined U-net: A new semantic technique on hand segmentation
abstract
Hand segmentation aims to segment the hand profile however the biggest challenge is the segmentation hand over the face or skin-related environments. To solve these problems, many previous papers rely on a very deep neural network or collect new large-scale datasets on real-life scenes to increase the diversity and complexity. To perform it on a standard GPU, the training and inference time is still long, and always requires a large amount of GPU memory. In this paper, we propose a new hand segmentation technique, Refined U-Net, based on the original U-Net [1]. The main objective of Refined U-Net is to perform with few parameters and increasing the inference speed while achieving high accuracy during the hand segmentation process. We substantially improve its performance by refining the segmentation results and reducing the gap of feature vectors. In inference time, our Refined U-Net can prune the refinement block to increase the speed and reduce the number of parameters. We also eliminate the disadvantages of the popular dataset and propose a reliable way to create the virtual dataset. Our Refined U-Net can achieve 200 frames per second (FPS) on the GPU (RTX 2080Ti) and outperforms in accuracy for the state-of-the-art designs in Egohands [2] and GTEA [3] datasets.
Tsung-Han Tsai 0001, Shih-An Huang
Neurocomputing1
2022 Design and implementation of filterbank for MPEG-2/4 AAC system
Tsung-Han Tsai 0001, Hsing-Chuang Liu
Integr.1
2022 Hardware design for blind source separation using fast time-frequency mask technique
Tsung-Han Tsai 0001, Pei-Yun Liu, Yu-He Chiou
Integr.1
2021 Live Demonstration: Real-Time Multi-Hand Segmentation on Exhibition
abstract
In this paper, we proposed a multi-hand segmentation on exhibition. In exhibition there are many objects with similar color such as skin clothes and the decoration close to skin color. First we made a lot of virtual image to make the datasets closed to the exhibition, and combined the palm and back of hand into same picture. Secondly we proposed a robustness neural network call "Unet- Encoder Network (Unet-EN)" to train this datasets. We use pruning method to reduce the parameter and increase the speed. We implemented on NVIDIA® Jetson™ TX2. As a result, it can be implemented on some skin color space and supported multi-hand segmentation.
Tsung-Han Tsai 0001, Shih-An Huang
ISCAS1
2021 Live Demonstration: Intelligent Voice Wake-up Human-Following Robot
abstract
In this paper, we propose an implementation of human following robot on Nvidia Jetson AGX Xavier with Intel̅ Realsense D435. Firstly, we used Wake-up word recognition algorithm to wake up the robot then select the following target. Secondly, we propose a robustness neural network by combining the tracking algorithm and detection algorithm. The target information is transferred to the motor control board to move the vehicle based on the position and distance of the target. As a result, it can be implemented on the smaller embedded system.
Tsung-Han Tsai 0001, Po-Hsien Wu, Xin-Hui Lin, Po-Ting Chi
ISCAS1
2021 A robust tracking algorithm for a human-following mobile robot
abstract
Abstract The capability of recognizing and tracking a specific human is considered a key technique for serving mobile robots. This paper proposes a real‐time tracker that uses a stereo camera to track specific people in different environments. A tracker has been designed to detect the selected target from multiple people by use of block matching algorithm, human detection, and colour histogram comparison. Considering the similar colour of the objects, the predictor of the Kalman filter determines the position of the object. Finally, the mobile robot is moved according to the tracking results with a relative distance between the robot and the target. The effectiveness of the proposed method is demonstrated by experiments on many video sequences and real environments.
Tsung-Han Tsai 0001, Chia-Hsiang Yao
IET Image Process.1
2021 Memory Access Optimization for On-Chip Transfer Learning
abstract
Training of Deep Neural Network (DNN) at the edge faces the challenge of high energy consumption due to the requirements of a large number of memory accesses for gradient calculations. Therefore, it is necessary to minimize data fetches to perform training of a DNN model on the edge. In this paper, a novel technique has been proposed to reduce the memory access for the training of fully connected layers in transfer learning. By analyzing the memory access patterns in the backpropagation phase in fully connected layers, the memory access can be optimized. We introduce a new method to update the weights by introducing the delta term for every node of output and fully connected layer. Delta term aims to reduce memory access for the parameters which are required to access repeatedly during the training process of fully connected layers. The proposed technique shows 0.13x-13.93x energy savings for the training of fully connected layers for famous DNN architectures on multiple processor architectures. The proposed technique can be used to perform transfer learning on-chip to reduce energy consumption as well as memory access.
Muhammad Awais Hussain, Tsung-Han Tsai 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2020 Live Demonstration: Vision-Based Real-Time Fall Detection System on Embedded System
abstract
In this paper, we proposed an implementation of fall detection system on Raspberry Pi with the Intel̅ Neural Compute Stick 2. Firstly, we used skeleton extraction algorithm to obtain the important skeleton information. Secondly, we proposed a robustness neural network using pruning method to reduce the parameter and calculation and combined with the neural compute stick to execute the module. The final result will transfer to the Raspberry Pi to display on the monitor. As a result, it can be implemented on the smaller embedded system.
Tsung-Han Tsai 0001, Chin-Wei Hsu, Wei-Chung Wan
ISCAS1
2020 Design of vision-based indoor positioning based on embedded system
abstract
Indoor positioning techniques have become very important in recent years. Due to the wide deployment of surveillance cameras, it has become feasible to use the videos for indoor positioning. The success of using this approach can also reduce the load of security persons of watching the monitors all the time. In this study, the authors propose a vision‐based indoor positioning system. The proposed method uses a frame processing technique and applies the Gaussian mixture learning for video background model. The foreground object can be extracted by using the background subtraction. Based on the foreground object, the objects can be tracked and used in the direct linear transform, and generate a bird's‐eye map with camera information. A real‐time demonstration has been also provided. It shows the tracing of the moving objects and the bird's‐eye view.
Tsung-Han Tsai 0001, Chih-Hao Chang, Shih-Wei Chen, Chia-Hsiang Yao
IET Image Process.1
2020 Design of hand gesture recognition system for human-computer interaction
Tsung-Han Tsai 0001, Chih-Chi Huang, Kung-Long Zhang
Multim. Tools Appl.1
2019 A Sketch Classifier Technique with Deep Learning Models Realized in an Embedded System
abstract
Since 2011, due to the growth in the amount of information, the innovation of learning algorithms and the improvement of computer technology make the application of artificial intelligence feasible in a wide range of fields. This paper presents a sketch classifier technique with deep learning models. We use the depth-wise convolution layer to lighten the deep neural network. The result shows the improvement in approximately 1/5 of computation. We use Google Quick Draw dataset to train and evaluate the network, which can have 98% accuracy in 10 categories and 85% accuracy in 100 categories. Finally, we realize it on STM32F469I Discovery development board for demonstration. The system can achieve real-time implementation of sketch classification.
Tsung-Han Tsai 0001, Po-Ting Chi, Kuo-Hsing Cheng
DDECS1
2019 Implementation of FPGA-based Accelerator for Deep Neural Networks
abstract
At present, there are many researches on deep neural network (DNN) applied in life. In the task of object recognition, deep convolutional neural network (CNN) has a good performance, but it relies on GPU to solve a large number of complex operations. Thus the hardware accelerator of DNN is concerned by many people. In order to implement the DNN model on hardware, complex connection relationship and memory usage scheduling are needed. This paper presnets the design of FPGA-based accelerator for DNN. The proposed architecture is implemented on Xilinx Zynq-7020 FPGA. It takes the advantages of low latency and low usage in the task of MNIST digital identification, and keeps the 96 % recognition rate.
Tsung-Han Tsai 0001, Yuan-Chen Ho, Ming-Hwa Sheu
DDECS1
2019 Efficient Lossless Compression Scheme for Multi-channel ECG Signal
abstract
Electrocardiogram (ECG) is the recording of the heart electrical activity and used to diagnose heart disease nowadays. The diagnosis requires a large amount of time for acquiring enough multi-channel data normally. Thus storage and transmission of 12 lead ECG data will result in massive cost. In this work, we propose a multi-channel ECG lossless compression which uses the adaptive linear prediction for intra and inter channel decorrelation. The proposed technique is based on the adaptive Golomb-Rice codec for entropy coding with adaptive linear prediction. Thus the coefficient of linear prediction and Golomb-Rice codec will make self-adjustment during the process. Finally we evaluate the proposed algorithm with MIT-BIH Arrhythmia database for single-channel compression, and PTB database for multichannel compression.
Tsung-Han Tsai 0001, Fong-Lin Tsai
ICASSP1
2018 Fast mode decision method based on edge feature for HEVC inter-prediction
abstract
High‐efficiency video coding (HEVC) is an international video coding standard. It can reduce bit rate by almost 50% when preceding state‐of‐the‐art standard H.264/advanced video coding (AVC) with the same objective quality but increase about 40% encoding complexity. The computation complexity of HEVC largely relies on the inter‐prediction, which consumes about 60–70% of the total encoding time. This study proposes a fast mode decision method for HEVC inter‐prediction. Since current coding unit (CU) is the basic coding zone for HEVC encoding process, the authors apply the edge detection method with Sobel operator to extract the edge information inside CU. It is based on the threshold‐oriented means to compare those edge values and then predict the best partition mode for this CU. Furthermore, they define a weighting factor to adopt various resolutions of test sequence from wide quarter video graphics array (WQVGA) to 1600 p. The experimental results show that the proposed algorithm can reduce computational complexity by 50% on average compared with HEVC reference software, and appear to have almost the same coding performance.
Tsung-Han Tsai 0001, Sheng-Shuan Su, Ting-Yu Lee
IET Image Process.1
2018 Single-Chip Design for Intelligent Surveillance System
Tsung-Han Tsai 0001, Shih-Wei Chen
IEEE Trans. Very Large Scale Integr. Syst.1
2016 WHDVI: A wireless high definition video interface technique for digital home
Tsung-Han Tsai 0001, Pei-Yun Tsai 0001, Meng-Yuan Huang, Li-Yang Huang
Integr.1
2015 Design for an intelligent surveillance system based on system-on-a-programmable-chip platform
abstract
The digital surveillance system becomes more and more popular in recent years. The systems attempt to raise amount of high resolution cameras, consequently those systems stupendously increase the computational load on central server. As in the intelligent object recognition processing flow, the technique on tracking multiple targets, such as tracking group of people through occlusion, is still challenging. In this paper, we present a platform-based System-on-a-Programmable-Chip (SOPC) surveillance system. We discuss the behavior of the moving objects with adaptive search method. With this system, we can track the moving people in successive frame by object boundary box and velocity without color cues or appearance model. The proposed system can still solve the problem on the detection error even the foreground is similar to the background. The overall architectures of the intelligent surveillance system consist of three parts: multiple background maintenance (MBM) accelerator, object labeling and an SOPC processor.
Tsung-Han Tsai 0001, Chih-Hao Chang
ISCAS1
2015 Content-based singer classification on compressed domain audio data
Tsung-Han Tsai 0001, Yu-Siang Huang, Pei-Yun Liu, De-Ming Chen
Multim. Tools Appl.1
2014 A High Performance Object Tracking Technique with an Adaptive Search Method in Surveillance System
abstract
In the video scene, the technique on tracking multiple targets, such as tracking group of people through occlusion, is still challenging. In this paper, we present an algorithm for multiple targets tracking system. We discuss the behavior of the moving objects with adaptive search method. Several cases are classified, including the no match case, only-one match case, split case and occlusion case respectively. With this system, we can track the moving people in successive frame by object boundary box and velocity without color cues or appearance model. Even though people are interacted with each other or the occlusion is caused by other foreground objects, the proposed algorithm can still perform well. Furthermore we consider the people movement with respect to the distance with the camera as an adaptive search range to deal with the condition. As the foreground is similar to the background, the proposed algorithm can still solve the problem on the detection error.
Tsung-Han Tsai 0001, Chih-Hao Chang
ISM1
2013 Memory-efficient scalable video encoder architecture for multi-source digital home environment
abstract
In this paper, a memory-efficient and low-complexity architecture is proposed for the scalable video encoder, achieving the requirement of the multi-source digital home environment. The proposed very-large-scale integration architecture of the scalable video encoder is implemented in TSMC 0.18-μm 1P6M CMOS technology. The proposed hardware is synthesized under 0.18-μm CMOS technology; resulting throughput is 93.3M samples/sec, occupying 182K gates. Resulting power dissipation is 42.13 mW, operating at 150 MHz clock source. The performance of proposed work (the throughput) meets 1080p@30 fps real-time encoding, constrained by wireless high definition video interface for mobile environment.
Tsung-Han Tsai 0001, Zong-Hong Li, Hsueh-Yi Lin, Li-Yang Huang
ISCAS1
2013 Hybrid Frame Rate Upconversion Method Based on Motion Vector Mapping
abstract
Liquid crystal displays (LCDs), which serve as receivers of high visual quality video, suffer from motion blur issues. One of the methods in terminating motion blur is motion compensated frame rate upconversion, which is widely adopted in LCDs. In the previous work, i.e., particle-based frame rate upconversion, the computation complexity is high while repeated operations and some improper cost evaluation setup are observed. Therefore, in this paper, hybrid frame rate upconversion is proposed with two features. First, the cost evaluation for particle-based motion trajectory calibration is modified based on the possible noise sources and video resolution variations. Second, repeated operations in particle-based motion trajectory calibration are observed. Therefore, original particle-based motion trajectory calibration is replaced by initial motion vector assignment and subsequent motion vector mapping to achieve computation complexity reduction, while the effective search range is relatively expanded. According to the experiment results, the visual quality is enhanced by 1.87 dB on average, compared with state-of-the-art unidirectional-based frame rate upconversion approaches. On the other hand, the computation complexity of the proposed design is reduced by 25-90% based on target video resolution, concluding a high visual quality and low computation complexity frame rate upconversion design.
Tsung-Han Tsai 0001, Hsueh-Yi Lin
IEEE Trans. Circuits Syst. Video Technol.1
2013 Algorithm and Architecture Design of Human-Machine Interaction in Foreground Object Detection With Dynamic Scene
abstract
In the field of intelligent visual surveillance, the topic of tolerating background motions while detecting foreground motions in dynamic scene is widely explored by the recent foreground detection literatures. Applying the sophisticated background modeling method is a common solution for such dynamic background problem. However, the sophisticated background modeling method is computation intensive and involves huge memory bandwidth on data access. Realizing such approach on a multicamera surveillance system for real-time application can dramatically increase the hardware cost. This paper presents a hardware-oriented foreground detection that is based on human-machine interaction in object level (HMIiOL) scheme. The HMIiOL can vary the conditions for a moving object been regarded as a foreground object. The conditions are depending on background environment and are derived from the information from human-machine interaction. By the HMIiOL scheme, adopting a simple background modeling method can achieve well foreground detection with significant background motions. A processor based on system-on-chip design is presented for the HMIiOL-based foreground detection. The presented processor consists of accelerators to increase throughput of the computationally intensive tasks in the algorithm, and a reduced instruction set computing unit to handle the interaction task and the noncomputation-intensive tasks. Pipelining and parallelism techniques are used to increase the throughput. The detecting capability of the processor reaches HD720 at 30 Hz. The maximum throughput can be up to 32.707 Mpixels/s. Performance evaluation and comparison with existed foreground detection hardware show the improvement of our design.
Tsung-Han Tsai 0001, Chung-Yuan Lin, Sz-Yan Li
IEEE Trans. Circuits Syst. Video Technol.1
2012 Tagging Webcast Text in Baseball Videos by Video Segmentation and Text Alignment
abstract
Sports video annotation, an active research area in the field of multimedia content understanding, is an essential process in applications, such as summarization, highlight extraction, event detection, and retrieval. This paper considers the issue in relation to the annotation of baseball videos. Conventional baseball video annotation frameworks are based primarily on video content analysis, such as scoreboard recognition and machine learning techniques, which require a substantial amount of human input to collect and organize training data. The performance of such frameworks might become unstable if they encounter audiovisual patterns not included in the training data. To address the issue, we propose a novel framework for baseball video annotation that aligns high-level webcast text with low-level video content. Several cues, which are derived from the video content and webcast text, are utilized for alignment by leveraging hierarchical agglomerative clustering and genetic algorithm optimization. In addition, we develop an unsupervised method to learn the pitch segment properties of baseball videos by Markov random walk, and thereby reduce the need for human intervention substantially. Our experiments demonstrate that the proposed framework yields a robust result against a variety of video content and enhances the automaticity in baseball video annotation.
Chih-Yi Chiu, Po-Chih Lin, Sheng-Yang Li, Tsung-Han Tsai 0001, Yu-Lung Tsai
IEEE Trans. Circuits Syst. Video Technol.4
2012 Design and Implementation of Efficient Video Stabilization Engine Using Maximum a Posteriori Estimation and Motion Energy Smoothing Approach
abstract
To smooth the video content caused by handheld devices, this paper designs a hardware-oriented engine for efficient video stabilization. This engine is realized based on the motion energy computation. Maximum a posteriori estimation derives the global motion. Significantly, the motion energy smoothing accomplishes video stabilization. The global motion is smoothed by calculating the continuous and curve energy of successive frames. In addition, to achieve real-time video stabilization, efficient hardware architecture is proposed. The novel data reuse scheme is designed for enhancing the speed of corner point detection. The estimation skip technique is manipulated for lowering the computation of local motion estimation. Double buffering and pipeline running is designed for efficiently deriving the global motion. With these approaches, the corresponding hardware architecture has the characteristics of high efficiency and high throughput. The experimental results show that the proposed video stabilization engine can produce well-smooth videos and have high precision. The proposed hardware architecture enhances the performance of video stabilization with real-time and large resolution-processing ability. The objective comparison also demonstrates our good performance on video stabilization.
Tsung-Han Tsai 0001, Chih-Lun Fang, Hui-Min Chuang
IEEE Trans. Circuits Syst. Video Technol.1
2012 Exploring Contextual Redundancy in Improving Object-Based Video Coding for Video Sensor Networks Surveillance
abstract
In recent years, intelligent video surveillance attempts to provide content analysis tools to understand and predict the actions via video sensor networks (VSN) for automated wide-area surveillance. In this emerging network, visual object data is transmitted through different devices to adapt to the needs of the specific content analysis task. Therefore, they raise a new challenge for video delivery: how to efficiently transmit visual object data to various devices such as storage device, content analysis server, and remote client server through the network. Object-based video encoder can be used to reduce transmission bandwidth with minor quality loss. However, the involved motion-compensated technique often leads to high computational complexity and consequently increases the cost of VSN. In this paper, contextual redundancy associated with background and foreground objects in a scene is explored. A scene analysis method is proposed to classify macroblocks (MBs) by type of contextual redundancy. The motion search is only performed on the specific type of context of MB which really involves salient motion. To facilitate the encoding by context of MB, an improved object-based coding architecture, namely dual-closed-loop encoder, is derived. It encodes the classified context of MB in an operational rate-distortion-optimized sense. The experimental results show that the proposed coding framework can achieve higher coding efficiency than MPEG-4 coding and related object-based coding approaches, while significantly reducing coding complexity.
Tsung-Han Tsai 0001, Chung-Yuan Lin
IEEE Trans. Multim.1
2011 A Novel Design of CAVLC Decoder With Low Power and High Throughput Considerations
abstract
This paper proposes a novel algorithm and its very large scale integration design for context-based adaptive variable length code (CAVLC) decoding. In order to improve through put of CAVLC decoder, we propose two new methods, which are multiple level decoding (MLD) and nonzero skipping for run_before decoding (NZS). By performing parallel operations on the level decoder, MLD can decode two levels in one cycle at most situations, and NZS can produce several values of run_ before in the same cycle. These two methods have the advantages of low complexity and regularity. The proposed architecture needs 141 cycles/macroblock. Moreover, the proposed CAVLC decoder can run at 33.5 MHz to meet the real time requirement for 1920 × 1088 resolution. The power consumption for the 1920 × 1088 resolution is about 1.83 mW. The operation frequency can be reduced about 29.1% to 71.5% compared with other architectures. With an aid on a lower operation frequency, it is suitable for many low power applications. The synthesis result shows that the gate count is 13175 gates, and the maximum frequency can archive 160 MHz.
Tsung-Han Tsai 0001, Te-Lung Fang, Yu-Nan Pan
IEEE Trans. Circuits Syst. Video Technol.1
2011 Text-Video Completion Using Structure Repair and Texture Propagation
abstract
Today, more superimposed text is embedded within videos. Usually some text is unnecessary. Thus, one requires an approach to remove the text and complete the video. However, few conventional approaches complete the video well due to the large-sized text, structure regions, and various types of videos. In response, this study designed a text-video completion algorithm that poses text-video completion as structure repair and texture propagation. To repair the structure regions, the structure interpolation uses the new model's rotated block matching to estimate the initial location of completed regions and later refine the coordinates of completed regions. The information in the neighboring frames then fills the structure regions. To complete the structure regions without tedious manual interaction, the structure extension utilizes the spline curve estimation. Afterwards, derivative propagation realizes the texture region completion. The experiment results are based on several real TV programs, where all of the text regions were completed with spatio-temporal consistency. Additionally, comparisons present that the performance of the proposed algorithm is superior to those of conventional approaches. Its advantages include the reduction of design complexity by only integrating the structure information in multi-frame and the demonstration of structure consistency for realistic videos.
Tsung-Han Tsai 0001, Chih-Lun Fang
IEEE Trans. Multim.1
2010 High level feature: Head and body co-trakcing by Kalman filter
abstract
Tracking multiple targets in complex situation is challenging. The difficulties are tackled multiple targets with occlusions, especially when multiple involved targets are grouped and moving together in appearance. In this paper, we present a multiple targets tracking system for the management of occlusion problem. The proposed algorithm introduces a geometric shape co-tracking strategy. It decomposes targets into geometric shapes located on body and head parts based on reasonable target geometry consideration. Features selected from the decomposed geometric shapes then can be used to track targets through intersections such as occlusion. Projection histogram and ellipsoid shape model are adopted to manage decomposed geometric shapes corresponding to each target. Tracking is done through Kalman filtering process with high efficient and low complexity issue. Experimental results show that the occlusion of grouped targets can be tracked successfully on recent challenging benchmark sequences.
Chung-Yuan Lin, Sz-Yan Li, Tsung-Han Tsai 0001
ICIP4
2010 A scalable parallel hardware architecture for connected component labeling
abstract
The parallel connected component labeling used in binary image analysis is reconsidered in this paper for the high throughput and intermediate memory requirements problem on high dimensional image sequence. It is based on a proposed dual-parallel connected component labeling method. The main idea is to break the sequentiality of the labeling procedure by separating image into slices and to correctly delimit the extent of all connected components locally, on each slice, simultaneously. According to the proposed method, a scalable architecture which can be adaptive to different throughput requirement is derived. The proposed architecture consists of local label assignment, local label fusion, and global process unit. The forest structure is introduced to cope with both global and local label equivalent. Based on the forest structure, find and union operations are implemented to complete the entire connected components labeling during two raster scans. Performance of the proposed architecture estimated in terms of the number of clocks and memory requirement are brought forward to justify the superiority of the novel design compared against previous implementation.
Chung-Yuan Lin, Sz-Yan Li, Tsung-Han Tsai 0001
ICIP3
2010 A bandwidth-efficient embedded compression algorithm using two-level rate control scheme for video coding system
abstract
In modern video coding system, the memory bandwidth of external memory becomes more and more important and almost dominates the system performance. To save the memory bandwidth, the embedded compression (EC) technique is integrated into video coding system. In this paper, the bandwidth-efficient EC algorithm is proposed. It comprises three core techniques: side-oriented prediction, adjusted binary code (ABC) entropy coding, and two-level rate control scheme. The side-oriented prediction can save the extra bits for the representation of prediction residual. The prediction residual is allocated to an efficient codeword by ABC entropy coding. With two-level rate control scheme, not only the compression ratio (CR) can be precisely controlled, but also the visual quality is maintained. Experiment results show this work achieves the PSNR drop of 1.33%, and the error between CR and Target CR (TCR) is as minor as 1.60%. Consequently, this work is quite suitable for saving memory bandwidth in video coding system.
Yu-Hsuan Lee, Tsung-Han Tsai 0001
ISCAS3
2010 A 6.4 Gbit/s Embedded Compression Codec for Memory-Efficient Applications on Advanced-HD Specification
abstract
The embedded compression (EC) technique is applied to reduce the memory bandwidth and capacity in a display system. In this paper, the high-speed EC algorithm is proposed for advanced-HD specification. It mainly comprises three features: 1) the associated geometric-based probability model is developed to construct context-modeling mechanism without context-table; 2) develop content-adaptive Golomb-Rice code and geometric-based binary code as the entropy coding with minor order of context; and 3) provide the rate control mechanism to guarantee the saving ratio of memory bandwidth and capacity. With competitive coding efficiency, the computation-efficiency of the proposed EC algorithm is about 44% and 40% of FELICS and JPEG-LS. The proposed very-large-scale integration architecture of entire codec is implemented in TSMC 0.18- 1P6M CMOS technology. Based on pixel-based parallelism and segment-based parallelism techniques, the encoding/decoding capability reaches Quad Full-high definition (QFHD) (3840 × 2160) at 30 Hz. The maximum throughput is as high as 6.4 Gbit/s. Furthermore, with multi-level parallelism, the performance can be extended to QHD (2560 × 1440) at 120 Hz and QFHD at 120 Hz for the double frame rate technique.
Tsung-Han Tsai 0001, Yu-Hsuan Lee
IEEE Trans. Circuits Syst. Video Technol.1
2009 Design and Integration for Background Subtraction and Foreground Tracking Algorithm
abstract
Automatic understanding of events happening at a site is the ultimate goal for many visual surveillance systems. Understanding of events requires that certain lower level computer vision tasks be performed. These include foreground detection, labeling foreground parts, and tracking targets. To achieve these tasks, it is necessary to build background subtraction and foreground tracking in the scene. This paper proposed a hardware-oriental algorithm for background subtraction and foreground tracking. To achieve real-time processing and flexibility, the system is then mapped to a SoC architecture with a single camera. The architecture contains two acceleration units and a programmable micro-processor unit. The usage of micro-processor can provide high flexibility for events understanding in different surveillance by user program. And the proposed accelerator hardware unit is used to increase the entire throughput. Simulation results show that the foreground detection and tracking results are satisfied. Performance of the proposed architecture estimated in terms of the number of clocks is brought forward to justify the real-time processing ability for 30 CIF frames per second.
Tsung-Han Tsai 0001, Chung-Yuan Lin, De-Zhang Peng, Giua-Hua Chen
IAS1
2009 Architecture design for a low-cost and low-complexity foreground object segmentation with Multi-model Background Maintenance algorithm
abstract
This paper presents an architecture design for a low cost and low complexity foreground object detection based on Multi-model Background Maintenance (MBM) algorithm . The MBM framework basically contains two principal features. These features consist of static and dynamic pixels to represent the characteristic of background. Under this framework, a pure time-varying background image is maintained and learned using the statistical information of the multiple Gaussian distribution with principal features. In the MBM architecture, look-up table based Gaussian density function architecture is proposed. Three look-up tables are used for exponential and division of the Gaussian density function. The characteristic of Gaussian density function is also used to enormously reduce the table size in a low cost and low complexity consideration. The total gate count of the foreground object detection architecture is about 14.4 K gates with TSMC 0.18 um technology. The operation frequency of this design is up to 100 MHz.
De-Zhang Peng, Chung-Yuan Lin, Wen-Tsai Sheu, Tsung-Han Tsai 0001
ICIP4
2009 The cycle-efficient idct algorithm for H.264/SVC with DSP platform
abstract
In this paper, the cycle-efficient IDCT algorithm is proposed for H.264/SVC in DSP platform. Owing to the data structure of IDCT in H.264/SVC JSVM, the extra memory access seriously degrades the performance of DSP platform. To overcome it, the proposed algorithm mainly incorporates three techniques, data structure reordering, symmetricalbased scheduling and interleaving-parallelism technique. For each 4middot4 IDCT, the proposed algorithm achieves only 20 cycles are consumed. With two spatial layers of 4CIF and CIF, the IDCT processing speed is accelerated as high as 18.6 times under 30 fps.
Huang-Chun Lin, Yu-Hsuan Lee, Tsung-Han Tsai 0001
ICME3
2009 Data-aware Platform Realization of Videotext Extraction for Content Integration
abstract
A request to simultaneously comprehend the video content of multi-channels has become. Therefore, a content-integration system was designed by extracting and integrating the embedded videotexts with other video content. This system was developed on a dual-core platform. For speeding up videotext extraction on hardware platforms, a data-aware transfer framework was proposed. In this framework, adaptive-size blocks are transferred based on the computing data, and single instruction multiple data (SIMD) architectures were manipulated to perform parallel data fetch and computation. The evaluation results demonstrated that the extracted videotexts from the informative channel can be integrated with the main-channel content, and the framework can significantly reduce the number of required cycles in the dual-core architecture. The comparison also presented that the performance of the proposed framework is higher than that of other approaches. Therefore, the proposed frameworks help to advance the data transfer in dual-core platforms and to benefit content integration.
Chih-Lun Fang, Ren-Chih Kuo, Tsung-Han Tsai 0001
ISCAS3
2009 Two-stage Method for Specific Audio Retrieval based on MP3 Compression Domain
abstract
In this paper, the content-based retrieval of one-singer audio example based on MP3 (MPEG 1 layer III) digital music archive is considered. In our proposed method, the Sub-Band Coefficients (SBC) in a MP3 frame is used for feature extraction. Both Quantization-Tree indexing (QT) approach and the Mel-Frequency Subband Coefficients (MFSCs) approach are proposed for indexing on MP3 objects. Finally, a Melody-Line contour comparison method is used to measure the similarity between MP3 objects. Evaluations on a content-based MP3 retrieval system are performed. Experimental results show that our proposed approach can perform good performance and high accuracy. At least 95.9% of accuracy can be achieved at the top-1 retrieval result.
Tsung-Han Tsai 0001, Wei-Chin Chang
ISCAS1
2009 2DVTE: A two-directional videotext extractor for rapid and elaborate design
Tsung-Han Tsai 0001, Yung-Chien Chen, Chih-Lun Fang
Pattern Recognit.1
2008 Design and Implementation of a Videotext Extractor on Dual-Core Platform
abstract
Many videotexts exist in TV programs. Some videotexts provide valuable information. Thus, an efficient design to extract these videotexts is requested. Existing videotext extractors work on the PC platform and they are difficult to achieve real-time extraction and integration. Therefore, this work designs a videotext extractor on a dual-core platform. A distributed design framework for a dual-core platform is proposed. The extraction task is dispatched to the ARM and the DSP. The ARM core executes capture, display, control, and extraction threads. The DSP core performs algorithms. The ARM and the DSP communicate by buffers and solid channels. On the DSP side, some techniques are manipulated to optimize the videotext extractor. They include software pipeline, internal memory, adjusted program, assembly optimization, and DMA. To achieve high performance, two transferred schemes of DMA are proposed. This system is implemented on the TI Davinci DM6446 platform. All input videos are 720 times 480 with 30 fps captured from real-time DVB-T system. The simulation result shows that this extractor can process the large-size frames, and all the videotext can be extracted. With this novel architecture, the extraction speed can be enhanced to 23 frames per second.
Chih-Lun Fang, Tsung-Han Tsai 0001, Ren-Chih Kuo
APSCC2
2008 Advertisement video completion using hierarchical model
abstract
With the rapid spread of video content, some embedded videotexts are superfluous advertisements needing an approach to remove. Because conventional video completion methods only deal with the limited video repairing, an adaptive advertisement video completion algorithm recovering various types of large and structural regions is designed. It poses the task of videotext removal as a hierarchical model. At the top of the hierarchy, the rotated block matching is derived, and temporal structure is recovered by the adaptive interpolation algorithm. Simultaneously, spatial structure is completed by the extension algorithm. At the bottom of the hierarchy, the spatial texture region completion is carried out by gradient derivative smooth propagation and duplication. The contribution is its comprehension and applicability to various types of videos. The experimental results show that all of the various videotext regions can be completed with temporal and spatial consistency, and the performance is superior to other existing methods.
Chih-Lun Fang, Tsung-Han Tsai 0001
ICME2
2008 The segment-based rate control algorithm in JPEG-LS for bandwidth-efficiency applications
abstract
The JPEG-LS with rate control mechanism is widely applied in buffer and bandwidth managements. In existing works, the compression ratio (CR) gradually reaches target compression ratio (TCR) based on the adjustment of information loss level at row by row. However, the CR-fluctuation is induced and limits its performance in bandwidth-efficiency applications. In this paper, the proposed segment-based rate control (SBRC), which comprises progressive adjustment of information loss level and scaling just noticeable distortion (JND) model, is applied in JPEG-LS. Experiment results show that the CR-fluctuation can be efficiently reduced and visual quality is also maintained.
Tsung-Han Tsai 0001, Shu-Ching Kao, Yu-Xuan Lee
ICME1
2008 An efficient embedded compression algorithm using adjusted binary code method
abstract
Both computation-intensive and bandwidth-intensive are primary characterizes in closed-loop video coding systems. With the rapid progress of semiconductor industry, the commutation-intensive issue can be properly handled by parallelism or pipeline processing. Therefore, bandwidth-intensive issue becomes more and more important in video coding. The embedded compression (EC) technology is widely applied to reduce frame memory size and bandwidth requirement. In this paper, an efficient EC algorithm, including ABC-based recompression, side-clipping mechanism and average-based prediction, is proposed. With proposed EC algorithm, the compression ratio (CR) of 50% can be guaranteed with minor PSNR degradation of 4.02% on average. Although EC is an additional function in video coding system, the extra encoding computation load is just 2.4% on average. Consequently, the proposed EC algorithm can be exploited to reduce the frame memory and bandwidth requirement in video coding systems.
Yu-Xuan Lee, Tsung-Han Tsai 0001
ISCAS2
2008 Optimization techniques of AAC decoder on PACDSP VLIW processor
abstract
MPEG AAC has been widely used in variant applications and there are several standards developed based on the AAC. Considering to the trade-off between flexibility and performance, the DSP is adopted for implementation. The power consumption is an important issue for portable devices and there are limited resources on the DSP. However, due to the complex algorithms in AAC, design optimizations are required to reduce the power consumption and the memory utilization. Besides, the traditional algorithms usually not addressed on the optimization of VLIW based DSP. In this paper, we propose optimization techniques for the AAC decoding blocks on a VLIW based PACDSP processor. The realized decoder can be operated at a lower frequency of only 15 MHz and needs only 27 Kbytes of program memory and 27 Kbytes of data memory.
Chun-Nan Liu, Jui Hong Hung, Tsung-Han Tsai 0001
ISCAS3
2008 High Efficiency Architecture of Fast Block Motion Estimation with Real-Time QFHD on H.264 Video Coding
abstract
The H.264/AVC inter-prediction is performed for variable block-size motion estimation (VBSME) such as 16times16, 16times8, 8times16, 8times8, 8times4, 4times8 and 4times4, it cause the high complexity for H.264 motion estimation (ME). This investigation develops an architecture for a combined fast motion estimation algorithm with edge information mode decision (EIMD) and predict hexagon search (PHS). Compared with other popular ME architecture, the proposed architecture has a large search range and low processing frequency. For the general specification of SDTV (720times480) with 4 reference frames, search range 256times256, the proposed architecture needs only 18.66 MHz. For the very high specification of QFHD (3840times2160) with 1 reference frame, search range 256times256, the proposed architecture only requires 112 MHz. The gate count of the proposed architecture is 300 K, and the memory usage is 12.6 KB.
Tsung-Han Tsai 0001, Yu-Nan Pan
ISM1
2007 Fast Motion Estimation and Edge Information Inter-Mode Decision on H.264 Video Coding
abstract
In the upcoming video coding standard, MPEG-4 AVC/JVT/H.264, motion estimation (ME) is allowed to use seven kinds of block sizes to improve the rate-distortion performance. This new feature has achieved significant coding gain compared to coding a macroblock (MB) using fixed block size. However, ME is computational intensive with complexity increasing linearly to the number of allowed block sizes. In this paper, a combined fast motion estimation algorithm with Predict Hexagon Search (PHS) and Edge Information Mode Decision (E1MD) is proposed. EIMD uses the edge information to predict the best matching block size. It can determine the best MB type quickly. The PHS can search the candidate MV efficiently. The analysis results show that the speed improvement of our proposed algorithm over some popular fast motion estimation algorithm is about 2~15 times. Compared with JM10.2, the speed-up is marvelously enhanced as 20~40 times and can still keep good quality as JM10.2.
Yu-Nan Pan, Tsung-Han Tsai 0001
ICIP (2)2
2007 A Hybrid CAVLD Architecture Design with Low Complexity and Low Power Considerations
abstract
In this paper, we proposed a hybrid high performance VLSI architecture for MPEG-4 AVC/H.264 CAVLC decoding. We introduce two techniques to improve throughput of CAVLC decoder, which is called MSLD (multi symbol for level decoding) and NZS (non zero skip for run_before decoding). Our proposed design can decode two levels and more than two run _befores in the same clock cycle. These two techniques have the advantages of low complexity and regularity design. According to the evaluation, our proposed design needs 137 cycles in average for one macroblock decoding. It says that this CAVLC decoder can run at 33.5 MHz to meet the real time requirement for H.264 video decoding on 1920x1088 resolution. Compared with the previous design, it can reduce around 56% work frequency for the same application. With a low working frequency, it will be suitable for a low power application.
Tsung-Han Tsai 0001, De-Lung Fang, Yu-Nan Pan
ICME1
2007 Complexity Reduction of H.263 to H.264 Transcoder with Fast Mode Decision
abstract
In the paper, the authors propose a new algorithm to dramatically reduce the complexity of transcoder. In our H.263 to H.264 transcoding, the time consumption for motion estimation during encoding process is shortened. The new algorithm uses the information extracted from decoder, and the modified motion vector decomposition algorithm is also proposed. During motion re-estimation, extensive experimental results and analysis show that the complexity is reduced significantly.
Tsung-Han Tsai 0001, Hsueh-Yi Lin, Yu-Xuan Lee, Pin-Hua Chen
ISCAS1
2006 Content-Based Retrieval of Mp3 Songs For One Singer Using Quantization Tree Indexing and Melody-Line Tracking Method
abstract
In this paper, the content-based retrieval with audio examples based on MP3 digital music archive for one singer is considered. In our proposed method, the Sub-Band Coefficients (SBC) in a MP3 frame is used for feature extraction. Both Quantization-Tree indexing (QT) approach and the Melody-Line Tracking (MLT) approach are proposed for indexing MP3 objects. Finally, a two-stage comparison method is used to measure the similarity between MP3 objects. We also employ the measurement of the similarity between query sample and database items for a singer. Evaluations on a content-based MP3 retrieval system are performed. Experimental results show that our proposed approach can achieve good performance and high accuracy.
Tsung-Han Tsai 0001, Jui-Hung Hung
ICASSP (5)1
2006 Low complexity architecture design of MDCT-based psychoacoustic model for MPEG 2/4 AAC encoder
abstract
In this paper, we proposed a low complexity architecture design for psycho-acoustic model (PAM). PAM is key component of MPEG-2/4 advanced audio coding (AAC) encoder. It occupies heavy computation load in AAC encoder and makes the AAC encoder hard to be implemented on portable devices for real-time condition. In order to conquer questions describe above, we propose a MDCT-based PAM algorithm and its dedicated architecture to accelerate PAM calculation. The main advantage of MDCT-based PAM is the filterbanks in AAC can be reduced form three to two. Furthermore, we use look-up table method to replace computation of spreading-function. Second, the logarithmic number system (LNS) is used to reduce computation load of many special functions and data word-length. In the hardware architecture design, the pipeline is used to increase throughput of the PAM. Besides, a logarithmic unit that converts data into log scale at one cycle uses in our design. The proposed PAM architecture is implemented in UMC 0.18 CMOS technology. The total gate count is 69476.
Tsung-Han Tsai 0001, Jia-Her Luo, Shih-Way Huang, Sung-Che Li
ISCAS1
2005 A 3D Predict Hexagon Search Algorithm for Fast Block Motion Estimation on H.264 Video Coding
abstract
In the upcoming video coding standard, MPEG-4 AVC/JVT/H. 264, motion estimation is allowed to use multiple references and multiple block sizes to improve the rate-distortion performance. However, full exhaustive search of all block sizes is computational intensive with complexity increasing linearly to the number of allowed reference frame and block size. In this paper, a novel search algorithm, Three Dimensional Predict Hexagon Search (3DPHS) is proposed. The 3DPHS patterns depend on the characteristics of motion vector distribution. It can predict the hexagon search pattern in horizontal or vertical direction. The proposed algorithm also considered the characteristics of multiple references and multiple block sizes in H. 264. It can save all the same level search points for higher definition video sequences. The analysis shows that the speed improvement of 3DPHS over the Diamond Search (DS) and the Hexagon Based Search (HEXBS) is about 58% and 53% respectively.
Tsung-Han Tsai 0001, Yu-Nan Pan
ICME1
2005 Fast decomposition of filterbanks for the state-of-the-art audio coding
abstract
This letter derives fast decomposition for the quadrature mirror filterbanks (QMFs) of the low power spectral band replication (SBR) tools in the MPEG high efficiency advanced audio coding (HE AAC) decoder. In contrast with the standard method where computation-intensive matrix operations are employed in the QMF, the proposed method decomposes the matrix operations into conventional discrete cosine transform of type II and III (DCT-II and DCT-III) and simple permutations for easy implementation. The computational complexity can be also reduced effectively by using fast algorithms for DCT.
Shih-Way Huang, Tsung-Han Tsai 0001, Liang-Gee Chen
IEEE Signal Process. Lett.2
2004 Content-based retrieval of audio example on MP3 compression domain
abstract
This paper considers content-based retrieval of a MP3-based (MPEG 1 layer III) digital music archive. In our approach, two kinds of value, scale-factor (SCF) and sub-band coefficient (SBC), in a MP3 frame are used. These two values are extracted from the MP3 decoder to compute the MP3 features for indexing the MP3 objects. Evaluations on a content-based MP3 retrieval system indicate that our approach can achieve a good performance.
Tsung-Han Tsai 0001, Yung-Tsung Wang
MMSP1
2004 A fast binary motion estimation algorithm for MPEG-4 shape coding
abstract
This paper presents a fast binary motion estimation (BME) algorithm using diamond search pattern for MPEG-4 shape coding, which is the key technology for supporting the content-based video coding. Based on the properties of binary shape information, a boundary mask for efficient search positions can be generated. Therefore, a large number of search points can be skipped. Simulation results show that our algorithm combined with diamond shaped zones takes equal bit rate in the same quality but reduces the number of search points marvelously in BME to 0.6% compared with full search algorithm, which is described in MPEG-4 verification mode. The proposed algorithm will reduce computational complexity of shape coding significantly and be suitable for real-time software and hardware applications of MPEG-4 shape coding.
Tsung-Han Tsai 0001, Chia-Pin Chen
IEEE Trans. Circuits Syst. Video Technol.1
2003 A low power VLSI implementation for variable length decoder in MPEG-1 layer III
abstract
MPEG layer III (MP3) audio coding algorithm was a widely used audio coding standard. It involves several complex coding techniques and is therefore difficult to create an efficient architecture design. The variable length decoding was an important part, which needs great amount of search and memory read/write operations. In this paper a data driven variable length decoding algorithm is presented, which can exploit the signal statistics of variable length codes to reduce power. The decoder was designed based on simplicity and low-cost, low power consumption while retaining the high efficiency requirements.
Tsung-Han Tsai 0001, Wen-Cheng Chen, Chun-Nan Liu
ICME1
2002 Architecture design for MPEG-2 AAC filterbank decoder using modified regressive method
abstract
MPEG-2 Advanced Audio Coding (AAC) is the widely used audio standard and offers the high-quality multi-channel surround audio. In the AAC, the filterbank is the most important part in MPEG-2 AAC audio coding. It occupies the high computation load in decoding flow and can't be solved efficiently by a simple architecture. This paper presents an effective architecture that is suitable for VLSI implementation. It performs the IMDCT with dynamic windowing and overlap-add algorithm and can reduces the cost and memory storage efficiently.
Tsung-Han Tsai 0001, Jiun-Nan Liu
ICASSP1
2001 A Novel Architecture Design for MP3 Audio Decoder
abstract
MPEG Layer III (MP3) Audio decoding algorithm is involved in several complex-coding techniques and is therefore difficult to create an efficient architecture design for. This paper presents an effective method that reduces the amount of required operation significantly. Based on the computation analysis, a novel architecture was developed, which performs two main functions, IMDCT with dynamic windowing and subband Synthesis efficiently with pipeline operation. The decoder was designed based on simplicity and low-cost, while retaining the high-efficiency requirement
Tsung-Han Tsai 0001, Ya-Chau Yang
ICME1
1997 A Flexible High-Throughput VLSI Architecture with 2-D Data-Reuse for Full-Search Motion Estimation
abstract
This paper describes a data-interlacing architecture with two-dimensional (2-D) data-reuse for full-search block-matching algorithm. Based on a one-dimensional processing element (PE) array and two data-interlacing shift-register arrays, the proposed architecture can efficiently reuse data to decrease external memory accesses and save the pin counts. It also achieves 100% hardware utilization and a high throughput rate. In addition, the same chips can be cascaded for different block sizes, search ranges, and pixel rates.
Yeong-Kang Lai, Liang-Gee Chen, Tsung-Han Tsai 0001, Po-Cheng Wu
ICIP (2)3
1996 Design strategy for three-dimensional subband filter banks
abstract
Since three-dimensional (3-D) subband coding has been introduced, most researches on 3-D subband coding perform temporal filtering first. In this paper, however, we investigate the best permutation strategy for temporal, vertical, and horizontal filtering to minimize the requirement of delay elements and find that the results are opposite to our expectation.
Po-Cheng Wu, Liang-Gee Chen, Yeong-Kang Lai, Tsung-Han Tsai 0001
ICIP (1)4