EDBT 2026 Demo / reviewers in the wild / expert
Shao-Yi Chien
dblp:99/4773
· DBLP profile ↗
192ranked-venue papers
20as first author
17since 2021 · last 2026
0000-0002-0634-6294ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 148 · 18 first-author · 12 since 2021Systems, architecture and hardware · 38 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 16 · 6 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Real-Time FPGA-Based Hardware-Algorithm Co-Design for Monocular Visual Odometry on Edge DevicesabstractSpatial computing has become a cornerstone of consumer electronics in the metaverse era, powering augmented reality (AR), virtual reality (VR) head-mounted displays (HMDs), and smart glasses. A key enabling technology for these mobile platforms is visual odometry (VO), which supports accurate motion tracking for seamless navigation and interaction. However, deploying deep learning-based VO on resource-constrained edge devices remains challenging due to high computational complexity, memory usage, and power demands. Moreover, critical operations such as feature matching, triangulation, and nonlinear optimization are notoriously intensive for embedded processors, underscoring the need for application-specific acceleration. This work presents a hardware-algorithm co-designed VO acceleration system for edge deployment, implemented on a Xilinx UltraScale+ MPSoC ZCU104. The system integrates an ARM Cortex-A53 processor, a neural network accelerator, and custom modules for feature matching and pose refinement. With hardware-aware algorithmic optimizations, the proposed design achieves a 255.6× speedup in neural inference, 13.7× acceleration for geometric modules, and an additional 2.1× gain through task-level parallelism, sustaining 30.6 FPS in real time. Compared to existing FPGA-based VO designs, our system offers the highest localization accuracy while maintaining real-time performance, demonstrating its practical viability for spatial computing in real-world scenarios. Li-Yang Huang, Yu-Kai Hsieh, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | A Pyramid-Free, Memory-Efficient RGB-D Visual Odometry Accelerator via Algorithm-Hardware CodesignabstractRecent advancements in spatial computing have transformed the interaction paradigm between digital information and physical environments, enabling context-aware applications that enhance daily activities across healthcare, education, entertainment, and industry. At the heart of spatial computing lies visual odometry (VO), which estimates real-time device motion to accurately align virtual content with the real world. However, achieving real-time RGB-D VO on mobile and wearable platforms remains challenging due to limited computing resources, particularly stringent memory constraints. In this article, we present a memory-efficient RGB-D VO accelerator developed via algorithm–hardware codesign. Our method significantly reduces memory consumption by employing a hybrid pipeline that integrates sparse, feature-based initialization with dense, direct-based optimization, thereby eliminating memory-intensive image pyramids. To the best of our knowledge, this is the first pyramid-free RGB-D VO accelerator that achieves real-time operation with only 128 kB of on-chip SRAM, demonstrating the effectiveness of algorithm–hardware codesign in memory-constrained environments. Implemented in TSMC 40-nm CMOS technology, the proposed architecture achieves a 32.78% reduction in memory usage, a 10.20% smaller chip area, and a$1.68\times $improvement in frame rate compared to baseline designs, efficiently processing RGB-D data at 34.4 f/s using only 128 kB of on-chip memory. Li-Yang Huang, Pin-Yi Lin, Shao-Yi Chien |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | Efficient Concertormer for Image Deblurring and Beyond
Pin-Hung Kuo, Jinshan Pan, Shao-Yi Chien, Ming-Hsuan Yang 0001 |
ICCV | 3 |
| 2025 | Memory-Efficient RGBD Visual Odometry for Mobile DevicesabstractThe spatial computing has been a popular topic in recent years, driving the development and use of consumer electronics such as AR/VR head-mounted displays (HMDs) and smart glasses. A critical component of these mobile devices is visual odometry (VO), which provides on-device motion tracking to allow users to interact with and move freely in virtual space. VO must be sufficiently efficient to handle real-time processing on resource-constrained mobile devices. To meet this requirement, we propose a memory-efficient algorithm from a hardware perspective, achieving over tenfold memory savings. Our architecture further reduces memory usage by 32.78%, area by 10.20%, and improves performance by 1.68x. Implemented in TSMC 40 nm technology, it demonstrates competitive results compared to other works, handling nearly three times more data due to processing the depth map. Li-Yang Huang, Pin-Yi Lin, Shao-Yi Chien |
ISCAS | 3 |
| 2025 | Hardware Accelerated Marker-Based Accurate Rigid Object 6-DoF Pose Tracking SystemabstractAugmented Reality (AR) and Mixed Reality (MR) have gained more and more awareness among consumers in recent years. However, pose accuracy, stability, and latency still left a lot to be desired when it comes to users’ experiences and acceptance. This work focuses on marker-based rigid object pose accuracy and computation latency. We developed a flexible marker-based accurate rigid object 6-DoF pose tracking system. First, a calibration algorithm for marker pose configuration is proposed after the multiple-marker system is setup for a target application. Next, a pose tracking system is developed to achieve sub-mm performance with the calibrated marker pose configuration. Moreover, a high-throughput pose engine hardware accelerator is proposed to achieve real-time performance. Hua-Yang Weng, Li-Yang Huang, Shao-Yi Chien |
ISCAS | 3 |
| 2025 | Efficient Non-Blind Image Deblurring With Discriminative Shrinkage Deep NetworksabstractMost existing non-blind deblurring methods formulate the problem into a maximum-a-posteriori framework and address it by manually designing a variety of regularization terms and data terms of the latent clear images. However, explicitly designing these two terms is quite challenging, which usually leads to complex optimization problems. In this paper, we propose a Discriminative Shrinkage Deep Network for fast and accurate deblurring. Most existing methods use deep convolutional neural networks (CNNs), or radial basis functions only to learn the regularization term. In contrast, we formulate both the data and regularization terms while splitting the deconvolution model into data-related and regularization-related sub-problems. We explore the properties of the Maxout function and develop a deep CNN model with Maxout layers to learn discriminative shrinkage functions, which directly approximate the solutions of these two sub-problems. Moreover, we develop a U-Net according to Krylov subspace method to restore the latent clear images effectively and efficiently, which plays a role but is better than the conventional fast-Fourier-transform-based or conjugate gradient method. Experimental results show that the proposed method performs favorably against the state-of-the-art methods regarding efficiency and accuracy. Pin-Hung Kuo, Jinshan Pan, Shao-Yi Chien, Ming-Hsuan Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Graph Attention Convolutional Network for 3D Human Pose and Shape Estimation from Point CloudsabstractWe propose Graph Attention Convolutional Network for 3D human pose and shape estimation. Unlike most deep-learning methods that utilize RGB images as input, we opt for 3D data, believing it can convey richer information. Our method comprises two-stage models. The first, named the Local Joint Network (LJN), employs grouping techniques to gather points and predict 3D joints. The second is Graph Attention Convolutional Network, which takes 3D joints as input, leveraging a combination of Graph Convolutional Neural Networks (GC-NNs) and Transformers. The key advantage lies in its consideration of both local and non-local interactions. We acquire point clouds using synthetic data and a Kinect v2 camera. Additionally, for RGB images lacking 3D information, we introduce a Point Cloud Generation System capable of synthesizing 3D data. To the best of our knowledge, we are the first to apply this mechanism to this field. Yung-Wei Fan, Sheng-Chun Huang, Shao-Yi Chien |
ICME | 3 |
| 2023 | DPDM: Feature-Based Pose Refinement with Deep Pose and Deep Match for Monocular Visual OdometryabstractIn recent years, the metaverse has been a popular topic, and it drives many consumer electronics like AR/VR HMDs (Head Mounted Displays) and smart glasses. In these mobile devices, a critical technology is visual odometry (VO), which provides on-device motion tracking so that the user can interact with and move freely in the virtual information. In this work, we propose a novel hybrid monocular visual odometry framework named DPDM (Deep Pose and Deep Match), which properly integrates deep learning into geometry-based methods. We revisit the traditional feature-based optimization and improve it by replacing its crucial components with deep prediction. With the powerful high-level information extraction ability of deep neural networks, DPDM can obtain robust and accurate results through a simple frame-to-frame sparse feature-based pose refinement module. Experiments show that DPDM can outperform traditional VO and pure learning-based VO. Compared to state-of-the-art hybrid VO, DPDM can achieve competitive performance and higher FPS (Frames Per Second). Li-Yang Huang, Shao-Syuan Huang, Shao-Yi Chien |
ICIP | 3 |
| 2022 | FedFR: Joint Optimization Federated Framework for Generic and Personalized Face RecognitionabstractCurrent state-of-the-art deep learning based face recognition (FR) models require a large number of face identities for central training. However, due to the growing privacy awareness, it is prohibited to access the face images on user devices to continually improve face recognition models. Federated Learning (FL) is a technique to address the privacy issue, which can collaboratively optimize the model without sharing the data between clients. In this work, we propose a FL based framework called FedFR to improve the generic face representation in a privacy-aware manner. Besides, the framework jointly optimizes personalized models for the corresponding clients via the proposed Decoupled Feature Customization module. The client-specific personalized model can serve the need of optimized face recognition experience for registered identities at the local device. To the best of our knowledge, we are the first to explore the personalized face recognition in FL setup. The proposed framework is validated to be superior to previous approaches on several generic and personalized face recognition benchmarks with diverse FL scenarios. The source codes and our proposed personalized FR benchmark under FL setup are available at https://github.com/jackie840129/FedFR. Chih-Ting Liu, Chien-Yi Wang, Shao-Yi Chien, Shang-Hong Lai |
AAAI | 3 |
| 2022 | Learning Discriminative Shrinkage Deep Networks for Image Deconvolution
Pin-Hung Kuo, Jinshan Pan, Shao-Yi Chien, Ming-Hsuan Yang 0001 |
ECCV (19) | 3 |
| 2022 | Incremental False Negative Detection for Contrastive Learning
Tsai-Shien Chen, Wei-Chih Hung, Hung-Yu Tseng, Shao-Yi Chien, Ming-Hsuan Yang 0001 |
ICLR | 4 |
| 2021 | How To Exploit the Transferability of Learned Image Compression to Conventional CodecsabstractLossy image compression is often limited by the simplicity of the chosen loss measure. Recent research suggests that generative adversarial networks have the ability to overcome this limitation and serve as a multi-modal loss, especially for textures. Together with learned image compression, these two techniques can be used to great effect when relaxing the commonly employed tight measures of distortion. However, convolutional neural network-based algorithms have a large computational footprint. Ideally, an existing conventional codec should stay in place, ensuring faster adoption and adherence to a balanced computational envelope.As a possible avenue to this goal, we propose and investigate how learned image coding can be used as a surrogate to optimise an image for encoding. A learned filter alters the image to optimise a different performance measure or a particular task. Extending this idea with a generative adversarial network, we show how entire textures are replaced by ones that are less costly to encode but preserve a sense of detail.Our approach can remodel a conventional codec to adjust for the MS-SSIM distortion with over 20% rate improvement without any decoding overhead. On task-aware image compression, we perform favourably against a similar but codec-specific approach. Jan Klopp, Keng-Chi Liu, Liang-Gee Chen, Shao-Yi Chien |
CVPR | 4 |
| 2021 | Online-trained Upsampler for Deep Low Complexity Video CompressionabstractDeep learning for image and video compression has demonstrated promising results both as a standalone technology and a hybrid combination with existing codecs. However, these systems still come with high computational costs. Deep learning models are typically applied directly in pixel space, making them expensive when resolutions become large.In this work, we propose an online-trained upsampler to augment an existing codec. The upsampler is a small neural network trained on an isolated group of frames. Its parameters are signalled to the decoder. This hybrid solution has a small scope of only 10s or 100s of frames and allows for a low complexity both on the encoding and the decoding side.Our algorithm works in offline and in zero-latency settings. Our evaluation employs the popular x265 codec on several high-resolution datasets ranging from Full HD to 8K. We demonstrate rate savings between 8.6% and 27.5% and provide ablation studies to show the impact of our design decisions. In comparison to similar works, our approach performs favourably. Jan Klopp, Keng-Chi Liu, Shao-Yi Chien, Liang-Gee Chen |
ICCV | 3 |
| 2021 | Interactive Object Segmentation With Dynamic Click TransformabstractIn the interactive segmentation, users initially click on the target object to segment the main body and then provide corrections on mislabeled regions to iteratively refine the segmentation masks. Most existing methods transform these user-provided clicks into interaction maps and concatenate them with image as the input tensor. Typically, the interaction maps are determined by measuring the distance of each pixel to the clicked points, ignoring the relation between clicks and mislabeled regions. We propose a Dynamic Click Transform Network (DCT-Net), consisting of Spatial-DCT and Feature-DCT, to better represent user interactions. Spatial-DCT transforms each user-provided click with individual diffusion distance according to the target scale, and Feature-DCT normalizes the extracted feature map to a specific distribution predicted from the clicked points. We demonstrate the effectiveness of our proposed method and achieve favorable performance compared to the state-of-the-art on three standard benchmark datasets. Chun-Tse Lin, Wei-Chih Tu, Chih-Ting Liu, Shao-Yi Chien |
ICIP | 4 |
| 2021 | Hard Samples Rectification for Unsupervised Cross-Domain Person Re-IdentificationabstractPerson re-identification (re-ID) has received great success with the supervised learning methods. However, the task of unsupervised cross-domain re-ID is still challenging. In this paper, we propose a Hard Samples Rectification (HSR) learning scheme which resolves the weakness of original clustering-based methods being vulnerable to the hard positive and negative samples in the target unlabelled dataset. Our HSR contains two parts, an inter-camera mining method that helps recognize a person under different views (hard positive) and a part-based homogeneity technique that makes the model discriminate different persons but with similar appearance (hard negative). By rectifying those two hard cases, the re-ID model can learn effectively and achieve promising results on two large-scale benchmarks. Chih-Ting Liu, Man-Yu Lee, Tsai-Shien Chen, Shao-Yi Chien |
ICIP | 4 |
| 2021 | A Dense Tensor Accelerator with Data Exchange Mesh for DNN and Vision WorkloadsabstractWe propose a dense tensor accelerator called VectorMesh, a scalable, memory-efficient architecture that can support a wide variety of DNN and computer vision workloads. Its building block is a tile execution unit (TEU), which includes dozens of processing elements (PEs) and SRAM buffers connected through a butterfly network. A mesh of FIFOs between the TEUs facilitates data exchange between tiles and promote local data to global visibility. Our design performs better according to the roofline model for CNN, GEMM, and spatial matching algorithms compared to state-of-the-art architectures. It can reduce global buffer and DRAM fetches by 2-22 times and up to 5 times, respectively. Wei-Chao Chen, Chia-Lin Yang, Shao-Yi Chien |
ISCAS | 4 |
| 2021 | Two-Way Recursive FilteringabstractWe present a new way of recursive image filtering called the two-way recursive filtering. Different from most recursive image filters that pose the 2D image filtering as iterative 1D row/column filtering, we consider the propagation on a 2D image plane directly. To compute the filtered value of a pixel, our method takes feedback from two neighboring pixels, one connected horizontally and the other connected vertically. In this way, the proposed filter has a dense receptive field, leading to fewer iterations needed to filter an image. We demonstrate its ability in high-quality edge-aware image smoothing without compromising the efficiency compared to 1D filters. We show that the proposed filter can also be formulated as a layer in a deep neural network. The feedback coefficients of the recursive filter are learned via backpropagating through the proposed filter. Our experiments on segmentation refinement show that our method produces favorable results against the three-way connection scheme in a previous segmentation refinement algorithm. To the best of our knowledge, this is the first work that explores the two-way pixel connectivity for recursive filtering. Wei-Chih Tu, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Orientation-Aware Vehicle Re-Identification with Semantics-Guided Part Attention Network
Tsai-Shien Chen, Chih-Ting Liu, Chih-Wei Wu, Shao-Yi Chien |
ECCV (2) | 4 |
| 2020 | Space-Time Guided Association Learning For Unsupervised Person Re-IdentificationabstractPerson re-identification (Re-ID) aims to match images of the same person across distinct camera views. In this paper, we propose the Space-Time Guided Association Learning (STGAL) for unsupervised Re-ID without ground truth identity nor image correspondence observed during training. By exploiting the spatial-temporal information presented in pedestrian data, our STGAL is able to identify positive and negative image pairs for learning Re-ID feature representations. Experiments on a variety of datasets confirm the effectiveness of our approach, which achieves promising performance when comparing to the state-of-the-art methods. Chih-Wei Wu, Chih-Ting Liu, Wei-Chih Tu, Yu Tsao 0001, Yu-Chiang Frank Wang, Shao-Yi Chien |
ICIP | 6 |
| 2020 | A Novel Gaming Video Encoding Process Using In-Game Motion VectorsabstractIn this paper, we propose a motion preprocessing method for the use in the game application pipeline, which is composed of two stages: the proposed preprocessing method and the High Efficiency Video Coding (HEVC) encoder. The method accepts the object information from the game application, preprocesses the motion vectors of objects, and pass the preprocessed motion data to the HEVC encoder. The HEVC encoder takes the motion data as the initial (or the dedicated) value of motion estimation. Therefore, the traditional diamond search can be skipped and hence increase the encoding performance of the HEVC encoder. In the motion preprocessing method, the following three steps are taken: a coordination system transformation, determining motion vectors for 4 × 4checkerboard blocks [Atomic Block (AB)], and the selection of proper motion vectors for all varieties of prediction units in the encoder. With a focus on the special issues, such as move-out zone and bi-directional prediction, we are able to further optimize the performance of the encoder. We examined two types of 2D gaming scenes in our experiments. The experimental results show that, as compared with the original diamond search method provided by the encoder, our algorithm is able to achieve up to 49.0% time reduction of video encoding. The Bjontegaard Delta bit rate can achieve up to -17.0% in the random_access mode while combining with the x265 encoder and up to -26.2% in the lowdelay mode while combining with the HM-16 encoder. Chi-Wei Lu, Sheng-De Wang, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Utilising Low Complexity CNNs to Lift Non-Local Redundancies in Video CodingabstractDigital media is ubiquitous and produced in ever-growing quantities. This necessitates a constant evolution of compression techniques, especially for video, in order to maintain efficient storage and transmission. In this work, we aim at exploiting non-local redundancies in video data that remain difficult to erase for conventional video codecs We design convolutional neural networks with a particular emphasis on low memory and computational footprint. The parameters of those networks are trained on the fly, at encoding time, to predict the residual signal from the decoded video signal. After the training process has converged, the parameters are compressed and signalled as part of the code of the underlying video codec. The method can be applied to any existing video codec to increase coding gains while its low computational footprint allows for an application under resource-constrained conditions. Building on top of High Efficiency Video Coding, we achieve coding gains similar to those of pretrained denoising CNNs while only requiring about 1% of their computational complexity Through extensive experiments, we provide insights into the effectiveness of our network design decisions. In addition, we demonstrate that our algorithm delivers stable performance under conditions met in practical video compression: our algorithm performs without significant performance loss on very long random access segments (up to 256 frames) and with moderate performance drops can even be applied to single frames in high-resolution low delay settings. Jan Klopp, Liang-Gee Chen, Shao-Yi Chien |
IEEE Trans. Image Process. | 3 |
| 2020 | MERIT: Tensor Transform for Memory-Efficient Vision Processing on Parallel ArchitecturesabstractComputationally intensive deep neural networks (DNNs) are well- suited to run on GPUs, but newly developed algorithms usually require the heavily optimized DNN routines to work efficiently, and this problem could be even more difficult for specialized DNN architectures. In this article, we propose a mathematical formulation that can be useful for transferring the algorithm optimization knowledge across computing platforms. We discover that data movement and storage inside parallel processor architectures can be viewed as tensor transforms across memory hierarchies, making it possible to describe many memory optimization techniques mathematically. Such transform, which we call memory-efficient ranged inner-product tensor (MERIT) transform, can be applied to not only DNN tasks but also many traditional machine learning and computer vision computations. Moreover, the tensor transforms can be readily mapped to existing vector processor architectures. In this article, we demonstrate that many popular applications can be converted to a succinct MERIT notation on GPUs, speeding up GPU kernels up to 20 times while using only half as many code tokens. We also use the principle of the proposed transform to design a specialized hardware unit called MERIT-z processor. This processor can be applied to a variety of DNN tasks as well as other computer vision tasks while providing comparable area and power efficiency to dedicated DNN application-specific integrated circuits (ASICs). Wei-Chao Chen, Shao-Yi Chien |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Spatially and Temporally Efficient Non-local Attention Network for Video-based Person Re-Identification
Chih-Ting Liu, Chih-Wei Wu, Yu-Chiang Frank Wang, Shao-Yi Chien |
BMVC | 4 |
| 2019 | Multi-scale Dense Network for Single-image Super-resolutionabstractRecently, deep neural networks have led to tremendous advances in image super-resolution. As a well-known one-to-many inverse problem, the deep learning based methods tackle this issue via large receptive field. By that, the deep network could infer each output pixel from sufficient context information. However, most existing studies use larger kernel size or design a very deep network model to attain sufficient receptive field. The computational cost dramatically increments along with the training difficulty. Concerning this problem, the goal of this paper is to design an effective and trainable convolutional neural network. We proposed a multi-scale dense network (MSDN) which is composed of deep concatenation and basic blocks, namely multi-scale dense block (MSDB). The proposed MSDB use different dilated convolutions to gather multi-scale information; meanwhile concatenating the different dilated convolution results magnify the receptive field of a single layer. To facilitate the training difficulty, there are the dense skip connections in the proposed MSDB. Moreover, the deep concatenation and global skip connection are also adopted for improving training furthermore. Consequently, we achieve a large receptive field network without deeper structure. The experiments indicate that the quality of the proposed MSDN yields the state-of-the-art result. Chia-Yang Chang, Shao-Yi Chien |
ICASSP | 2 |
| 2019 | Convolutional Neural Network Accelerator with Vector QuantizationabstractDeep neural networks (DNNs) have demonstrated impressive performance in many edge computer vision tasks, causing the increasing demand for DNN accelerator on mobile and internet of things (IoT) devices. However, the massive power consumption and storage requirement make the hardware design challenging. In this paper, we introduce a DNN accelerator based on a model compression technique vector quantization (VQ), which can reduce the network model size and computation cost simultaneously. Moreover, a specialized processing element (PE) is designed with various SRAM bank configurations as well as dataflows such that it can support different codebook/kernel sizes, and keep high utilization under small input or output channel numbers. Compared to the state-of-the-art, the proposed accelerator architecture achieves 4.2 times reduction in memory access and 2.05 times throughput per cycle for batch-one inference. Heng Lee, Yi-Heng Wu, Shao-Yi Chien |
ISCAS | 4 |
| 2019 | Increasing Compactness of Deep Learning Based Speech Enhancement Models With Parameter Pruning and Quantization TechniquesabstractThe most recent studies on deep learning based speech enhancement (SE) are focused on improving denoising performance. However, successful SE applications require striking a desirable balance between the denoising performance and computational cost in real scenarios. In this study, we propose a novel parameter pruning (PP) technique, which removes redundant channels in a neural network. In addition, parameter quantization (PQ) and feature-map quantization (FQ) techniques were also integrated to generate even more compact SE models. The experimental results show that the integration of PP, PQ, and FQ can produce a compacted SE model with a size of only 9.76% compared to that of the original model, resulting in minor performance losses of 0.01 (from 0.85 to 0.84) and 0.03 (from 2.55 to 2.52) for STOI and PESQ scores, respectively. These promising results confirm that the PP, PQ, and FQ techniques can be used to effectively reduce the storage of an SE system on edge devices. Jyun-Yi Wu, Szu-Wei Fu, Chih-Ting Liu, Shao-Yi Chien, Yu Tsao 0001 |
IEEE Signal Process. Lett. | 5 |
| 2018 | Back-Projection Lightweight Network for Accurate Image Super Resolution
Chia-Yang Chang, Shao-Yi Chien |
ACCV (5) | 2 |
| 2018 | Learning a Code-Space Predictor by Exploiting Intra-Image-Dependencies
Jan Klopp, Yu-Chiang Frank Wang, Shao-Yi Chien, Liang-Gee Chen |
BMVC | 3 |
| 2018 | Learning Superpixels With Segmentation-Aware Affinity LossabstractSuperpixel segmentation has been widely used in many computer vision tasks. Existing superpixel algorithms are mainly based on hand-crafted features, which often fail to preserve weak object boundaries. In this work, we leverage deep neural networks to facilitate extracting superpixels from images. We show a simple integration of deep features with existing superpixel algorithms does not result in better performance as these features do not model segmentation. Instead, we propose a segmentation-aware affinity learning approach for superpixel segmentation. Specifically, we propose a new loss function that takes the segmentation error into account for affinity learning. We also develop the Pixel Affinity Net for affinity prediction. Extensive experimental results show that the proposed algorithm based on the learned segmentation-aware loss performs favorably against the state-of-the-art methods. We also demonstrate the use of the learned superpixels in numerous vision applications with consistent improvements. Wei-Chih Tu, Ming-Yu Liu 0001, Varun Jampani, Deqing Sun, Shao-Yi Chien, Ming-Hsuan Yang 0001, Jan Kautz |
CVPR | 5 |
| 2018 | Minimum Spanning Distance for Image SegmentationabstractIn this paper, we review the design of Minimum Barrier Distance and propose a new path-wise distance metric called Minimum Spanning Distance (MSD). Unlike most existing distance metrics, which only define distance between two pixels on gray-scale images, the proposed distance metric conceptually estimates the color space spanned by the colors on the path of interest. Therefore, the MSD takes into consideration the three channels on color images at the same time to compute distance. Compared with other distance metrics, MSD can not only achieve the highest numerical scores but also produce visually good segmentation maps in our experiment of interactive segmentation on the Gulshan dataset. Chao-Te Chou, Wei-Chih Tu, Shao-Yi Chien |
ICASSP | 3 |
| 2018 | Speech Dereverberation Based on Integrated Deep and Ensemble Learning AlgorithmabstractReverberation, which is generally caused by sound reflections from walls, ceilings, and floors, can result in severe performance degradation of acoustic applications. Due to a complicated combination of attenuation and time-delay effects, the reverberation property is difficult to characterize, and it remains a challenging task to effectively retrieve the anechoic speech signals from reverberation ones. In the present study, we proposed a novel integrated deep and ensemble learning algorithm (IDEA) for speech dereverberation. The IDEA consists of offline and online phases. In the offline phase, we train multiple dereverberation models, each aiming to precisely dereverb speech signals in a particular acoustic environment; then a unified fusion function is estimated that aims to integrate the information of multiple dereverberation models. In the online phase, an input utterance is first processed by each of the dereverberation models. The outputs of all models are integrated accordingly to generate the final anechoic signal. We evaluated the IDEA on designed acoustic environments, including both matched and mismatched conditions of the training and testing data. Experimental results confirm that the proposed IDEA outperforms single deep-neural-network-based dereverberation model with the same model architecture and training data. Wei-Jen Lee, Syu-Siang Wang, Fei Chen 0011, Xugang Lu, Shao-Yi Chien, Yu Tsao 0001 |
ICASSP | 5 |
| 2018 | SRIANN: Sphere Ring Intersection for Approximate Nearest Neighbor Search in VideosabstractIn the field of approximate nearest neighbor (ANN) search, rare of the existing approaches are tailored for video applications. The Ring Intersection Approximate Nearest Neighbor (RIANN) is the first ANN search algorithm for videos. It achieves real-time by performing the ANN search on the sparse grid and interpolating others. For some applications, the dense ANN search is needed to ensure the searching accuracy. To achieve dense ANN search in real-time, we consider the parallel computing as a solution. However, the RIANN algorithm is not suitable for parallel computing as the algorithm itself suffers from bad thread coherency. In this paper, we propose the Sphere Ring Intersection Approximate Nearest Neighbor (SRIANN), which solves the problem of bad thread coherency and improves the accuracy of ANN search compared to the original RIANN method. The experimental results show that the proposed method is the only one able to perform dense ANN search for CIF videos in real-time. Wei-Chih Tu, Shao-Yi Chien |
ICIP | 3 |
| 2018 | Computation-Performance Optimization of Convolutional Neural Networks with Redundant Kernel RemovalabstractDeep Convolutional Neural Networks (CNNs) are widely employed in modern computer vision algorithms, where the input image is convolved iteratively by many kernels to extract the knowledge behind it. However, with the depth of convolutional layers getting deeper and deeper in recent years, the enormous computational complexity makes it difficult to be deployed on embedded systems with limited hardware resources. In this paper, we propose two computation-performance optimization methods to reduce the redundant convolution kernels of a CNN with performance and architecture constraints, and apply it to a network for super resolution (SR). Using PSNR drop compared to the original network as the performance criterion, our method can get the optimal PSNR under a certain computation budget constraint. On the other hand, our method is also capable of minimizing the computation required under a given PSNR drop. Chih-Ting Liu, Yi-Heng Wu, Shao-Yi Chien |
ISCAS | 4 |
| 2018 | Direct pose estimation for planar objects
Po-Chen Wu, Hung-Yu Tseng, Ming-Hsuan Yang 0001, Shao-Yi Chien |
Comput. Vis. Image Underst. | 4 |
| 2017 | Unrolled Memory Inner-Products: An Abstract GPU Operator for Efficient Vision-Related ComputationsabstractRecently, convolutional neural networks (CNNs) have achieved great success in fields such as computer vision, natural language processing, and artificial intelligence. Many of these applications utilize parallel processing in GPUs to achieve higher performance. However, it remains a daunting task to optimize for GPUs, and most researchers have to rely on vendor-provided libraries for such purposes. In this paper, we discuss an operator that can be used to succinctly express computational kernels in CNNs and various scientific and vision applications. This operator, called Unrolled-Memory-Inner-Product (UMI), is a computationally-efficient operator with smaller code token requirement. Since a naive UMI implementation would increase memory requirement through input data unrolling, we propose a method to achieve optimal memory fetch performance in modern GPUs. We demonstrate this operator by converting several popular applications into the UMI representation, and achieve 1.3x-26.4x speedup against frameworks such as OpenCV and Caffe. Wei-Chao Chen, Shao-Yi Chien |
ICCV | 3 |
| 2017 | Distributed video codec with spatiotemporal side informationabstractIn this paper, a distributed video coding (DVC) system with spatiotemporal side information is proposed. The proposed framework addresses the problem of poor compression performance of DVC for high-motion video sequences by integrating temporal and spatial prediction schemes in one framework. Super-resolution techniques are employed for spatially-predicted side information generation, and a support vector machine is trained to adaptively select the coding structure with both spatial and temporal prediction. In addition, an encoder-driven coding mode selection at different granularities, including frame, block and coefficient levels, is adopted to further improve the coding performance for various video conditions. Experimental results show that the average BD rate reduction of the proposed framework is 12.93% compared with the DISCOVER DVC, and the coding gain is significant, especially for high-motion sequences. Moreover, the average computing complexity is only 92.26% of the DISCOVER DVC. Yueh-Ying Lee, Pin-Hung Kuo, Chia-han Lee, Yen-Kuang Chen, Shao-Yi Chien |
ISCAS | 5 |
| 2017 | Object-based on-line video summarization for internet of video thingsabstractIn order to address the high transmission bandwidth requirement of an Internet-of-Video-Things (IoVT), an object-based on-line video summarization algorithm is proposed to summarize the captured video information at the sensor nodes before being transmitted to the server. It is composed of two stages: intra-view and inter-view stages. In the intra-view stage, human object detector is employed with the proposed human object descriptor. In the inter-view stage, an on-line clustering algorithm with a two-layer K-nearest-neighbor model is also proposed for object clustering. Experimental results show that significant improvement can be achieved when compared with state-of-the-art works. Shih-Ting Lin, Yuan-Hsin Liao, Yu Tsao 0001, Shao-Yi Chien |
ISCAS | 4 |
| 2017 | VLSI architecture design of layer-based bilateral and median filtering for 4k2k videos at 30fpsabstractBilateral filtering (BLF) and median filtering (MF) are key components in many applications. As the image resolution grows rapidly, implementation of efficient filtering is highly demanded. In this paper, we present a unified VLSI architecture that is able to compute both kinds of filters for 4k2k videos at 30fps. One feature of this design is that we leverage an emerging layer-based algorithm for both BLF and MF, which turns the non-linear BLF or MF into a set of linear filtering followed by output interpolation. It allows more accurate approximation than previous works. Moreover, to address the introduced design challenges of filter units and comparison units, we propose an architecture with moving sum method (MSM) of box filtering and a parallel component selector (PCS), respectively. The MSM requires only one line buffer, and the PCS architecture allows us to complete the comparison in one cycle. Furthermore, a shared pre-computation unit is also proposed to support both kinds of filters. As a result, our design costs only 6.33KB on-chip memory and 161.5K logic gates to support both BLF and MF. It is much efficient than existing architectures which support BLF only. Ming-Yi Tai, Wei-Chih Tu, Shao-Yi Chien |
ISCAS | 3 |
| 2017 | D-PET: A direct 6 DoF pose estimation and tracking system on graphics processing unitsabstractReal-time recovering an accurate 6 DoF pose of a known planar target is essential for augmented reality and robotics applications. Despite several pose estimation tracking systems have been proposed over recent years, there is still the need for a more efficient and more accurate solution for general planar objects. In this work, we develop an innovative GPU implementation of a real-time pose estimation and tracking system. It consists of a pose estimation unit and a pose tracker unit. While the former computes an initial pose of a target using direct method, the latter realizes accurate pose tracking with a hierarchical search scheme. Experiments on both synthetic and real datasets demonstrate that the proposed algorithm performs favorably with various planar targets. By implementing our method on an embedded GPU, the system achieves to work at 11 FPS and is suitable for real-time applications. Hung-Yu Tseng, Po-Chen Wu, Shao-Yi Chien |
ISCAS | 4 |
| 2017 | Learning to Compose with Professional Photographs on the WebabstractPhoto composition is an important factor affecting the aesthetics in photography. However, it is a highly challenging task to model the aesthetic properties of good compositions due to the lack of globally applicable rules to the wide variety of photographic styles. Inspired by the thinking process of photo taking, we formulate the photo composition problem as a view finding process which successively examines pairs of views and determines their aesthetic preferences. We further exploit the rich professional photographs on the web to mine unlimited high-quality ranking samples and demonstrate that an aesthetics-aware deep ranking network can be trained without explicitly modeling any photographic rules. The resulting model is simple and effective in terms of its architectural design and data sampling method. It is also generic since it naturally learns any photographic rules implicitly encoded in professional photographs. The experiments show that the proposed view finding network achieves state-of-the-art performance with sliding window search strategy on two image cropping datasets. Yi-Ling Chen 0004, Jan Klopp, Min Sun 0001, Shao-Yi Chien, Kwan-Liu Ma |
ACM Multimedia | 4 |
| 2017 | Occlusion-aware Video Temporal ConsistencyabstractImage color editing techniques such as color transfer, HDR tone mapping, dehazing, and white balance have been widely used and investigated in recent decades. However, naively employing them to videos frame-by-frame often leads to flickering or color inconsistency. To solve it generally, earlier methods rely on temporal filtering or warping from the previous frame, but they still fail in the cases of occlusion and produce blurry results. We introduce a new framework for these challenges: (1) We develop an online keyframe strategy to keep track of the dynamic objects, where more temporal information can be acquired than a single previous frame. (2) To preserve image details, local color affine model is employed. The main concept of this post-processing step is to capture the color transformation from editing algorithms and maintain the detail structures of the raw image simultaneously. Practically, our approach takes a raw video and its per-frame processed version, and generates a temporally consistent output. In addition, we propose a video quality metric to evaluate temporal coherence. Extensive experiments and subjective test are done to show the superiority of the proposed framework with respect to color fidelity, detail preservation, and temporal consistency. Chun-Han Yao, Chia-Yang Chang, Shao-Yi Chien |
ACM Multimedia | 3 |
| 2017 | DodecaPen: Accurate 6DoF Tracking of a Passive StylusabstractWe propose a system for real-time six degrees of freedom (6DoF) tracking of a passive stylus that achieves sub-millimeter accuracy, which is suitable for writing or drawing in mixed reality applications. Our system is particularly easy to implement, requiring only a monocular camera, a 3D printed dodecahedron, and hand-glued binary square markers. The accuracy and performance we achieve are due to model-based tracking using a calibrated model and a combination of sparse pose estimation and dense alignment. We demonstrate the system performance in terms of speed and accuracy on a number of synthetic and real datasets, showing that it can be competitive with state-of-the-art multi-camera motion capture systems. We also demonstrate several applications of the technology ranging from 2D and 3D drawing in VR to general object manipulation and board games. Po-Chen Wu, Robert Wang 0002, Kenrick Kin, Christopher D. Twigg, Shangchen Han, Ming-Hsuan Yang 0001, Shao-Yi Chien |
UIST | 7 |
| 2017 | Quantizing Intersections Using Compact VoxelsabstractAbstract Efficient intersection queries are important for ray tracing. However, building and maintaining the acceleration structures is demanding, especially for fully dynamic scenes. In this paper, we propose a quantized intersection framework based on compact voxels to quantize the intersection as an approximation. With high‐resolution voxels, the scene geometry can be well represented, which enables more accurate simulation of global illumination, such as detailed glossy reflections. In terms of memory usage in our graphics processing unit implementation, voxels are binarized and compactly encoded in a few 2D textures. We evaluate the rendering quality at various voxel resolutions. Empirically, high‐fidelity rendering can be achieved at the voxel resolution of 1 K3 or above, which produces images very similar to those of ray tracing. Moreover, we demonstrate the feasibility of our framework for various illumination effects with several applications, including first‐bounce indirect illumination, glossy refraction, path tracing, direct illumination, and ambient occlusion. Yu-Jung Chen, Shao-Yi Chien |
Comput. Graph. Forum | 3 |
| 2017 | Algorithm and Architecture Design of Multirate Frame Rate Up-conversion for Ultra-HD LCD SystemsabstractIn current liquid crystal display (LCD) systems, the frame size becomes larger than the ultra-HD (3840 × 2160) resolution, and the refresh rate becomes higher than 120 Hz or more. However, the available video frame rates are usually at 24, 30, or 60 frames/s only, which are lower than the refresh rate of LCDs. To fill the gap between the video and LCD systems, frame interpolation techniques are usually adopted. Although frame rate up-conversion (FRUC) is regarded as the most efficient method, many design challenges are encountered in the current high-resolution and high-frame-rate LCD systems. In this paper, we developed a hardware-efficient multirate FRUC, which is capable of increasing the video frame rate from 24 or 60 to 120 frames/s. We improved the accuracy of motion vectors (MVs) between the video frames using predictive square search motion estimation (ME), followed by Markov random field (MRF) correction. Subsequently, we applied the block-based forward motion compensation (MC) to interpolate the intermediate frames. Thereafter, we performed subblock refinement to enhance the visual quality. The experiments show that the proposed algorithm performed well in both subjective and objective evaluations. We also designed hardware architecture for our multirate FRUC to support the current ultra-HD LCD systems. We proposed ping-pong two-way scheduling to eliminate the dependence among blocks. We realized 54%, 70%, and 35% cycle reduction and 62%, 82%, and 35% bandwidth reduction in the ME, MRF MV correction, and MC, respectively. The SRAM was shared by all modules, and its total size was reduced by 88%. We also implemented our design to a chip using the TSMC 90-nm cell library. Yung-Lin Huang, Fu-Chen Chen, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Feasible and Robust Optimization Framework for Auxiliary Information Refinement in Spatially-Varying Image EnhancementabstractIn content-based image processing, the precise inference of auxiliary information dominates various image enhancement applications. Given the rough auxiliary information provided by users or inference algorithms, a common scenario is to refine it with respect to the image content. Quadratic Laplacian regularization is generally used as the refinement framework because of the availability of closed-form solutions. However, solving the resultant large linear system imposes a great burden on commodity computing hardware systems in the form of computational time and memory consumption, so efficient computing algorithms without losing precision are required, especially for large images. In this paper, we first analyze the geometric nature of the quadratic Laplacian regularization associated with the algebraic property of the corresponding linear system, which clarifies the essential issues causing ineffective solutions for conventional optimization algorithms. Correspondingly, we propose an optimization scheme that is capable of approaching the closed-form solution in an efficient manner using existing fast local filters, and we perform a spectral analysis to validate the robustness of this method in severe conditions. Finally, experimental results show that the proposed scheme is more feasible for large input images and is more robust to obtain the effective refinement than conventional algorithms. Chia-Liang Tsai, Shao-Yi Chien |
IEEE Trans. Image Process. | 2 |
| 2016 | Optimized Regressor Forest for Image Super-Resolution
Chia-Yang Chang, Wei-Chih Tu, Shao-Yi Chien |
BMVC | 3 |
| 2016 | Real-Time Salient Object Detection with a Minimum Spanning TreeabstractIn this paper, we present a real-time salient object detection system based on the minimum spanning tree. Due to the fact that background regions are typically connected to the image boundaries, salient objects can be extracted by computing the distances to the boundaries. However, measuring the image boundary connectivity efficiently is a challenging problem. Existing methods either rely on superpixel representation to reduce the processing units or approximate the distance transform. Instead, we propose an exact and iteration free solution on a minimum spanning tree. The minimum spanning tree representation of an image inherently reveals the object geometry information in a scene. Meanwhile, it largely reduces the search space of shortest paths, resulting an efficient and high quality distance transform algorithm. We further introduce a boundary dissimilarity measure to compliment the shortage of distance transform for salient object detection. Extensive evaluations show that the proposed algorithm achieves the leading performance compared to the state-of-the-art methods in terms of efficiency and accuracy. Wei-Chih Tu, Shengfeng He, Qingxiong Yang, Shao-Yi Chien |
CVPR | 4 |
| 2016 | Fast video super-resolution via approximate nearest neighbor searchabstractImage super-resolution has gained much attention in these years, while video super-resolution remains almost unchanged. In this paper, we propose a fast super-resolution method for video. We exploit recent development of learning-based technique that achieves state-of-the-art in accuracy and efficiency for image super-resolution. We leverage the temporal coherency of video contents to approximate the nearest neighbor search in learning-based SR. Experimental results show that our method is able to produce visually similar or better results while being 20 times faster than baseline frame-by-frame fast image SR and being orders of magnitude faster than complex optimization-based video SR. Wei-Chih Tu, Shao-Yi Chien |
ICIP | 3 |
| 2016 | Constant time bilateral filtering for color imagesabstractBilateral filtering is a commonly used technique in image processing. However, being nonlinear, it is computationally expensive. The situation gets worse while the filter radius grows up. Several works have been proposed to accelerate the computation. Nevertheless, most techniques are tailored for grayscale image bilateral filtering or confined to specific kernel functions. In this paper, we propose a constant time bilateral filter for color images. Specifically, we extend an existing constant time bilateral filtering technique, which has been demonstrated the state-of-the-art for gray-scale images. We generalize the original approach as a three-stage algorithm. Based on the generalization, we explain the drawbacks of its naïve extension for color images and propose a solution that adapts to the image content. Experimental results demonstrate the effectiveness of our solution. Wei-Chih Tu, Ying-An Lai, Shao-Yi Chien |
ICIP | 3 |
| 2016 | Patch-based face hallucination with multitask deep neural networkabstractFace hallucination technique generates high-resolution face images from low-resolution ones. In this paper, we propose a patch based multitask deep learning method for face hallucination, which is robust to blurring of images. Our method is based on fully connected feedforward neural network, and the weights of the final layers are fine-tuned separately on different clusters of patches. Experimental results show that our system outperforms the prior state-of-the-art methods by a significant margin, while using less testing computation time. Wei-Jen Ko, Shao-Yi Chien |
ICME | 2 |
| 2016 | Example-based video color transferabstractColor transfer is an image processing technique commonly used to fix images with wrong colors, enhance the lighting conditions, or produce special styles to express specific emotions. With the aid of a reference image, the intended color characteristics can be properly transferred to the source images or videos. Since the applications start growing popular, many image color transfer methods have emerged; However, only few video color transfer algorithms have been proposed so far. In this paper, we propose an efficient example-based color transfer algorithm for both images and videos, which preserves the gradient details of the source by building Laplacian pyramids and improves the spatio-temporal consistency by employing patch matching method. Other than the existing quality metrics, we further propose a video quality metric regarding temporal consistency and compare our algorithm with other color transfer methods. The experimental results show that our algorithm generally produces outputs with high fidelity in terms of colors, scene details, and also spatiotemporal consistency. Chun-Han Yao, Chia-Yang Chang, Shao-Yi Chien |
ICME | 3 |
| 2016 | Perceptual HEVC/H.265 system with local just-noticeable-difference modelabstractIn this paper, a perceptual video coding system is developed based on the newest coding standard-HEVC/H.265. We promote a local Just-Noticeable-Difference (JND) model, which simultaneously considers the contrast sensitivity model, luminance masking, temporal masking, residue-dependent adjustment, and the summation effect of quantization distortion to get the weighting of importance of human eye perception for each coding unit (CU) in a video frame. The proposed algorithm achieves better bit allocation for video coding systems by changing quantization parameters at CU level. Simulation and subjective experiments show that our system achieves average 14% bitrate saving in the QP range of 27-37 without perceptual quality degradation. Wen-Wei Chao, Shao-Yi Chien |
ISCAS | 3 |
| 2016 | Learning patch-based anchors for face hallucinationabstractWith the goal of increasing the resolution of face images, recent face hallucination methods advance learning techniques which observe training low and high-resolution patches for recovering the output image of interest. Since most existing patch-based face hallucination approaches do not consider the location information of the patches to be hallucinated, the resulting performance might be limited. In this paper, we propose an anchored patch-based hallucination method, which is able to exploit and identify image patches exhibiting structurally and spatially similar information. With these representative anchors observed, improved performance and computation efficiency can be achieved. Experimental results demonstrate that our proposed method achieves satisfactory performance and performs favorably against recent face hallucination approaches. Wei-Jen Ko, Yu-Chiang Frank Wang, Shao-Yi Chien |
MMSP | 3 |
| 2016 | Direct 3D pose estimation of a planar targetabstractEstimating 3D pose of a known object from a given 2D image is an important problem with numerous studies for robotics and augmented reality applications. While the state-of-the-art Perspective-n-Point algorithms perform well in pose estimation, the success hinges on whether feature points can be extracted and matched correctly on targets with rich texture. In this work, we propose a robust direct method for 3D pose estimation with high accuracy that performs well on both textured and textureless planar targets. First, the pose of a planar target with respect to a calibrated camera is approximately estimated by posing it as a template matching problem. Next, the object pose is further refined and disambiguated with a gradient descent search scheme. Extensive experiments on both synthetic and real datasets demonstrate the proposed direct pose estimation algorithm performs favorably against state-of-the-art feature-based approaches in terms of robustness and accuracy under several varying conditions. Hung-Yu Tseng, Po-Chen Wu, Ming-Hsuan Yang 0001, Shao-Yi Chien |
WACV | 4 |
| 2016 | Lighting-driven voxels for memory-efficient computation of indirect illumination
Shao-Yi Chien |
Vis. Comput. | 2 |
| 2015 | Distributed computing in IoT: System-on-a-chip for smart cameras as an exampleabstractThere are four major components in application systems with internet-of-things (IoT): sensors, communications, computation and service, where large amount of data are acquired for ultra-big data analysis to discover the context information and knowledge behind signals. To support such large-scale data size and computation tasks, it is not feasible to employ centralized solutions on cloud servers. Thanks for the advances of silicon technology, the cost of computation become lower, and it is possible to distribute computation on every node in IoT. In this paper, we take video sensing network as an example to show the idea of distributed computing in IoT. Existing related works are reviewed and the architecture of a system-on-a-chip solution for distributed smart cameras is proposed with coarse-grained reconfigurable image stream processing architecture. It can accelerate various computer vision algorithms for distributed smart cameras in IoT. Shao-Yi Chien, Wei-Kai Chan, Yu-Hsiang Tseng, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen |
ASP-DAC | 1 |
| 2015 | Real-time eye localization, blink detection, and gaze estimation system without infrared illuminationabstractGaze tracking systems have high potential to be used as natural user interface devices; however, the mainstream systems are designed with infrared illumination, which may be harmful for human eyes. In this paper, a real-time eye localization, blink detection, and gaze estimation system is proposed without infrared illumination. To deal with various lighting conditions and reflections on the iris, the proposed system is based on a continuously updated color model for robust iris detection. Moreover, the proposed algorithm employs both the simplified and the original eye images to achieve the balance between robustness and accuracy. Experimental results show that the proposed system can achieve the accuracy of 96.8% for blink detection and the accuracy of 1.973 degree for gaze estimation with the processing speed of 10-11fps. The performance is comparable to previous works with infrared illumination. Bo-Chun Chen, Po-Chen Wu, Shao-Yi Chien |
ICIP | 3 |
| 2015 | Efficient natural color image denoising based on guided filterabstractImage denoising is always an active problem in image processing over the years. Though the principles to gray-scale image denoising have been extensively explored in recent advances. Relatively few works focus on the problem of color image denoising. In this paper, we address the denoising problem for natural color images. We propose a simple yet effective denoising algorithm based on the local color line assumption for natural color images. We further show that the proposed denoising strategy can be easily realized by slight modification to the guided filter. Moreover, the intermediate data in guided filter computation is utilized to make the denoising algorithm highly adaptable to local noise level. Experimental results show that the proposed method is effective to natural color image denoising and comparable to other complex methods with low complexity. Chia-Liang Tsai, Wei-Chih Tu, Shao-Yi Chien |
ICIP | 3 |
| 2015 | Painted face effect removal by a projector-camera system with dynamic ambient light adaptabilityabstractPainted face effect, where the textured contents are projected on the presenter's face, frequently occurs during presentations with projectors. To remove this annoying effect, in this paper, a projector-camera system is designed to modify the projection content according to the location information of the presenter generated with a background segmentation technique. A new non-linear photometric model for a projector-camera system is also introduced as well as an ambient light adaptation scheme to generate the background model for object segmentation under varied lighting conditions. Experimental results show that this system can successfully remove painted face effect without introducing any artifacts. This technique can improve the user experience for presentations with projectors. Po-Jung Chiu, Shao-Yi Chien |
ICME | 2 |
| 2015 | 3D Background Modeling in Multi-view RGB-D VideoabstractIn this paper, we proposed a 3D background modeling system for multi-view 3D video. We first reconstructed a 3D model, and we updated the subsequent frames into it using our proposed updating strategy. The results show that dynamic objects in the model can be excluded, leaving behind a compact 3D background model. Yung-Lin Huang, Ku-Chu Wei, Shao-Yi Chien |
ACM Multimedia | 3 |
| 2014 | Low complexity on-line video summarization with Gaussian mixture model based clusteringabstractTechniques of video summarization have attracted significant research interests in the past decade due to the rapid progress in video recording, computation, and communication technologies. However, most of the existing methods analyze the video in an off-line manner, which greatly reduces the flexibility of the system. On-line summarization, which can progressively process video during video recording, is then proposed for a wide range of applications. In this paper, an on-line summarization method using Gaussian mixture model is proposed. As shown in the experiments, the proposed method outperforms other on-line methods in both summarization quality and computational efficiency. It can generate summarization with a shorter latency and much lower computation resource requirements. Shun-Hsing Ou, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen, Shao-Yi Chien |
ICASSP | 5 |
| 2014 | Collaborative noise reduction using color-line modelabstractRecently, more and more natural image statistics are found useful for image restoration problems. In this paper, we propose a noise reduction technique by use of color-line assumption for natural color images. Based on the color-line model, we propose an algorithm to analyse local color statistics and recover the original image by promoting color linearity of a local patch. Moreover, the proposed method is employed on superpixels to alleviate the boundary effect of the denoising operation. The experimental results show that the proposed method can collaborate with existing noise reduction methods to successfully further boost the quality in both perceptual and objective evaluations. Wei-Chih Tu, Chia-Liang Tsai, Shao-Yi Chien |
ICASSP | 3 |
| 2014 | Eigen-patch: Position-patch based face hallucination using eigen transformationabstractFace hallucination increases the resolution of facial images and can be employed in video surveillance applications. Conventional approaches based on principal component analysis or position-patch suffers from the artifacts, low sharpness, and significant quality degradation for dis-aligned input images. In this paper, a novel face hallucination approach called eigen-patch is proposed. It combines eigen transformation with the concept of position-patch to increase local details while maintaining computation efficiency. Moreover, an image alignment procedure is proposed to align the input image to the database with multiple hypothesis verification. In addition, a re-projection procedure is also proposed to maintain the fidelity of the whole system. Experimental results show that the proposed scheme can improve the resolution of the input facial image faithfully, and it is more robust to dis-alignment between the input image and database. Hong-Yuh Chen, Shao-Yi Chien |
ICME | 2 |
| 2014 | Error resilience for key frames in distributed video coding with rate-distortion optimized mode decisionabstractDistributed video coding (DVC) is a potential solution for distributed video sensors in wireless visual sensor and machine-to-machine (M2M) networks. However, the error resilience schemes have not been fully investigated in literatures. In this paper, we propose several error resilience schemes for key frame transmission of DVC, including resending, refining, and hybrid modes. The resending mode asks the transmitters to resend the lost packets as automatic repeat-request, the refining mode conceals the corrupted key frames by temporal error concealment techniques with considering the characteristics of DVC, and the hybrid mode can adaptively select the better mode packet-by-packet with the proposed rate-distortion optimized mode decision. Experimental results show that the proposed schemes can achieve around 6-dB gain in PSNR compared with a baseline approach with intra error concealment. Furthermore, the hybrid mode can achieve 0.5-dB gain in PSNR for some sequences compared with the better one among the resending and refining modes. Hsin-Fang Wu, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen, Shao-Yi Chien |
ISCAS | 5 |
| 2014 | Automatic high dynamic range hallucination in inverse tone mappingabstractNowadays the dynamic range of displays has been higher and higher, which means that contents can be recorded and displayed with more detail. However, the original low dynamic range contents were recorded in a lower dynamic range. Such contents will be unsatisfying compared to high dynamic range contents, especially in the saturated, or overexposed region. This paper proposes an algorithm to compensate such exposed regions, which is called automatic high dynamic range image hallucination for inverse tone mapping. Inverse tone-mapping is the process of creating a high dynamic range image from a single low dynamic range image. In this work, high dynamic range image hallucination is used as the key method to reproduce the information which is lost in the low dynamic range image capturing. Previous methods require user interaction as a hallucination criteria, and is not practical in some applications where user interaction is not available. In this paper, the hallucination is performed automatically with the assistance of luminance and texture decoupling process. This scheme produces visually satisfying results and has the potential to be applied to video inverse tone-mapping with its automatic property. Pin-Hung Kuo, Huai-Jen Liang, Chi-Sun Tang, Shao-Yi Chien |
MMSP | 4 |
| 2014 | Stable pose tracking from a planar target with an analytical motion model in real-time applicationsabstractObject pose tracking from a camera is a well-developed method in computer vision. In theory, the pose can be determined uniquely from a calibrated camera. However, in practice, most real-time pose estimation algorithms experience pose ambiguity. We consider that pose ambiguity, i.e., the detection of two distinct local minima according to an error function, is caused by a geometric illusion. In this case, both ambiguous poses are plausible, but we cannot select the pose with the minimum error as the final pose. Thus, we developed a real-time algorithm for correct pose estimation for a planar target object using an analytical motion model. Our experimental results showed that the proposed algorithm effectively reduced the effects of pose jumping and pose jittering. To the best of our knowledge, this is the first approach to address the pose ambiguity problem using an analytical motion model in real-time applications. Po-Chen Wu, Yao-Hung Tsai, Shao-Yi Chien |
MMSP | 3 |
| 2014 | Edge-aware depth completion for point-cloud 3D scene visualization on an RGB-D cameraabstractNowadays, 3D scene reconstruction using RGB-D videos becomes more popular because of the widely-available off-the-shelf RGB-D camera. However, the depth information from current RGB-D camera still need improved in order to reconstruct the 3D scene with better quality. In this paper, an edge-aware depth completion method aims to recover more accurate depth information is proposed. There are mainly two parts in our proposed method. The first part is the edge-aware color image analysis, and the second part is depth image processing including unreliable depth pixel invalidation and filling. The depth image processing can retrieve more accurate depth information using our proposed edge-aware color image analysis. Consequently, we can not only preserve the reliable depth information, but also fill in the appropriate depth values to align edges of depth image with edges of its corresponding color image. Besides, the experimental results show that the visualization of the reconstructed point-cloud 3D scene benefits from our proposed edge-aware depth completion. Finally, the PSNR evaluation using ground truth depth information is presented. Yung-Lin Huang, Tang-Wei Hsu, Shao-Yi Chien |
VCIP | 3 |
| 2014 | Communication-efficient multi-view keyframe extraction in distributed video sensorsabstractVideo sensors are widely used in many applications such as security monitoring and home care. However, the growth of the number of sensors makes it impractical to stream all videos back to a central server for further processing, due to communication bandwidth and server storage constraints. Multi-view video summarization allows us to discard redundant data in the video streams taken by a group of sensors. All prior multi-view summarization methods, however, process video data in an off-line and centralized manner, which means that all videos are still required to be streamed back to the server before conducting the summarization. This paper proposes an on-line, distributed multi-view summarization system, which integrates the ideas of Maximal Marginal Relevance (MMR) and MS-Wave, a bandwidth-efficient distributed algorithm for finding k-nearest-neighbors and k-farthest-neighbors. Empirical studies show that our proposed system can discard redundant videos and keep important keyframes as effectively as centralized approaches, while transmitting only 1/6 to 1/3 as much data. Shun-Hsing Ou, Yu-Chen Lu, Jui-Pin Wang, Shao-Yi Chien, Shou-De Lin, Mi-Yen Yeh, Chia-han Lee, Phillip B. Gibbons, V. Srinivasa Somayazulu, Yen-Kuang Chen |
VCIP | 4 |
| 2014 | VLSI Architecture Design of Guided Filter for 30 Frames/s Full-HD VideoabstractFiltering is widely used in image and video processing for various applications. Recently, the guided filter has been proposed and became one of the popular filtering methods. In this paper, to achieve the computation demand of guided filtering in full-HD video, a double integral image architecture for guided filter ASIC design is proposed. In addition, a reformation of the guided filter formula is proposed, which can prevent the error resulted from truncation in the fractional part and modify the regularization parameter ε on user's demand. The hardware architecture of the guided image filter is then proposed and can be embedded in mobile devices to achieve real-time HD applications. To the best of our knowledge, this paper is also the first ASIC design for guided image filter. With a TSMC 90-nm cell library, the design can operate at 100 MHz and support for Full-HD (1920 × 1080) 30 frame/s with 92.9K gate counts and 3.2 KB on-chip memory. Moreover, for the hardware efficiency, our architecture is also the best compared to other previous works with bilateral filter. Chieh-Chi Kao, Jui-Hsin Lai, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Algorithm and Architecture Design of High-Quality Video Upscaling Using Database-Free Texture SynthesisabstractBecause of real-time requirements and low hardware-cost constraints, conventional TV scalers can only employ basic interpolation technique and thus introduces some artifacts that degrade the viewing quality of the output sequences. In this paper, a low-complexity super-resolution (SR) algorithm, which can provide vivid output image with rich details and sharp edges, and its associated hardware architecture, is proposed. There are two main contributions in this paper. The first is the development of the database-free texture synthesis technique. With the fractal property of nature images, it is possible to find proper high-resolution patches in a low-resolution input image itself. Therefore, the texture synthesis can be performed without database to provide proper and rich details. In addition, a faithful reconstruction constraint is used to maintain the temporal consistency. The second contribution is the hardware architecture design of the database-free texture synthesis. Partial-sum reuse technique is developed to reduce 76% of computation in the texture synthesis, and a tile-based processing technique is proposed to dramatically reduce the on-chip memory and off-chip memory bandwidth requirements. Experimental results show that the proposed SR algorithm outperforms other ones with reasonable hardware cost, where 766k in gate count and 22 kB in on-chip SRAM are required to achieve full high-definition processing ability at the working frequency of 240 MHz. The results show that the proposed algorithm and architecture are able to provide high-quality output in real-time while solving the problems of zigzag and blurred effects caused by the conventional scalers. Yi-Nung Liu, Yi-Chun Lin, Yung-Lin Huang, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | HD video decoding scheme based on mobile heterogeneous system architectureabstractEfficiency for mobile devices becomes important than ever due to the demanding multimedia applications. Take HD video decoding for example, it is getting challenging as the resolution increases. To enhance efficiency, HSA (Heterogeneous System Architecture) is proposed to drive a new era of computation. Based on this idea, we develop a heterogeneous architecture, which is composed of a mobile GPU (Graphics Processing Unit) with a proposed configurable filtering unit and a CPU to coordinate the decoding flow. The feasibility of H.264/MPEG-4 HD video decoding pipeline for the proposed architecture is verified. Furthermore, according to the corresponding GPU hardware model, the estimated processing time of inverse quantization, inverse transform and motion compensation for decoding a 1080p video can reach to 16ms and 11ms per frame for MPEG-4 and H.264 respectively. Yu-Jung Chen, Hsin-Fang Wu, Shao-Yi Chien |
ICASSP | 5 |
| 2013 | Quantization error reduction in depth mapsabstractSince most depth maps are quantized to 8-bit numbers in current 3D video systems, the induced cardboard effects can disturb human perception. Moreover, depth maps with larger resolution suffer more from the quantization error. Therefore, this paper proposes an optimization approach to reduce the depth quantization error with well-preserved structure of the depth maps. The experimental results demonstrate that the proposed approach can successfully recover the structure characteristics from the quantized depth maps. Evaluation in mean square error (MSE) and mean structural similarity index (MSSIM) also strongly support our theory and algorithm. Through enhancing the quality of the depth maps from the very beginning, this work can benefit most 3D processing applications, such as 3D modeling, shape registration, and view synthesis. Ku-Chu Wei, Yung-Lin Huang, Shao-Yi Chien |
ICASSP | 3 |
| 2013 | Efficient view synthesis scheme with ray casting and pull-push techniquesabstractView synthesis, composed of depth-image-based rendering followed by hole-filling, is a crucial technology for 3D TV and free-viewpoint TV. To realize view synthesis in practical systems, the efficiency of view synthesis must be considered to achieve good a trade-off between the image quality and the computational complexity. We propose a efficient view synthesis scheme that, when compared to state-of-the-art backward warping, requires only half of the runtime with comparable quality. Specifically, the proposed scheme uses ray casting and pull-push processing to render in one pass, which can be regarded as applying 3D filters in the depth-image-based rendering. Moreover, the proposed scheme can benefit the hole-filling process to further improve the efficiency of view synthesis. Ku-Chu Wei, Yung-Lin Huang, Shao-Yi Chien |
ICME | 3 |
| 2013 | Low-complexity feedback-channel-free distributed video coding with enhanced classifierabstractDistributed video coding (DVC) is an emerging video coding paradigm due to its flexibility to introduce much lower encoding complexity than conventional predictive codecs, which is beneficial for some applications such as wireless video surveillance, wireless sensor networks, and disposable video cameras. Although DVC systems without feedback channel address a wider range of applications, it is not commonly discussed in literatures due to its lower coding performance. In this paper, on the basis of PRISM DVC architecture, a low-complexity feedback-channel-free DVC system is proposed with a new classifier to improve the coding performance. Experimental results show that the proposed system can provide about 0.8 dB gain in PSNR to the state-of-the-art and 2 dB gain to IST-PRISM. It is also competitive regarding other feedback-channel-free DVC systems and can even achieve similar performance to DVC systems with feedback channel for low-motion sequences. Yuh-Jiun Wang, Szu-Lu Hsu, Teng-Yuan Cheng, Chia-han Lee, Shao-Yi Chien |
ISCAS | 5 |
| 2013 | Algorithm adaptive video deinterlacing using self-validation frameworkabstractDeinterlacing is well known as an ill-posed problem of video restoration. In this paper, a self-validation framework is proposed to solve the problem. First, various algorithms with different assumptions will be performed to generate candidate results. Then, a method called double interpolation is applied to test the consistency of each algorithm. Finally, the value of each missing pixel will be chosen among the results of those algorithms according to their performances in the previous test. By doing so, each situation can be handled by the most suitable algorithm and the overall result can be better than any of them. Experimental results demonstrate the validity of the proposed framework in both subjective and objective assessments. Ting-Chun Wang, Yi-Nung Liu, Shao-Yi Chien |
ISCAS | 3 |
| 2013 | Video Object Segmentation and Tracking Framework With Improved Threshold Decision and Diffusion DistanceabstractVideo object segmentation and tracking are two essential building blocks of smart surveillance systems. However, there are several issues that need to be resolved. Threshold decision is a difficult problem for video object segmentation with a multi-background model. In addition, some conditions make robust video object tracking difficult. These conditions include nonrigid object motion, target appearance variations due to changes in illumination, and background clutter. In this paper, a video object segmentation and tracking framework is proposed for smart cameras in visual surveillance networks with two major contributions. First, we propose a robust threshold decision algorithm for video object segmentation with a multi-background model. Second, we propose a video object tracking framework based on a particle filter with the likelihood function composed of diffusion distance for measuring color histogram similarity and motion clue from video object segmentation. The proposed framework can track nonrigid moving objects under drastic changes in illumination and background clutter. Experimental results show that the presented algorithms perform well for several challenging sequences, and our proposed methods are effective for the aforementioned issues. Shao-Yi Chien, Wei-Kai Chan, Yu-Hsiang Tseng, Hong-Yuh Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Brain-Inspired Framework for Fusion of Multiple Depth Cuesabstract2-D-to-3-D conversion is an important step for obtaining 3-D videos, as a variety of monocular depth cues have been explored to generate 3-D videos from 2-D videos. As in a human brain, a fusion of these monocular depth cues can regenerate 3-D data from 2-D data. By mimicking how our brains generate depth perception, we propose a reliability-based fusion of multiple depth cues for an automatic 2-D-to-3-D video conversion. A series of comparisons between the proposed framework and the previous methods is also presented. It shows that significant improvement is achieved in both subjective and objective experimental results. From the subjective viewpoint, the brain-inspired framework outperforms earlier conversion methods by preserving more reliable depth cues. Moreover, an enhancement of 0.70-3.14 dB and 0.0059-0.1517 in the perceptual quality of the videos is realized in terms of the objective-modified peak signal-to-noise ratio and disparity distortion model, respectively. Chung-Te Li, Yen-Chieh Lai, Chien Wu, Sung-Fang Tsai, Tung-Chien Chen, Shao-Yi Chien, Liang-Gee Chen |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2013 | Perceptual Quality-Regulable Video Coding System With Region-Based Rate Control SchemeabstractIn this paper, we discuss a region-based perceptual quality-regulable H.264 video encoder system that we developed. The ability to adjust the quality of specific regions of a source video to a predefined level of quality is an essential technique for region-based video applications. We use the structural similarity index as the quality metric for distortion-quantization modeling and develop a bit allocation and rate control scheme for enhancing regional perceptual quality. Exploiting the relationship between the reconstructed macroblock and the best predicted macroblock from mode decision, a novel quantization parameter prediction method is built and used to achieve the target video quality of the processed macroblock. Experimental results show that the system model has only 0.013 quality error in average. Moreover, the proposed region-based rate control system can encode video well under a bitrate constraint with a 0.1% bitrate error in average. For the situation of the low bitrate constraint, the proposed system can encode video with a 0.5% bit error rate in average and enhance the quality of the target regions. Guan-Lin Wu, Yu-Jie Fu, Sheng-Chieh Huang, Shao-Yi Chien |
IEEE Trans. Image Process. | 4 |
| 2012 | Power optimization of wireless video sensor nodes in M2M networksabstractLow-power wireless video sensor nodes play important roles for applications in machine-to-machine (M2M) network. Several design issues to optimize the power consumption of a video sensor node are addressed in this paper. For the video coding engine selection, the comparison between conventional video coding system and distributed video coding (DVC) system shows that although the rate-distortion performance of existing DVC codec still has room to improve, it can provide lower power consumption with a noisy transmission channel. Furthermore, it also demonstrated that video analysis unit can help to filter out video contents without event-of-interest to reduce transmission power. Finally, several future research directions are addressed, and the trade-off between the video analysis unit, video coding unit, and data transmission should be further studied to design wireless video sensors with optimized power consumption. Shao-Yi Chien, Teng-Yuan Cheng, Chieh-Chuan Chiu, Pei-Kuei Tsung, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen |
ASP-DAC | 1 |
| 2012 | Hybrid distributed video coding with frame level coding mode selectionabstractDistributed video coding (DVC), a new video coding paradigm based on Slepian-Wolf and Wyner-Ziv theories, is a promising solution for implementing low-power and low-cost distributed wireless video sensors since most of the computation load is moved from the encoder to the decoder. It has been showed that there is still room to improve the coding efficiency of the current DVC codec. In this paper, we propose a hybrid coding structure with frame-level coding mode selection (CMS) to allow the DVC encoder to flexibly choose channel coding or entropy coding to code each band. The experimental results show the significant improvement on the rate-distortion (R-D) performance with only slight increase in the encoding complexity compared to the DISCOVER codec. The proposed DVC system performs comparably to H.264 No Motion with much lower encoding complexity. Chieh-Chuan Chiu, Shao-Yi Chien, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen |
ICIP | 2 |
| 2012 | Sampling Technique Analysis of Nyström Approximation in Pixel-Wise Affinity MatrixabstractSpectral graph methods are widely employed in image segmentation, and they exhibit excellent performance. However, for high-resolution images, it is impractical to directly calculate the eigenvectors of the affinity matrix owing to the high computational requirements. The Nystrom method provides an efficient way to approximate the large-scale affinity matrix by low-rank approximation. In the machine learning field, previous studies have mainly focused on less data points with high dimensional features. To the best of our knowledge, this is the first study to discuss the performance of sampling methods for Nystrom approximation, in which we focus on the pixel-wise affinity matrix for a single image. In this paper, we propose a mean-shift segmentation-based Nystrom sampling technique for image analysis. The experimental results show that for images with simple compositions and backgrounds, k-means sampling performs better, whereas for images with more complicated compositions and backgrounds, the proposed method can perform better. Chieh-Chi Kao, Jui-Hsin Lai, Ja-Ling Wu, Shao-Yi Chien |
ICME | 4 |
| 2012 | Color Filter Array Demosaicking Using Self-validation FrameworkabstractColor demosaicking is well known as an ill-posed problem of sensor image restoration. In this paper, a self-validation framework for color demosaicking is proposed. In the proposed self-validation framework, multiple algorithms under different hypotheses will be performed to generate multiple candidates. Then the final estimation of a missing color sample will be decided by evaluating the local consistency of each algorithm with double interpolation. With this framework, the strengths of different algorithms can be combined and thus eliminate color artifacts. Experimental results demonstrate that the proposed framework can improve the image quality in both subjective and objective measures. Ting-Chun Wang, Yi-Nung Liu, Shao-Yi Chien |
ICME | 3 |
| 2012 | System Design of Perceptual Quality-Regulable H.264 Video EncoderabstractIn this work, a perceptual quality-regulable H.264 video encoder system has been developed. Exploiting the relationship between the reconstructed macro block and its best predicted macro block from mode decision, a novel quantization parameter prediction method is built and used to regulate the video quality according to a target perceptual quality. An automatic quality refinement scheme is also developed to achieve a better usage of bit budget. Moreover, with the aid of salient object detection, we further improve the quality on where human might focus on. The proposed algorithm achieves better bit allocation for video coding system by changing quantization parameters at macro block level. Compared to JM reference software with macro block layer rate control, the proposed algorithm achieves better and more stable quality with higher average SSIM index and smaller SSIM variation. Guan-Lin Wu, Yu-Jie Fu, Shao-Yi Chien |
ICME | 3 |
| 2012 | Stable Pose Estimation with a Motion Model in Real-Time ApplicationabstractEstimation of a object pose from camera is a well-developing topic in computer vision. In theory, the pose from a calibrated camera can be uniquely determined. But in practice, most of the real-time pose estimation algorithms suffer from pose ambiguity due to low accuracy of the target object. We think that pose ambiguity¡Xtwo distinct local minima of the according error function¡Xexist because of the phenomenon of geometric illusions. Both of the ambiguous poses are plausible. After obtaining the solution of two minima (pose candidates), we develop a real-time algorithm for stable pose estimation of a target objects with a motion model. In the experimental results, the proposed algorithm diminish the significance of pose jumping and pose jittering effectively. To the best of our knowledge, this is the first work to solve the pose ambiguity problem with motion model in real-time application. Po-Chen Wu, Jui-Hsin Lai, Ja-Ling Wu, Shao-Yi Chien |
ICME | 4 |
| 2012 | Universal embedded compression engine for LCD TV system-on-a-chip with Band-Expansion Progressive Wavelet CodingabstractFor high-resolution and high-frame-rate LCD TVs, the large memory bandwidth and large off-chip memory size become the bottleneck of efficient SoC design. To solve this problem, a universal embedded compression engine is proposed in this paper. A new algorithm named as Band-Expansion Progressive Wavelet Coding (BE-PWC) is first proposed to provide two coding modes, line and block modes, with precise rate control ability. Next, the associate hardware architectures are also proposed to achieve high throughput and low latency. Implementation results show that it can achieve high throughput requirements of 1080p at 120fps with only 143k in logic gate count under the working frequency of 200MHz. Keng-Hsien Huang, Shao-Yi Chien |
ISCAS | 2 |
| 2012 | TCU: Thread compaction unit for GPGPU applications on mobile graphics hardwareabstractThread divergence frequently occurs in various GPGPU applications while parallel threads encounter branching, especially for multimedia algorithms involving high-level image processing and computer vision algorithms. Either predicate or branch instruction degrades the parallel processing performance as a result of synchronization cost. In this work, we propose a configurable thread compaction unit (TCU) to relieve such execution overhead. Through early-stage compacting divergent threads by evaluating a compacting function with a compacting map, TCU can prevent redundant executions caused by predicate instruction, or repetitively invoking processors to fetch instructions and validate effective branches. Our simulation results show that, with TCU, GPUs can improve up to 24.5x and 1.8x performance for Viola-Jones face detection framework compared to predicate and branch instruction. Furthermore, 4.3x and 1.4x improvement in salient region linear feature extraction can be achieved as well. Finally, cache issue in TCU architecture is also discussed in detail. Yu-Jung Chen, Pai-Shun Ting, Meng-Lin Yu, Shao-Yi Chien |
MMSP | 5 |
| 2012 | Combination of SSIM and JND with content-transition classification for image quality assessmentabstractImage quality assessment (IQA) is a crucial feature of many image processing algorithms. The state-of-the-art IQA index, the structural similarity (SSIM) index, has been able to accurately predict image quality by assuming that the human visual system (HVS) separates structural information from non-structural information in a scene. However, the precision of SSIM is relatively lacking when used to access blurred images. This paper proposes a novel metric of image quality assessment, the JND-SSIM, which adopts the just-noticeable difference (JND) algorithm to differentiate between plain, edge, and texture blocks and obtain a visibility threshold map. Based on varying block transition types between the reference and distorted image, SSIM values are assigned respective weights and scaled down by visibility threshold map. We then test our algorithm on the LIVE and TID Image Quality Database, thereby demonstrating that our improved IQA index is much closer to human opinion. Ming-Chung Hsu, Guan-Lin Wu, Shao-Yi Chien |
VCIP | 3 |
| 2012 | Content-adaptive inverse tone mappingabstractTone mapping is an important technique used for displaying high dynamic range (HDR) content on low dynamic range (LDR) devices. On the other hand, inverse tone mapping enables LDR content to appear with an HDR effect on HDR displays. The existing inverse tone mapping algorithms usually focus on enhancing the luminance in over-exposed regions with less (or even no) effort on the process of the wellexposed regions. In this paper, we propose an algorithm with not only enhancement in the over-exposed regions but also in the remaining well-exposed regions. This paper provides an ”histogram-based” method for inverse tone mapping. The proposed algorithm contains a content-adaptive inverse tone mapping operator, which has different responses with different scene characteristics. Scene classification is included in this algorithm to select the environment parameters. Lastly, enhancement of the over-exposed regions, which reconstructs the truncated information, is performed. Pin-Hung Kuo, Chi-Sun Tang, Shao-Yi Chien |
VCIP | 3 |
| 2012 | Region-Based perceptual quality regulable bit allocation and rate control for video coding applicationsabstractIn this paper, a perceptual quality regulable H.264 video encoder system has been developed. We use structure similarity index as the quality metric for distortion-quantization modeling and develop a bit allocation and rate control scheme for enhancing regional perceptual quality. Exploiting the relationship between the reconstructed macroblock and its best predicted macroblock from mode decision, a novel quantization parameter prediction method is built and used to regulate the video quality of the processing macroblock according to a target perceptual quality. Experimental results show that the model can achieve high accurate. Compared to JM reference software with macroblock layer rate control, the proposed encoding system can effectively enhance perceptual quality for target video regions. Guan-Lin Wu, Yu-Jie Fu, Shao-Yi Chien |
VCIP | 3 |
| 2012 | Semantic scalability using tennis videos as examples
Jui-Hsin Lai, Shao-Yi Chien |
Multim. Tools Appl. | 2 |
| 2012 | Low-Decoding-Latency Buffer Compression for Graphics Processing UnitsabstractPower consumption is the key design factor for graphics processing units (GPUs), especially for mobile applications. The increasing bandwidth required to produce more realistic graphics is a major power draw. To address this factor, in this paper, we present a new universal buffer compression method that can handle both color and depth data with the same hardware unit. In contrast to the current state-of-art technologies, which mainly focus on achieving higher and higher compression ratios but discarded the decompression latency, our method reaches a good compromise between the two, which are factors critical to system performance. With spatial prediction and bitstream rearrangement, the data dependencies between different samples are reduced, which enables a parallel decoding process and makes the proposed system have 6.78 times lower decoding latency. Moreover, by adopting a similar concept for the color/depth compression in the DXT5 texture compression method, better quality in terms of PSNR can be achieved without introducing any decoding latency when retrieving a texel. Shao-Yi Chien, Ka-Hang Lok, Yen-Chang Lu |
IEEE Trans. Multim. | 1 |
| 2012 | Tennis Real PlayabstractTennis Real Play (TRP) is an interactive tennis game system constructed with models extracted from videos of real matches. The key techniques proposed for TRP include player modeling and video-based player/court rendering. For player model creation, we propose the process for database normalization and the behavioral transition model of tennis players, which might be a good alternative for motion capture in the conventional video games. For player/court rendering, we propose the framework for rendering vivid game characters and providing the real-time ability. We can say that image-based rendering leads to a more interactive and realistic rendering. Experiments show that video games with vivid viewing effects and characteristic players can be generated from match videos without much user intervention. Because the player model can adequately record the ability and condition of a player in the real world, it can then be used to roughly predict the results of real tennis matches in the next days. The results of a user study reveal that subjects like the increased interaction, immersive experience, and enjoyment from playing TRP. Jui-Hsin Lai, Chieh-Li Chen, Po-Chen Wu, Chieh-Chi Kao, Min-Chun Hu 0001, Shao-Yi Chien |
IEEE Trans. Multim. | 6 |
| 2012 | Preference-Aware View Recommendation System for Scenic Photos Based on Bag-of-Aesthetics-Preserving FeaturesabstractIn this paper, the framework for a real-time view recommendation system is proposed. The proposed system comprises two parts: offline aesthetic modeling stage and efficient online aesthetic view finding process. A preference-aware aesthetic model is proposed to suggest views according to varied user-favorite photographic styles, where a bottom-up approach is developed to construct an aesthetic feature library with bag-of-aesthetics-preserving features instead of top-down methods that implement the heuristic guidelines (rule-specific features) listed in photography literatures, which is employed in previous works. A collection of scenic photos is used as the test set; however, the proposed method can be employed to other types of photo collection according to different application scenarios. The proposed model can cover both implicit and explicit aesthetic features and can adapt to users' preferences with a learning process. In the second part, the learned model is employed in a view finder to help the user to locate the most aesthetic view while taking a photograph. The experimental results show that the proposed features in the library (92.06% in accuracy) outperform the state-of-the-art rule-specific features (83.63% in accuracy) significantly in the photo aesthetic quality classification task, and the rule-specific features are also proved to be encompassed by the proposed features. Meanwhile, it is observed from experiments that the features extracted for contrast information are more effective than those for absolute information, which is consistent with the properties of human visual systems. Furthermore, the user studies for the view recommendation task confirm that the suggested views are consistent with users' preferences (81.25% agreements). Hsiao-Hang Su, Tse-Wei Chen 0001, Chieh-Chi Kao, Winston H. Hsu, Shao-Yi Chien |
IEEE Trans. Multim. | 5 |
| 2012 | Visual Vocabulary Processor Based on Binary Tree Architecture for Real-Time Object Recognition in Full-HD ResolutionabstractFeature matching is an indispensable process for object recognition, which is an important issue for wearable devices with video analysis functionalities. To implement a low-power SoC for object recognition, the proposed visual vocabulary processor (VVP) is employed to accelerate the speed of feature matching. The VVP can transform hundreds of 128-D SIFT vectors into a 64-D histogram for object matching by using the binary-tree-based architecture, and 16 calculators for the computations of the Euclidean distances are designed for each of the two processors in each level. A total of 126 visual words can be saved in the six-level hierarchical memory, which instantly offers the data required for the matching process, and more than 5 times of bandwidth can be saved compared with the non-binary-tree-based architecture. As a part of the recognition SoC, the VVP is implemented with the 65-nm CMOS technology, and the experimental results show that the gate count and the average power consumption are 280 K and 5.6 mW, respectively. Tse-Wei Chen 0001, Yu-Chi Su, Keng-Yen Huang, Yi-Min Tsai, Shao-Yi Chien, Liang-Gee Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2011 | New optimization scheme for L2-norm total variation semi-supervised image soft labelingabstractIn image/video context processing, such as clustering, matting, or further editing and context-aware enhancement, the probability model on the basis of Markov property is usually employed, where the neighbors around the center have stronger connection. To realize the optimization of such probability models encounters to solve a large linear system under the objective functional of L2-norm total variation (TV). The existing feasible methods can deal with the problems with small or very large neighborhood, but there lacks of feasible method for solving linear system with intermediate neighborhood in an efficient and accurate way. In this paper, based on the theoretical analysis, we transform the optimization problem to a process with accumulated joint bilateral filtering. Both efficiency and accuracy are achieved with appropriate prove of validation. Finally, taking image soft segmentation as an example, the proposed optimization scheme is implemented on GPU with existing fast bilateral filter to show the feasibility. Chia-Liang Tsai, Shao-Yi Chien |
ICIP | 2 |
| 2011 | Coarse-to-fine temporal optimization for video retargeting based on seam carvingabstractIn this paper, a new video retargeting method based on temporal information and seam carving is presented. Two video energy functions, motion weight prediction and pixel-based optimization, are proposed to take the temporal information into account and make dynamic programming available during the process of retargeting. The motion weight prediction exploits both the block-based motion estimation and Gaussian masks to predict the coarse location of seams in the current frame and reduce the search range of dynamic programming. The pixel-based optimization then utilizes the concept of pixel-based optical flow to explore better temporal relations between the current frame and previous frames in the reduced search range. The experimental results show that combining these two video energy functions as well as dynamic programming, the proposed method could achieve content-aware and temporal smoothing retargeting results with less computational complexity. Wei-Lun Chao, Hsiao-Hang Su, Shao-Yi Chien, Winston H. Hsu, Jian-Jiun Ding |
ICME | 3 |
| 2011 | Automatic object segmentation with salient color modelabstractImage segmentation is a well-developing topic in the image processing, and a number of previous works have been proposed and achieved high performance. However, most previous works needed user-assistance to provide the prior information of the target object in the segmentation. In this paper we propose an unsupervised scheme, combining the salient object detection and segmentation method, to segment the target object without any prior information from users. The experimental results show that the proposed salient color model derived with salient features can provide a prior information with high confidence to generate precise segmentation automatically. The proposed color model of salient objects can not only be applied with Min-Cut algorithm, but also extended to more segmentation algorithms, like matting or non-parametric model. Chieh-Chi Kao, Jui-Hsin Lai, Shao-Yi Chien |
ICME | 3 |
| 2011 | Architecture design and analysis of image-based rendering engineabstractImage-based rendering (IBR) is a technique to render the video from images, and it provides users to have more interaction and immersive experience in watching a video. In this paper, we integrate the computation of several IBR applications, analyze the bandwidth of memory access, and design an architecture to process the computation of IBR. Experimental results show that the proposed IBR Engine is able to render a video with resolution 720×480 and 30 frames per second, which is 12.7 times faster than a Core2Due 2.83 GHz CPU. For the extensions, IBR Engine can be embedded in the television system and lets viewers enjoy the functions from IBR. Jui-Hsin Lai, Chieh-Li Chen, Shao-Yi Chien |
ICME | 3 |
| 2011 | Tennis real play: an interactive tennis game with models from real videosabstractTennis Real Play (TRP) is an interactive tennis game system constructed with models extracted from videos of real matches. The key techniques proposed for TRP include player modeling and video-based player/court rendering. For player model creation, we propose a database normalization process and a behavioral transition model of tennis players, which might be a good alternative for motion capture in the conventional video games. For player/court rendering, we propose a framework for rendering vivid game characters and providing the real-time ability. We can say that image-based rendering leads to a more interactive and realistic rendering. Experiments show that video games with vivid viewing effects and characteristic players can be generated from match videos without much user intervention. Because the player model can adequately record the ability and condition of a player in the real world, it can then be used to roughly predict the results of real tennis matches in the next days. The results of a user study reveal that subjects like the increased interaction, immersive experience, and enjoyment from playing TRP. Jui-Hsin Lai, Chieh-Li Chen, Po-Chen Wu, Chieh-Chi Kao, Shao-Yi Chien |
ACM Multimedia | 5 |
| 2011 | Scenic photo quality assessment with bag of aesthetics-preserving featuresabstractIn this paper, an aesthetic modeling method for scenic photographs is proposed. A bottom-up approach is developed to construct an aesthetic library with bag-of-aesthetics preserving features instead of top-down methods that implement the heuristic guidelines (rule-specific features) listed in the photography literature, which is employed in previous works. The proposed method can cover both implicit and explicit aesthetic features with a learning process. The experimental results show that the proposed features in the library (92.06% in accuracy) outperform the state-of-the-art rule-specific features (83.63% in accuracy) significantly in the aesthetic quality assessment for scenic photos, and the rule-specific features are also proved to be encompassed by the proposed features. Meanwhile, it is observed from experiments that the features extracted for contrast information are more effective than those for absolute information, which is consistent with the properties of human visual systems. Hsiao-Hang Su, Tse-Wei Chen 0001, Chieh-Chi Kao, Winston H. Hsu, Shao-Yi Chien |
ACM Multimedia | 5 |
| 2011 | Distributed video coding: A promising solution for distributed wireless video sensors or not?abstractLow-power and low-cost distributed wireless video sensors play important roles for applications in machine-to-machine (M2M) and wireless sensor networks. Distributed video coding (DVC), an emerging coding technology based on Wyner-Ziv theory, seems to be a possible solution for implementing low-power video sensors since most of the computational complexity is moved from the encoder to the decoder. In this paper, existing works on DVC are discussed with rate-distortion and power consumption analyses compared with H.264/AVC-based approaches. We show that, since more transmission power is required for compensating the lower rate-distortion performance, the power consumption of sensor nodes using DVC is just similar to that of using H.264/AVC with zero motion vectors. Therefore, there is still a room for improvement to make DVC applicable for distributed wireless video sensors. Based on our analysis results, several possible research directions, such as studies on the trade- off between hardware cost and system power consumption, are also addressed in this paper under a unified DVC framework. Chieh-Chuan Chiu, Shao-Yi Chien, Chia-han Lee, V. Srinivasa Somayazulu, Yen-Kuang Chen |
VCIP | 2 |
| 2011 | Tennis Video 2.0: A new presentation of sports videos with content separation and rendering
Jui-Hsin Lai, Chieh-Li Chen, Chieh-Chi Kao, Shao-Yi Chien |
J. Vis. Commun. Image Represent. | 4 |
| 2011 | Algorithm and Architecture Design of Image Inpainting Engine for Video Error Concealment ApplicationsabstractError concealment techniques can improve subjective video quality in a decoder when the video bitstream is corrupted during transmission. In this paper, to achieve perceptually pleasant results, an image inpainting technique in which structure information generated from edge information is adopted as the spatial error concealment method. In addition, a modified boundary matching algorithm for temporal error concealment is proposed for temporal frames. To maintain low hardware costs as regards the error concealment engine, the processing iteration number of each macroblock is limited to four based on the proposed inpainting algorithm. Block-based pipeline scheduling is also proposed to reduce the number of processing cycles and the on-chip memory size. Moreover, a cache-based data reuse scheme is developed to reduce the processing cycles and external bandwidth. Moreover, the two concealment modes share the same computational core to reduce hardware costs. A prototype chip is implemented by using the UMC 90 nm process. The total gate count is approximately 121 k at 200 MHz. The maximum processing capability can support 244.8 k macroblocks per second or 1920 × 1080 4:2:0 30 Hz video. The core size is 1.30 × 1.30 mm2. The average power dissipation is 131.4 mW at 200 MHz. Compared to other error concealment methods, the proposed design can achieve better perceptual quality at an acceptable additional hardware cost. Guan-Lin Wu, Ching-Yi Chen, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Photo Retrieval Based on Spatial Layout with Hardware Acceleration for Mobile DevicesabstractA new photo retrieval system for mobile devices is proposed. The system can be used to search for photos with similar spatial layouts efficiently, and it adopts an image segmentation algorithm that extracts features of image regions based on K-Means clustering. Since K-Means is computationally intensive for real-time applications and prone to generate clustering results with local optima, parallel hardware architectures are designed to meet the real-time requirement of the retrieval process. Experiments show that the proposed algorithm in the photo retrieval system obtains better mean average precision than other methods, and it is tested with image recognition problems. The robustness of the algorithm is also evaluated with noise and image blurring. Besides, the proposed K-Means hardware can provide a trade-off between the execution time and the retrieval performance on the software and hardware cosimulation platform. The contribution of this work is twofold. The first is the development of a photo retrieval framework for mobile devices, where a new texture feature is employed in the algorithm to enhance the retrieval performance. The other is the integration of the K-Means hardware accelerator and the photo retrieval system. The hardware architecture is analyzed, and the specifications are compared with previous works. Tse-Wei Chen 0001, Yi-Ling Chen 0008, Shao-Yi Chien |
IEEE Trans. Mob. Comput. | 3 |
| 2011 | Algorithm and Architecture Design of Perception Engine for Video Coding ApplicationsabstractIn image and video coding field, an effective compression algorithm should remove not only the spatial, temporal, and statistical redundancy but also the perceptual redundancy information from the pictures. Many perceptual models are presented in the literature to cooperate with video coding system to obtain significant bit rate reduction without perceptual distortion. One of the critical issues for those perceptual models is their high computational complexity to apply to real-time applications. To alleviate this problem, this paper aims at hardware architecture design of perception engine for video coding applications. The adopted perceptual models include the structural similarity model, visual attention models, and just-noticeable-distortion model, and contrast sensitivity function. Moreover, those models are further developed and modified to be suitable for hardware implementation. Macroblock-based processing with data reuse scheme is used to save the system bandwidth. The architecture of parallel processing for each visual model with sharing the on-chip memory and buffers is developed to reduce the chip area. Subjective experiment results show that the adopted model achieves about 7%-41% bit-rate saving in the QP range of 24-36 without visual quality degradation. For the hardware implementation of the perception engine, the chip is taped out using 0.18 m technology. The chip size is about 3.3 3.3 mm , and the power consumption is 83.9 mW. The processing capability is HDTV720p. Guan-Lin Wu, Tung-Hsing Wu, Shao-Yi Chien |
IEEE Trans. Multim. | 3 |
| 2011 | Flexible Hardware Architecture of Hierarchical K-Means Clustering for Large Cluster NumberabstractK-Means is an important clustering algorithm that is widely applied to different applications, including color clustering and image segmentation. To handle large cluster numbers in embedded systems, a hardware architecture of hierarchical K-Means (HK-Means) is proposed to support a maximum cluster number of 1024. It adopts 10 processing elements for the Euclidean distance computations and the level-order binary-tree traversal. Besides, a hierarchical memory structure is integrated to offer a maximum bandwidth of 1280 bit/cycle to processing elements. The experiments show that applications such as video segmentation and color quantization can be implemented based on the proposed HK-Means hardware. Moreover, the gate count of the hardware is 414 K, and the maximum frequency achieves 333 MHz. It supports the highest cluster number and has the most flexible specifications among our works and related works. Tse-Wei Chen 0001, Shao-Yi Chien |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2010 | Edge-adaptive image segmentation based on seam processing and K-Means clusteringabstractA new image segmentation method is proposed to combine the edge information with the feature-space method, K-Means clustering. A procedure called seam processing, which is computationally efficient, is employed to search for horizontal and vertical seams that contain edge information. By transforming the spatial coordinates based on the seam detection results, the edge information can be added to the feature vectors, which are the inputs of K-Means algorithm. The experiments show that the proposed method can achieve edge-adaptive segmentation results, which can not be obtained using traditional methods based on K-Means clustering. Tse-Wei Chen 0001, Hsiao-Hang Su, Yi-Ling Chen 0008, Shao-Yi Chien |
ICIP | 4 |
| 2010 | Vivid tennis player rendering system using broadcasting game videosabstractImage-based rendering has been highly developed for its wide applications such as view synthesis and special effects in movies. In this paper, we proposed a tennis player rendering system synthesizing diverse player action/motion based on extracted database from broadcasting game videos. The system gathers database by retrieving the player from videos and synthesizes various kinds of player action/motion according to the user's instructions. The results show that the proposed rendering system can render smooth action/motion transition with satisfactory visual effect. For further applications, the proposed system can be used in interactive tennis games with image textures. Chieh-Li Chen, Jui-Hsin Lai, Shao-Yi Chien |
ICME | 3 |
| 2010 | Low latency universal buffer compression and decompression for mobile graphics applicationsabstractDue to the thirst for bandwidth resource and the limitation of device size, mobile GPUs face more challenges than their desktop counterparts do when seeking realistic visual quality. To alleviate the severely restricted situation mobile GPUs in, we propose a new universal buffer compression method to handle both color and depth data with the same hardware unit. While the previous state-of-the-arts mainly focus on achieving high compression ratio but leave the decompression latency out of consideration, our method tends to reach a good compromise between two of them. Moreover, by adopting similar concept used in color/depth compression to DXT5 texture compression method, better quality can be achieved without introducing any decoding latency in retrieving a texel. Ka-Hang Lok, Yen-Chang Lu, Shao-Yi Chien |
ICME | 3 |
| 2010 | Perception-aware H.264/AVC encoder with hardware perception analysis engineabstractIn order to increase coding efficiency, a perception-aware H.264/AVC encoder is proposed in this paper. With a different perception models for intra-frames and inter-frames, the initial quantization parameter (QP) is perceptually adjusted to remove perceptual redundancy. Moreover, the associated hardware architectures of the whole encoder and the perception analysis engine are also proposed with hardware sharing and pipelining design techniques. Experimental results show that our encoder achieves about 6-27% bit-rate saving in the QP range of 28-36. Furthermore, the hardware cost of the perception analysis engine is 120K gates, which is acceptable to be integrated into an H.264/AVC encoder chip. Guan-Lin Wu, Tung-Hsing Wu, Yu-Jie Fu, Shao-Yi Chien |
ICME | 4 |
| 2010 | Support Vector Machines on GPU with Sparse Matrix FormatabstractEmerging general-purpose Graphics Processing Unit (GPU) provides a multi-core platform for wide applications, including machine learning algorithms. In this paper, we proposed several techniques to accelerate Support Vector Machines (SVM) on GPUs. Sparse matrix format is introduced into parallel SVM to achieve better performance. Experimental results show that the speedup of 55x-133.8x over LIBSVM can be achieved in training process on NVIDIA GeForce GTX470. Tsung-Kai Lin, Shao-Yi Chien |
ICMLA | 2 |
| 2010 | Image information splitting framework with importance sampling for robust transmissionabstractIn the traditional compression schemes of visual multimedia contents, such as video and image, scalable compression is usually achieved by splitting the information into one base layer and several enhanced layers. Such principle for scalable compression encounters a problem in the presence of transmission loss that the unpredictable loss in the base layer causes tremendous reduction of recognizable information by human perception. Multiple Description Coding (MDC) is defined to be robust in the presence of transmission loss and several MDC schemes have been proposed as extensions of the existing image and video compression standards. In this paper, a new presentation of image information based on the principle of MDC is proposed. The image is presented as samples whose density is variable according to the frequency energy. Moreover, the samples are able to be split into several partitions, promising that a recognizable approximation of the original content can be reconstructed by any partition, and more high frequency information can be reconstructed with more partitions, which fulfills the motivation of MDC. The reconstruction is realized by interpolating samples using normalized radial basis function (RBF) network. The result shows that the proposed framework of image information splitting provides acceptable quality scalability. Chia-Liang Tsai, Shao-Yi Chien |
ISCAS | 2 |
| 2010 | Direction-adaptive image upsampling using double interpolationabstractDouble interpolation quality evaluation can be used as a measurement of an interpolation operation. By using this double interpolation framework, a high efficient direction-adaptive upsampling algorithm is proposed without any threshold setting and post-processing. With the proposed upsampling algorithm, the problem of zigzagging artifacts on the edge no longer exists. Moreover, the proposed algorithm has low computation complexity. The experimental results show that the proposed algorithm has high quality image output. Yi-Chun Lin, Yi-Nung Liu, Shao-Yi Chien |
PCS | 3 |
| 2010 | Efficient Spatial-Temporal Error Concealment Algorithm and Hardware Architecture Design for H.264/AVCabstractThis paper presents an efficient error concealment algorithm for video bitstream over error-prone channel suffering from damage. Moreover, hardware architecture design and chip implementation of the proposed error concealment algorithm are also presented. For spatial error concealment, a mode selection algorithm considering the reuse of intra mode information embedded in bitstream is developed for the adaptation of bilinear and directional interpolation. It suffers only 0.08 dB video quality drop in average but the speedup measured on a general purpose processor is up to 40 times compared with the conventional methods. It is also more suitable for low cost hardware design. For temporal error concealment, the decoded motion vectors of the neighboring blocks of the corrupted macroblock are reused to provide hints to estimate the motion vector of the corrupted macroblock. Moreover, for real-time applications, a data and computational results reuse scheme of motion vector estimation is proposed and 96% computation and memory bandwidth can be reduced compared with the conventional methods with 0.18 dB quality drop in average. With the UMC 90 nm 1P9M process, the proposed error concealment engine can process HDTV1080P 30 frames per second video data and the power consumption is 15.77 mW at 125 MHz operation frequency. Guan-Lin Wu, Ching-Yi Chen, Tung-Hsing Wu, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Bandwidth Adaptive Hardware Architecture of K-Means Clustering for Video AnalysisabstractK-Means is a clustering algorithm that is widely applied in many fields, including pattern classification and multimedia analysis. Due to real-time requirements and computational-cost constraints in embedded systems, it is necessary to accelerate K-Means algorithm by hardware implementations in SoC environments, where the bandwidth of the system bus is strictly limited. In this paper, a bandwidth adaptive hardware architecture of K-Means clustering is proposed. Experiments show that the proposed hardware can be used in applications such as image segmentation, and it has the maximum clock speed 400-MHz and 440-K gate count with TSMC 90-nm technology. Moreover, the throughput of the proposed hardware reaches 16 dimension/cycle, and it can deal with feature vectors with different dimensions using five parallel modes to utilize the input bandwidth efficiently. Tse-Wei Chen 0001, Shao-Yi Chien |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Bandwidth adaptive hardware architecture of K-Means clustering for intelligent video processingabstractK-means is a clustering algorithm that is widely applied in many fields, including pattern classification and multimedia analysis. Due to real-time requirements and computational-cost constraints in embedded systems, it is necessary to accelerate k-means algorithm by hardware implementations in SoC environments, where the bandwidth of the system bus is strictly limited. In this paper, a bandwidth adaptive hardware architecture of k-means clustering is proposed. Experiments show that the proposed hardware has the maximum clock speed 400 MHz with TSMC 90 nm technology, and it can deal with feature vectors with different dimensions using five parallel modes to utilize the input bandwidth efficiently. Tse-Wei Chen 0001, Shao-Yi Chien |
ICASSP | 2 |
| 2009 | Cache-based integer motion/disparity estimation for quad-HD H.264/AVC and HD multiview video codingabstractTo provide more vivid perception, more and more advanced features, like the 4k×2k resolution and the multiview functionality, are emerging for TV. For a multiview video coding (MVC) encoder, motion and disparity estimation (ME/DE) take at least half the hardware requirement. To solve these challenges, a cache-based integer ME/DE algorithm is proposed. With a cache memory as the search window buffer, a predictor-centered ME/DE algorithm is presented. The search range can be reduced to ±16 pixels with less than 0.1dB quality drop compared with full search algorithm. Based on this algorithm, an integer ME/DE chip design is realized. It can reduce 82% on-chip SRAM and 39% system bandwidth. Moreover, the search candidate requirement is also reduced by 79%. As the result, an ME/DE chip design for 4k×2k quad-HD H.264 and HDTV MVC is implemented. Pei-Kuei Tsung, Wei-Yin Chen, Li-Fu Ding, Shao-Yi Chien, Liang-Gee Chen |
ICASSP | 4 |
| 2009 | Coarse-grained reconfigurable image stream processor architecture for embedded image/video processing and analysisabstractTo achieve low-cost low-power high-performance embedded image/ video processing and analysis, in this paper, a coarse-grained reconfigurable image stream processor (CRISP) architecture is proposed. With the reconfigurable datapath and interconnection optimized for image/video processing, it can be viewed as an application-specific reconfigurable device for image/video processing, and several design examples are shown to demonstrate the high efficiency of CRISP architecture. Shao-Yi Chien, Tsung-Huang Chen, Jason C. Chen, Tung-Yuan Cheng, Wei-Kai Chan |
ICME | 1 |
| 2009 | Color filter array demosaicking using joint bilateral filterabstractBilateral filter has shown its outstanding performance in image denoising and other multimedia applications. In this paper, a new color interpolation technique named joint bilateral demosaicking is proposed. Considering the image gradient, an edge-sensing initialization step is performed. In addition, joint bilateral filter exploits the correlation between color channels with the information from initialization while preserving edges. Experimental results show the superiority of the proposed technique in suppressing demosaicking artifacts and producing more visually acceptable images. Meng-Che Chuang, Yi-Nung Liu, Tsung-Huang Chen, Shao-Yi Chien |
ICME | 4 |
| 2009 | Algorithm and architecture design of multi-layer video coding enginewith hybrid scheme for wireless video linksabstractFor a wireless linking system, the bandwidth might be limited. Therefore, it needs a codec to compress the data transmitted in the system. However, tranditional video codecs can not be used directly in the wireless linking system because there are several requirements in the system. First, the codec should support various applications. Second, low latency is required for real-time applications. Third, error robustness is important in the wireless linking system. In this paper, a coding engine with hybrid scheme for wireless video links is proposed. It consists of two codecs which are JPEG-LS and H.264 intra prediction to process various applications and provides two type of bitstreams for wireless bandwidth variation. 20% of hardware cost is reduced by putting mode decision in front in the proposed encoder. The proposed codec is error robust by not using temporal information and has a low latency of 1.84 us by macroblock pipelining architectures. The processing capability of the proposed architecture achieves full-HD 60 fps. Keng-Hsien Huang, Han-Ru Chen, Shao-Yi Chien |
ICME | 3 |
| 2009 | Super-resolution sprite with foreground removalabstractSprite is an image constructed from video clips and is also a medium for multimedia applications. An automatic sprite generation with foreground removal and super-resolution is proposed in this paper. To remove the foreground objects, each pixel-value on the sprite is iteratively updated by the value with maximum appearance probability on temporal and spatial distribution. By storing the half-pixel, superresolution sprite has less blurring-defect from source video. In the result, the generated sprite preserves the complete scenes of background and has higher image quality, and it can used to increase the visual quality in current sprite applications and also employed to facilitate video segmentation. Jui-Hsin Lai, Chieh-Chi Kao, Shao-Yi Chien |
ICME | 3 |
| 2009 | Universal Rasterizer with edge equations and tile-scan triangle traversal algorithm for graphics processing unitsabstractThe rasterization stage in a graphics processing unit (GPU), which consists of triangle setup, rasterization, and parameter interpolation with plane equations, always requires huge operations and is usually the bottleneck of the performance. For real-time applications, a Universal Rasterizer (UR) with edge equations and a tile-scan triangle traversal algorithm are proposed for low cost graphics rendering. In UR, the basic functions for parameter interpolation and rasterization can be executed with a universal shared hardware to reduce the cost. The result shows that it can minimize the processing time of triangle traversal and guarantee no reiteration when traverse. With the hardware sharing and architecture design techniques of pipelining and scheduling, it can achieve the real-time requirements for graphics applications with reasonable hardware cost. Chih-Hao Sun, You-Ming Tsao, Ka-Hang Lok, Shao-Yi Chien |
ICME | 4 |
| 2009 | Bandwidth and Local Memory Reduction of Video Encoders using Bit Plane Partitioning Memory ManagementabstractThis paper presents a new memory management scheme for the reference frame buffer in a video encoder. The proposed BPPMM (Bit Plane Partitioning Memory Management) scheme changes the format of data, and hence we can access different number of the bit-planes of the pixel data on-demand. This BPPMM technique is especially suitable for motion estimation with bit-truncation. Experiments show that when BPPMM scheme is integrated in an H.264 hardware encoder, more than 46% of local SRAM size and 31% of the external memory bandwidth could be reduced with only a little quality degradation. Yi-Nung Liu, Meng-Che Chuang, Shao-Yi Chien |
ISCAS | 3 |
| 2009 | Bio-inspired Perceptual Video Encoding based on H.264/AVCabstractA bio-inspired perceptual video coding algorithm which can reduce the bit rate of the encoded bit stream while maintaining equal perceptual (visual) quality is proposed in this paper. Instead of applying the complicated bio-inspired perceptual model, we propose an algorithm which automatically analyzes the spatial and temporal information in the video sequence, and achieves better bit allocation for video coding systems by changing the quantization parameter (QP) in macroblock (MB) basis. Experiments show that our algorithm achieves about 5-30% bit-rate saving in the QP range of 28-36 without perceptual (visual) quality degradation. Tung-Hsing Wu, Guan-Lin Wu, Shao-Yi Chien |
ISCAS | 3 |
| 2009 | Tennis Video with Semantic ScalabilityabstractScalable video is the research topic to provide different size of video bitstream under different transmission bandwidth. In this paper, the semantic scalability is proposed that provides the scalable videos in semantic domain, and the tennis videos are used as the experiments. Contrary to decreasing the video quality to reduce the bitrates, the lower bitstream size is achieved by abandoning the video contents with less semantic importance. The experimental results show that the proposed semantic scalability provides four levels of the scalable videos and maintains the visual quality in watching the game video. The study of the scalability in semantic domain provides a new aspect for the scalable video. Jui-Hsin Lai, Shao-Yi Chien |
ISM | 2 |
| 2009 | Cooperative Surveillance System with Fixed Camera Object Localization and Mobile Robot Target Tracking
Chih-Chun Chia, Wei-Kai Chan, Shao-Yi Chien |
PSIVT | 3 |
| 2009 | Efficient Content Analysis Engine for Visual Surveillance NetworkabstractIn the next-generation visual surveillance systems, content analysis tools will be integrated. In this paper, to accelerate these tools, it is proposed to integrate a hardware content analysis engine into a smart camera system-on-a-chip (SoC). A smart camera SoC hardware architecture with the proposed visual content analysis engine is first presented. This engine consists of dedicated accelerators and a programmable morphology coprocessor. Stream processing design concept, frame-level pipelining, and subword level parallelism are employed together to efficiently utilize the bandwidth of the system bus and achieve high throughput. The implementation results show that, with 168 K logic gates and 40.63 Kb on-chip memory, a processing speed of 30 640 x 480 frames/s can be achieved, while the operations of video object segmentation, object description and tracking, and face detection and scoring are supported. Wei-Kai Chan, Jing-Ying Chang, Tse-Wei Chen 0001, Yu-Hsiang Tseng, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2009 | High-Quality Mipmapping Texture Compression With Alpha Maps for Graphics Processing UnitsabstractTexture compression is an important technique in graphics processing units (GPUs) for saving memory bandwidth. This paper presents a high-quality mipmapping texture compression (MTC) system with alpha maps. Based upon the wavelet transform, a hierarchical approach is adopted for mipmapping textures in the YCbCr color space and alpha channel. By inspecting the similarity between the alpha and luminance channels, the two channels are efficiently encoded together with linear prediction in the differential mode. In addition, the split mode manages textures with no strong relationship between the alpha and luminance channels. A layer overlapping technique is also proposed to reduce the texture memory bandwidth. Simulation results show that MTC can reduce the texture access traffic by 80% to 90% and provides high image quality as well. Compared with DirectX texture compression (DXTC), the most well-known texture compression with alpha maps, MTC reduces the texture access bandwidth by 30% more. VLSI implementation results show that the hardware cost of MTC is similar to that of DXTC and that MTC is suitable for integration in GPUs to provide high-quality textures with low memory bandwidth requirements. Chih-Hao Sun, You-Ming Tsao, Shao-Yi Chien |
IEEE Trans. Multim. | 3 |
| 2008 | Fast motion estimation with inter-view motion vector prediction for stereo and multiview video codingabstract3-D video will become one of the most important video technologies in the next generation of television. Due to ultra high data bandwidth requirement for 3-D video, effective compression technology becomes an essential part in the infrastructure. Thus stereo and multiview video coding (MVC) plays a critical role. However, MVC systems require much more computational complexity relative to mono-view video coding systems. Therefore, an ef..cient prediction scheme is necessary for encoding. In this paper, a new fast motion estimation (ME) algorithm is proposed. By utilizing disparity estimation (DE) to ..nd corresponding blocks between different views, the coding information such as motion vectors can be effectively shared and reused from the coded view channel. Therefore, the computation for ME in most view channels can be greatly reduced. Experimental results show that compared with the full search block matching algorithm applied to both ME and DE, the proposed algorithm saves 95% computation with near-FSBMA quality. Li-Fu Ding, Pei-Kuei Tsung, Wei-Yin Chen, Shao-Yi Chien, Liang-Gee Chen |
ICASSP | 4 |
| 2008 | Fast fingertip positioning by combining particle filtering with particle random diffusionabstractA new and efficient algorithm to find the fingertips for human computer interface is proposed in this paper.The Fingertip Positioning algorithm of this paper combines particle filtering with Particle Random Diffusion to find the different fingertips quickly and robustly. There are two special methods used in this algorithm, which are Particle Random Diffusion and Fingertip Particle Selection. Without checking every pixel in the image of video sequence, this algorithm works well with low computational cost. This algorithm comes out with good performance in cluttered backgrounds and processes in real time. Based on the accurate positions of fingertips, hand gesture recognition can be done efficiently and robustly. Ko-Jen Hsiao, Tse-Wei Chen 0001, Shao-Yi Chien |
ICME | 3 |
| 2008 | Frame rate up-conversionwith global-to-local iterative motion compensated interpolationabstractVideo frame rate up conversion (FRUC) has many applications, especially it has been proved to effectively reduce the motion blur problem and improve the visual quality on LCD. In this paper, we propose a global-to-local iterative algorithm to resolve the problems occurring in the existing block based algorithms. In addition, the iterative approach, where interpolated blocks are generated in confident order, is employed in order to find the applying order and combination of different motion compensated interpolation(MCI) algorithms. Experimental results show that the proposed algorithm outperforms other MCI algorithms in perceptual and can generate smooth high quality interpolated frames. Kung-Yen Hsu, Shao-Yi Chien |
ICME | 2 |
| 2008 | Content-aware image resizing using perceptual seam carving with human attention modelabstractIn this paper, a new image resizing technique, perceptual seam carving, is proposed. With considering both face map and saliency map as human attention model in the energy function, it can keep important information in perceptual when the image is downsized. Moreover, a switching scheme between seam carving and resampling is also proposed to avoid excessively distorting the images. Experiments show that the proposed algorithm can generate more desirable resized images than cropping, resampling, and conventional seam carving techniques. Daw-Sen Hwang, Shao-Yi Chien |
ICME | 2 |
| 2008 | Spatial-temporal consistent labeling for multi-camera multi-object surveillance systemsabstractFor an intelligent multi-camera multi-object surveillance system, object correspondence across time and space is important to many smart visual applications. In this paper, we propose a temporal and spatial consistent labeling algorithm for this demand. In the algorithm, an object corresponding database records the temporal and spatial consistency information for each segmented mask. With the database, the object-mask correlations are propagated through the propagation rules by analyzing mask splitting/merging conditions. In the spatial consistent labeling method, the homography warping and the earth mover’s distance are adopted to match same objects across different views. The earth mover’s distance solves the double matching problem, allows the algorithm to work normally under a small deviation of detected object locations, and makes pairing results have minimum global matching distances. The concept trusting-former-pairs-more is also adopted to avoid frequent pair switching if two objects are too close. The correct spatial labeling rate is about 89.25% in average. For online processing applications, the algorithm need not trace back to the past frames. The overall processing speed is about 10.24 frame per second (fps) with CIF size video running on a 2.8GHz general purpose CPU. Jing-Ying Chang, Tzu-Heng Wang, Shao-Yi Chien, Liang-Gee Chen |
ISCAS | 3 |
| 2008 | Architectural analyses of K-Means silicon intellectual property for image segmentationabstractK-Means is a clustering algorithm that is widely applied in many fields, including pattern classification, multimedia analysis, and image retrieval. Due to real-time requirements of image segmentation in embedded systems, it is necessary to accelerate K-Means algorithm by hardware implementations. The contribution of this paper includes a series of K-means hardware analyses and a newly proposed SIP for image segmentation in SoC environments. Experiments show that the proposed SIP has the maximum clock speed 200MHz with TSMC 0.18μm technology, and that it can be successfully used for image segmentation on an FPGA board with AMBA AHB. Tse-Wei Chen 0001, Chih-Hao Sun, Jun-Ying Bai, Han-Ru Chen, Shao-Yi Chien |
ISCAS | 5 |
| 2008 | Hardware-oriented image inpainting for perceptual I-frame error concealmentabstractError concealment technique can improve subjective video quality in decoder when the video bitstream is corrupted during transmission. In this paper, a new error concealment algorithm for I-frames is proposed. To achieve perceptually comfortable results, image inpainting algorithm is adopted with structure information generated from edge information. In addition, mixed spatial-temporal inpainting scheme is also proposed for I-frames when the previous frame is available. Moreover, a block-based scheduling is also developed for the proposed algorithm to make it simpler for hardware implementation. Experimental results show that the proposed algorithm outperforms the official H.264 decoder and can deal with various error patterns. It also shows that it is suitable for hardware implementation in terms of on-chip memory requirement and predictable processing cycles. Ching-Yi Chen, Guan-Lin Wu, Shao-Yi Chien |
ISCAS | 3 |
| 2008 | A cost effective reconfigurable memory for multimedia multithreading streaming architectureabstractData flow plays an important factor when design a multimedia multithreading streaming system. In this paper, a reconfigurable memory architecture is proposed to smooth the stream data flow to decrease maximum stream data bandwidth from the external memory. The reconfigurable memory not only can configure as data cache but also can configure as tightly couple memory depending on the application. In addition, heterogenous threads level configurability can adjust the optima on chip thread number in different pipeline stage. Experiment shows the proposed architecture can reduce at least 27% external bandwidth and increase at least 25% throughput in a graphic streaming example. You-Ming Tsao, Ka-Hang Lok, Chih-Hao Sun, Shao-Yi Chien, Liang-Gee Chen |
ISCAS | 5 |
| 2008 | Enhanced temporal error concealment algorithm with edge-sensitive processing orderabstractError Concealment techniques are widely used in video decoder with error-prone communication channels. In this paper, an enhanced edge-sensitive processing order for temporal error concealment algorithm is proposed. Side information of neighboring macroblocks of the corrupted macroblocks are considered to derive a suitable processing order for error concealment, and a new motion vector searching algorithm is also proposed for temporal error concealment. Experimental results prove that the processing order plays an important role in error concealment. With considering the processing order, the proposed algorithm outperforms existing algorithms in terms of PSNR and perceptual artifacts, and the improvement of 2.45dB in PSNR can be achieved compared with the same system with raster-scan order. Tung-Hsing Wu, Guan-Lin Wu, Ching-Yi Chen, Shao-Yi Chien |
ISCAS | 4 |
| 2008 | Fast image segmentation based on K-Means clustering with histograms in HSV color spaceabstractA fast and efficient approach for color image segmentation is proposed. In this work, a new quantization technique for HSV color space is implemented to generate a color histogram and a gray histogram for K-Means clustering, which operates across different dimensions in HSV color space. Compared with the traditional K-Means clustering, the initialization of centroids and the number of cluster are automatically estimated in the proposed method. In addition, a filter for post-processing is introduced to effectively eliminate small spatial regions. Experiments show that the proposed segmentation algorithm achieves high computational speed, and salient regions of images can be effectively extracted. Moreover, the segmentation results are close to human perceptions. Tse-Wei Chen 0001, Yi-Ling Chen 0008, Shao-Yi Chien |
MMSP | 3 |
| 2008 | Tennis video enrichment with content layer separation and real-time rendering in sprite planeabstractSport video enrichment can provide viewers more interaction and user experiences. In this paper, with tennis sport video as an example, two techniques are proposed for video enrichment: content layer separation and real-time rendering. The video content is decomposed into different layers, like field, players and ball, and the enriched video is rendered by re-integrated these layers information. They are both executed in sprite plane to avoid complex 3D model construction and rendering. Experiments shows that it can generate nature and seamless edited video by viewerspsila requests, and the real-time processing speed of 30 720times480 frames per second can be achieved on a 3 GHz CPU. Jui-Hsin Lai, Shao-Yi Chien |
MMSP | 2 |
| 2008 | Baseball and tennis video annotation with temporal structure decompositionabstractSport video annotation can help viewers easily browse sport video content and quickly find the hot events and highlights in a game. Although many annotation algorithms have been proposed, they are not suitable for practical implementation since the high complexity and the low precision rates are not acceptable. In this paper, a method of sport video temporal structure decomposition, which decomposes the sport video into many video clips, is proposed. Then score box information and additional semantic information are important clues for event annotation. Experimental results show that the proposed algorithm can successfully and effectively decompose video into clips. The annotation results also have extremely high precision and recall rates for both baseball and tennis videos. Jui-Hsin Lai, Shao-Yi Chien |
MMSP | 2 |
| 2008 | CRISP: Coarse-Grained Reconfigurable Image Stream Processor for Digital Still Cameras and CamcordersabstractTo design the hardware for image signal processing pipelines in digital still cameras (DSCs) and video camcoders, it is a dilemma for conventional solutions, such as application-specific integrated circuits (ASICs) and digital signal processors (DSPs), to achieve high processing capability at low cost while maintaining high flexibility for various algorithms. With the observation of the characteristics of image signal-processing pipelines, including the different requirements for different operation modes and the algorithmic similarity of image-processing tasks, a new coarse-grained reconfigurable image stream processor (CRISP) is proposed in this paper. The design idea is to devote low-cost hardware for the requirements in the preview mode and add some hardware resources for higher flexibility and processing capability in the picture-taking mode. With the coarse-grained reconfigurable stage processing elements designed for image signal-processing tasks and the reconfigurable interconnection unit with unified communication protocol, CRISP can be reconfigured as an efficient dedicated hardware in the preview mode, and it can act like a flexible DSP for the picture-taking mode with different contexts. Implementation result shows that the core (die) size is 5 mm2(7.72 mm2) with TSMC 0.18-mum process, and the power consumption is 218 mW at 1.8 V. At the working frequency of 115 MHz, the processor is capable of processing 11 M-pixel still images at 10 fps for DSCs or 1920 times 1080 video frames at 55 fps for camcorders. CRISP can execute image pipelines 83 times faster than the state-of-the-art DSP with only about one-tenth die size. Jason C. Chen, Shao-Yi Chien |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Efficient Architecture Design of Motion-Compensated Temporal Filtering/Motion Compensated Prediction EngineabstractSince motion-compensated temporal filtering (MCTF) becomes an important temporal prediction scheme in video coding algorithms, this paper presents an efficient temporal prediction engine which not only is the first MCTF hardware work but also supports traditional motion-compensated prediction (MCP) scheme to provide computation scalability. For the prediction stage of MCTF and MCP schemes, modified extended double current Frames is adopted to reduce the system memory bandwidth, and a frame-interleaved macroblock pipelining scheme is proposed to eliminate the induced data buffer overhead. In addition, the proposed update stage architecture with pipelined scheduling and motion estimation (ME)-like motion compensation (MC) with level C+ scheme can also save about half external memory bandwidth and eliminate irregular memory access for MC. Moreover, 76.4% hardware area of the update stage is saved by reusing the hardware resources of the prediction stage. This MCTF chip can process CIF 30 fps in real-time, and the searching range is [-32, 32) for 5/3 MCTF with four-decomposition level and also support 1/3 MCTF, hierarchical B-frames, and MCP coding schemes in JSVM and H.264/AVC. The gate count is 352-K gates with 16.8 KBytes internal memory, and the maximum operating frequency is 60 MHz. Yi-Hau Chen, Chih-Chi Cheng, Tzu-Der Chuang, Ching-Yeh Chen, Shao-Yi Chien, Liang-Gee Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2008 | Content-Aware Prediction Algorithm With Inter-View Mode Decision for Multiview Video Codingabstract3-D video will become one of the most significant video technologies in the next-generation television. Due to the ultra high data bandwidth requirement for 3-D video, effective compression technology becomes an essential part in the infrastructure. Thus multiview video coding (MVC) plays a critical role. However, MVC systems require much more memory bandwidth and computational complexity relative to mono-view video coding systems. Therefore, an efficient prediction scheme is necessary for encoding. In this paper, a new fast prediction algorithm, content-aware prediction algorithm (CAPA) with inter-view mode decision, is proposed. By utilizing disparity estimation (DE) to find corresponding blocks between different views, the coding information, such as rate-distortion cost, coding modes, and motion vectors, can be effectively shared and reused from the coded view channel. Therefore, the computation for motion estimation (ME) in most view channels can be greatly reduced. Experimental results show that compared with the full search block matching algorithm (FSBMA) applied to both ME and DE, the proposed algorithm saves 98.4-99.1% computational complexity of ME in most view channels with negligible quality loss of only 0.03-0.06 dB in PSNR. Li-Fu Ding, Pei-Kuei Tsung, Shao-Yi Chien, Wei-Yin Chen, Liang-Gee Chen |
IEEE Trans. Multim. | 3 |
| 2007 | Depth Map Generation for 2D-to-3D Conversion by Short-Term Motion Assisted Color SegmentationabstractThis paper presented a novel depth map generation method - the short-term motion assisted color segmentation, which combines the pictorial, monocular and binocular depth cues of human vision. The proposed method utilizes a motion/edge registration technique to avoid the motion jitter error in common motion segmentation. And the motion/image segment adaptation algorithm matches the connected components with the motion segments. Even for static scene, the connected component algorithm is still working for the depth map generation. The experimental results show that the adaptation of motion and image segmentation improves quality and smoothness of the depth map both in the spatial and temporal domain. Yu-Lin Chang, Chih-Ying Fang, Li-Fu Ding, Shao-Yi Chien, Liang-Gee Chen |
ICME | 4 |
| 2007 | Robust Video Object Segmentation Based on K-Means Background Clustering and Watershed in Ill-Conditioned Surveillance SystemsabstractA robust video object segmentation algorithm for complex conditions in surveillance systems is proposed in this paper. This algorithm contains an unsupervised K-Means background clustering technique to model the temporal distribution in RGB domain for each spatial position. Based on the proposed background model, the object mask generation process integrates noise reduction, cast shadow cancellation, and improved watershed transform to obtain satisfying object masks. Experiments show that it can be applied on low-fame-rate and noisy video sequences in surveillance systems in which temporal tracking becomes impractical, and achieve better segmentation results than the previous works for complex lighting conditions and outdoor scenes. Tse-Wei Chen 0001, Shou-Chieh Hsu, Shao-Yi Chien |
ICME | 3 |
| 2007 | Motion Adaptive Spatio-Temporal Gaussian Noise Reduction Filter for Double-Shot ImagesabstractThe high performance of the conventional spatio-temporal image noise reduction filters comes at the cost of high computationally intensity. In this paper, a motion adaptive spatio-temporal filter is proposed for double-shot images to achieve high performance with low computing power requirement. Two images are used simultaneously to separate the static regions and dynamic regions, where different noise reduction approaches are employed. Temporal average and 2-D adaptive filter is applied for static regions, where spatio-temporal filter with motion compensation is only applied for dynamic regions. Experiments show that the proposed noise reduction filter can achieve better noise reduction performance with less computationally intensity than other previous spatio-temporal filters. Shao-Yi Chien, Tse-Wei Chen 0001 |
ICME | 1 |
| 2007 | 3D Video Applications and Intelligent Video Surveillance Camera and its VLSI DesignabstractIn this demonstration, the core processing engines of two video applications, 3D video and intelligent video surveillance, are demonstrated. The developed algorithms and its VLSI design results are shown with hardware prototypes processing input video on-the-fly. In addition to the processing engine design, the development tools for efficiently designing these chips are also demonstrated. Shao-Yi Chien, Chi-Sheng Shih 0001, Mong-Kai Ku, Chia-Lin Yang, Yao-Wen Chang, Tei-Wei Kuo, Liang-Gee Chen |
ICME | 1 |
| 2007 | Multi-Pass and Frame Parallel Algorithms of Motion Estimation in H.264/AVC for Generic GPUabstractIn this paper, multi-pass and frame parallel algorithms are proposed to accelerate various motion estimation (ME) tools in H.264 with the graphics processing unit (GPU). By the multi-pass method to unroll and rearrange the multiple nested loops, the integer-pel ME can be implemented with two-pass process on GPU. Moreover, fractional ME needs six passes for frame interpolation with six-tap filter and motion vector refinement. Motion estimation with multiple reference frames can be implemented with two-pass process with frame-level parallel scheme by use of SIMD vector operations of GPU. Experimental results show that, compared to implementations with only CPU, about 6 times to 56 times speed-up can be achieved for different ME algorithms. Chuan-Yiu Lee, Chi-Ling Wu, Chin-Hsiang Chang, You-Ming Tsao, Shao-Yi Chien |
ICME | 6 |
| 2007 | Virtual Conduction System with Multi-Resolution Wall DisplayabstractThe virtual conduction system (VCS) allows a user to conduct a photo-realistic pseudo orchestra. The VCS includes four modules: (1) gesture recognition module, (2) audio rendering module, (3) video rendering module, and (4) multi-resolution display module. With gesture recognition and tempo adjustment, the user not only can change the playback rate of an audio and video recording, but also can control the volume from different portion of an orchestra in real time. With the video rendering module and multi-resolution display module, the VCS provides a new visual experience with an interactive multi-resolution wall-size display. Wei-Ting Peng, En-Wei Huang, Wei-Lun Chang, Po-Chung Huang, Jun-Ying Bai, Han-Ru Chen, Shao-Yi Chien, Shyh-Kang Jeng, Yi-Ping Hung, Li-Chen Fu, Lin-Shan Lee |
ICME | 7 |
| 2007 | Cost Effective Color Filter Array Demosaicking with Chrominance Variance Weighted InterpolationabstractDemosaicking is a color interpolation process that converts a raw image generated by a color filter array to a full color image. For most of the proposed demosaicking methods, the design only focus on image quality without considering the VLSI hardware cost. In this paper, the hardware cost required for many demosaicking algorithms is first analyzed. According to this analysis, a cost effective method is proposed by use of chrominance variance weighting scheme. Experimental results show that the proposed method can achieve better image quality in PSNR than the existing methods on variety of test images while low hardware cost is still maintained. It shows that this method can be a good compromise between image quality and hardware cost. Tsung-Huang Chen, Shao-Yi Chien |
ISCAS | 2 |
| 2007 | Video Segmentation with Model-Based Sprite Generation for Panning Surveillance CamerasabstractVideo segmentation is one of the most important techniques for advanced video surveillance systems. Most of the existing video segmentation algorithms suffer from high computational complexity to achieve acceptable results for surveillance systems and cannot deal with moving camera situations. In this paper, we propose a low complexity video segmentation algorithm with model-based sprite generation for panning cameras. The characteristics of surveillance systems are utilized to reduce the complexity of sprite generation, which is employed to register background information for panning cameras. In addition, several novel techniques are also proposed to improve the quality of segmentation results. Experimental results show that the proposed algorithm can effectively segment moving objects with good subjective quality. Der-Chun Cherng, Shao-Yi Chien |
ISCAS | 2 |
| 2007 | Coding Mode Analysis of MPEG-2 to H.264/AVC Transcoding for Digital TV ApplicationsabstractMPEG-2 to H.264/AVC transcoding is an important module for video recoding in digital TV applications. For pixel domain transcoding, MPEG-2 bitstream is decoded and then re-encoded by H.264 encoder. Since the behavior of the decoded video is different from the original video, in this paper, the performance of each coding mode is analyzed to select the effective coding tools. It is shown from the analysis that transcoding with motion estimation with only one reference frame and 16times16 block size and deblocking filter can achieve almost the same video quality with only 19% of the computation. The analysis result could be an important reference for the implementation of MPEG-2 to H.264 transcoder. Yi-Nung Liu, Chi-Sun Tang, Shao-Yi Chien |
ISCAS | 3 |
| 2007 | Flexible and Cost Effective Transport Stream Processor for DTVabstractA flexible transport stream processor for DTV which is also designed under cost-effective consideration is proposed in this paper. A RISC micro-controller is allocated as the core of transport stream processor for flexibly extending or changing the functions of the transport stream processor. For the consideration of cost-effective design, the functions of the transport stream processor are partitioned into ones which are suitable for hardware implementation and the others suitable for the software executed by the micro-controller. Special enhancement of the instruction set of the micro-controller is proposed, with which the code efficiency of bit-level data-field processing could be improved. A general parsing engine for parsing grouping data-fields is also proposed. With the features described above, about 50% of total cost of transport stream processor with baseline functions can be saved. Chia-Liang Tsai, Shao-Yi Chien |
ISCAS | 2 |
| 2007 | System Bandwidth Analysis of Multiview Video Coding with Precedence ConstraintabstractMultiview video coding (MVC) systems require much more bandwidth and computational complexity relative to mono-view video systems. Thus, when designing a VLSI architecture for MVC systems, the hardware resource allocation is a critical issue. In this paper, we propose a new system bandwidth analysis scheme for various and complicated MVC structures. The precedence constraint in the graph theory is adopted for deriving the processing order of frames in a MVC system. In addition, current block centric scheduling (CBCS) and search window centric scheduling (SWCS) are proposed for MVC bandwidth analysis. By adopting data reuse schemes, several design points are explored with the aid of the proposed analysis scheme. The suitable hardware resource allocation can be easily determined. Pei-Kuei Tsung, Li-Fu Ding, Wei-Yin Chen, Shao-Yi Chien, Tung-Chien Chen, Liang-Gee Chen |
ISCAS | 4 |
| 2007 | Automatic Feature-Based Face Scoring in Surveillance SystemsabstractFacial images with low resolution in surveillance sequences are hard to detect with traditional approaches, and the quality of these faces is a significant factor for human face recognition. A new technique called face scoring, which determines the face scores based on face quality, is proposed. It combines spirits of image-based face detection and essences of video object segmentation to filter out face candidates. Besides, the face scoring technique includes eight scoring functions based on feature extraction technique, integrated by a single layer neural network training system to obtain an optimal linear combination to select high-quality faces. In the proposed algorithm, the way to choose input vector is quite different from traditional approaches and has good properties. Experiments show that the proposed algorithm effectively extracts low-resolution human faces, which traditional algorithm cannot handle well. It can also rank face candidates according to face scores, which is useful for surveillance video summary and indexing. Tse-Wei Chen 0001, Shou-Chieh Hsu, Shao-Yi Chien |
ISM | 3 |
| 2007 | Spatial-Temporal Error Detection Scheme for Video Transmission over Noisy ChannelsabstractError detection plays an important role in an error- robust video decoder. In this paper, a spatial-temporal error detection scheme for a video decoder is proposed. By considering inherently spatial and temporal similarities in video sequences, the visually corrected macroblocks in the decoded frames are detected by employing a set of error detection procedures, where one cross-boundary similarity index and one cross-frame similarity index are defined for spatial and temporal error detection, respectively. An adaptive threshold scheme is also proposed to make the proposed error detection method suitable for different video sequences. After being integrated with an H.264 decoder with error concealment techniques, the video quality improvement of 0.5-2.4 dB in PSNR is achieved. This method can also be integrated with other video codecs to improve the decoded video quality over noisy channels. Guan-Lin Wu, Shao-Yi Chien |
ISM | 2 |
| 2007 | Tennis video 2.0: a new framework of sport video applicationsabstractThis video demo presents a new framework of sport video applications called as Tennis Video 2.0. The proposed information extraction scheme retrieves the temporal structure of a video and separates the video foreground and background objects into different layers. With the structure and layer information, the new multimedia is generated. Contrary to the conventional video contents, the proposed new multimedia enables users to generate their own contents and feedback requests to the video players for more interaction. Users even can share their created contents with friends in different transmission bandwidth with considering the semantic. Jui-Hsin Lai, Shao-Yi Chien |
ACM Multimedia | 2 |
| 2007 | Real-Time Memory-Efficient Video Object Segmentation in Dynamic Background with Multi-Background Registration TechniqueabstractBackground subtraction video segmentation is the important first step for video surveillance applications with fixed camera. There are many existing methods in the literature. However, most of them are either too simple to handle complex environment, such as dynamic background, or too complex to be executed in real-time. Based on the proposed multi-background registration technique, this paper presents a real-time video object segmentation algorithm. The proposed algorithm can better handle non-static background cases compared with original single background registration segmentation. Compared with other works that can handle dynamic background cases, it is efficient in memory usage. Wei-Kai Chan, Shao-Yi Chien |
MMSP | 2 |
| 2007 | Efficient Face Detection with Segmentation and Feature-based Face Scoring in Surveillance SystemsabstractFacial images with low resolution in surveillance sequences are hard to detect with traditional approaches. An efficient face detection and face scoring technique in surveillance systems is proposed. It combines spirits of image-based face detection and essences of video object segmentation to filter out high-quality faces. The proposed face scoring technique, which is useful for surveillance video summary and indexing, includes four scoring functions based on feature extraction and is integrated by a neural network training system to select high-quality face. Experiments show that the proposed algorithm effectively extracts low-resolution human faces, which traditional face detection algorithm cannot handle well. It can also rank face candidates according to face scores, which determine face quality. Tse-Wei Chen 0001, Wei-Kai Chan, Shao-Yi Chien |
MMSP | 3 |
| 2007 | Fast Algorithm and Architecture Design of Low-Power Integer Motion Estimation for H.264/AVCabstractIn an H.264/AVC video encoder, integer motion estimation (IME) requires 74.29% computational complexity and 77.49% memory access and becomes the most critical component for low-power applications. According to our analysis, an optimal low-power IME engine should be a parallel hardware architecture supporting fast algorithms and efficient data reuse (DR). In this paper, a hardware-oriented fast algorithm is proposed with the intra-/inter-candidate DR considerations. In addition, based on the systolic array and 2-D adder tree architecture, a ladder-shaped search window data arrangement and an advanced searching flow are proposed to efficiently support inter-candidate DR and reduce latency cycles. According to the implementation results, 97% computational complexity is saved by the proposed fast algorithm. In addition, 77.6% memory bandwidth is further saved with the proposed DR techniques at architecture level. In the ultra-low-power mode, the power consumption is 2.13 mW for real-time encoding CIF 30-fps videos at 13.5-MHz operating frequency Tung-Chien Chen, Sung-Fang Tsai, Shao-Yi Chien, Liang-Gee Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2006 | Relative Depth Layer Extraction for Monoscopic Video by Use of Multidimensional FilterabstractThis paper presents a relative depth layer extraction system for monoscopic video, using multi-line filters and a layer selection algorithm. Main ideas are to extract multiple linear trajectory signals from videos and to determine their relative depths using the concept of motion parallax. The proposed superficial line model used for detecting slow moving objects provides sufficient taps within few frames to reduce frame buffer, while the closest-hit line model used for detecting fast motion objects provides few enough taps to prevent blurring. To increase the correctness of layer map, three-level layer map co-decision is used to compensate low texture region defect Jing-Ying Chang, Chao-Chung Cheng, Shao-Yi Chien, Liang-Gee Chen |
ICME | 3 |
| 2006 | Real-Time Depth Image based Rendering Hardware Accelerator for Advanced Three Dimensional Television Systemabstract3D TV will become a prominent technology in the next generation. In this paper, a depth image based rendering system is proposed from algorithm level to hardware architecture level. We propose a novel depth image based rendering algorithm with edge-dependent Gaussian filter and interpolation to improve the rendered stereo image quality. Based on our proposed algorithm, a fully-pipelined depth image based rendering hardware accelerator is proposed to support real-time rendering. The proposed hardware accelerator is optimized in three steps. First, we analyze the effect of fixed point operation and choose the optimal wordlength to keep the stereo image quality. Second, a three-parallel edge-dependent Gaussian filter architecture is proposed to solve the critical problem of memory bandwidth. Finally, we optimize the hardware cost by the proposed hardware architecture. Only 1/21 amounts of vertical PEs and 1/11 amounts of horizontal PEs is needed by the proposed folded edge-dependent Gaussian filter architecture. Furthermore, by the proposed check mode, the whole Z-buffer can be eliminated during 3D image warping. In additions, the on-chip SRAMs can be reduced to 66.7 percent compared with direct implementation by global and local disparity separation scheme. A prototype chip can achieve real-time requirement under the operating frequency of 80 MHz for 25 SDTV frames per second (fps) in left and right channel simultaneously. The simulation result also shows the hardware cost is quite small compared with the conventional rendering architecture Wan-Yu Chen, Yu-Lin Chang, Hsu-Kuang Chiu, Shao-Yi Chien, Liang-Gee Chen |
ICME | 4 |
| 2006 | Human Object Tracking Algorithm with Human Color Structure Descriptor for Video Surveillance SystemsabstractSegmentation, tracking, and description extraction are important operations in smart camera surveillance systems. In this paper, a robust segmentation-and-descriptor based tracking algorithm is proposed. Segmentation is applied first, and description for each connected component is extracted for object classification to generate the video object masks. It can do segmentation, tracking, and description extraction with a single algorithm without redundant computation. In addition, a new descriptor for human objects, human color structure descriptor (HCSD), is also proposed for this algorithm. Experimental results show that the proposed algorithm can provide precise video object masks and trajectories. It is also shown that the proposed descriptor, HCSD, can achieve better performance than scalable color descriptor and color structure descriptor of MPEG-7 for human objects Shao-Yi Chien, Wei-Kai Chan, Der-Chun Cherng, Jing-Ying Chang |
ICME | 1 |
| 2006 | CRISP: coarse-grain reconfigurable image signal processor for digital still camerasabstractThis paper presents a novel preview-based coarse-grain reconfigurable image signal processor (CRISP) for digital still cameras (DSCs). The two modes in DSCs, which have quite different hardware considerations, make traditional implementation methods inefficient. One is preview mode, which needs realtime constraints and the other one is picture-taking mode, which requires high flexibility and capability for various algorithms in it. Low cost design of CRISP considers simpler image pipelines in preview mode and extends flexibility required in picture-taking mode with proper hardware resources devotion. Algorithmic similarity in image pipelines and successful hardware classification lead it to a combination of low cost and high efficiency. Coarse-grain modules connected by reconfigurable interconnection make it a good compromise between dedicated hardware and DSPs, which are suitable for only one, not all of two modes in DSCs respectively. The experimental results show that the total gate count of it is 38.6K with 5.8K byte memory. It can save more than 75% area from high end DSP, such as Trimedia TM1300. Besides, CRISP reduces execution cycle number of image pipeline tasks, such as 2-D filters to only 0.17% of that required by TM1300. Jason C. Chen, Chun-Fu Shen, Shao-Yi Chien |
ISCAS | 3 |
| 2006 | Multi-pass algorithm of motion estimation in video encoding for generic GPUabstractThe importance of video encoding has boomed rapidly since video data communication was widely needed. In this paper, we propose a multi-pass algorithm to accelerate the motion estimation (ME), the dominant part in video encoding, with the graphics processing unit (GPU). By the multi-pass method to unroll and rearrange the multiple nested loops, the complex ME can be implemented on GPU. Besides, ME can be executed efficiently with the built-in parallel processing and texture filter of GPU. Experimental results show that, by utilizing the computing power of GPU, about two times and 14 times speed-up can be achieved for integer-pel ME and MPEG-1/2 half-pel ME, respectively Pei-Lun Li, Chin-Hsiang Chang, Chi-Ling Wu, You-Ming Tsao, Shao-Yi Chien |
ISCAS | 6 |
| 2006 | Algorithm and hardware architecture design for weighted prediction in H.264/MPEG-4 AVCabstractWeighted prediction (WP) is a tool to compensate the brightness difference in video sequences with brightness variations. In this paper, some weight parameter determination methods are surveyed, and a weighted prediction algorithm together with the hardware architecture design is proposed. The main idea of the algorithm is to limit the number of weight parameters transmitted by quantizing the parameter into levels and using only offset as the parameter. As a result, the extra parameters sent in each slice header is thus limited by the number of levels, and the parameter determination process requires much less computations. By further utilizing sub-sampling of brightness levels and estimated offset sum, a simplified architecture is also proposed. Simulation result shows that the later architecture achieves a coding gain of about 0.5 dB over weighted prediction method in JM9.6 and has a minimum overhead to hardware implementation. Chi-Sun Tang, Chen-Han Tsai, Shao-Yi Chien, Liang-Gee Chen |
ISCAS | 3 |
| 2006 | Adaptive tile depth filter for the depth buffer bandwidth minimization in the low power graphics systemsabstractDepth buffer bandwidth minimization with depth filter plays an important role in designing low power graphics processors. In this paper, an efficient adaptive tile depth filter (ATDF) algorithm is proposed. The key concept is to consider more occlusion conditions in order to achieve better performance. Two existing algorithms, Z/sub max/ and Z/sub max/ algorithms, are integrated to detect both occluded and non-occluded fragments. Moreover, two new techniques, coverage mask and adaptive tile mode, are proposed to further improve the performance. Experiments show that the proposed algorithm can filter out up to 85% fragments to reduce 40% memory bandwidth for test applications with complex scenes, which outperforms the previous works. You-Ming Tsao, Chi-Ling Wu, Shao-Yi Chien, Liang-Gee Chen |
ISCAS | 3 |
| 2006 | High Performance Low Cost Video Analysis Core for Smart Camera Chips in Distributed Surveillance NetworkabstractConventional smart camera cannot achieve realtime processing for high-performance video content analysis algorithms with only RISCs. In literatures, DSPs or coprocessors are employed to implement video content analysis functions. In this paper, a video content analysis core with specially-designed hardware accelerators is proposed to realize the content analysis functions in smart camera with low cost. The resulting smart camera in this paper can then provide high performance content analysis functions, including video segmentation, video object description, and video object tracking, for real-time surveillance applications. Moreover, design techniques such as frame-level pipelining and subword level parallelism are also applied on the design of these special hardware accelerators in order to achieve high throughput rate, high hardware utilization, and saving the bus bandwidth. When considering the total hardware cost of smart camera chip, the proposed video analysis core is very low cost and is suitable to be integrated in the next generation surveillance systems Wei-Kai Chan, Shao-Yi Chien |
MMSP | 2 |
| 2006 | Analysis and architecture design of an HDTV720p 30 frames/s H.264/AVC encoderabstractH.264/AVC significantly outperforms previous video coding standards with many new coding tools. However, the better performance comes at the price of the extraordinarily huge computational complexity and memory access requirement, which makes it difficult to design a hardwired encoder for real-time applications. In addition, due to the complex, sequential, and highly data-dependent characteristics of the essential algorithms in H.264/AVC, both the pipelining and the parallel processing techniques are constrained to be employed. The hardware utilization and throughput are also decreased because of the block/MB/frame-level reconstruction loops. In this paper, we describe our techniques to design the H.264/AVC video encoder for HDTV applications. On the system design level, in consideration of the characteristics of the key components and the reconstruction loops, the four-stage macroblock pipelined system architecture is first proposed with an efficient scheduling and memory hierarchy. On the module design level, the design considerations of the significant modules are addressed followed by the hardware architectures, including low-bandwidth integer motion estimation, parallel fractional motion estimation, reconfigurable intrapredictor generator, dual-buffer block-pipelined entropy coder, and deblocking filter. With these techniques, the prototype chip of the efficient H.264/AVC encoder is implemented with 922.8 K logic gates and 34.72-KB SRAM at 108-MHz operation frequency. Tung-Chien Chen, Shao-Yi Chien, Yu-Wen Huang, Chen-Han Tsai, Ching-Yeh Chen, To-Wei Chen, Liang-Gee Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Joint Prediction Algorithm and Architecture for Stereo Video Hybrid Coding Systemsabstract3-D video will be the most prominent video technology in the next generation. Among the 3-D video technologies, stereo video systems are considered to be realized first in the near future. Stereo video systems require double bandwidth and more than twice the computational complexity relative to mono-video systems. Thus, an efficient coding scheme is necessary for transmitting stereo video. In this paper, a new structure of prediction core in stereo video coding systems is proposed from the algorithm level to the hardware architecture level. The joint prediction algorithm (JPA), which combines three prediction schemes, is proposed for high coding efficiency and low computational complexity. It makes the system outperform MPEG-4 temporal scalability and simple profile by 2-3 dB in rate-distortion performance. Besides, JPA also utilizes the characteristics of stereo video and successfully reduces about 80% computational complexity. Then, a new hardware architecture of the prediction core based on JPA and a modified hierarchical search block-matching algorithm is proposed. With a special data flow, no bubble cycles exist during the block-matching process. The proposed architecture also adopts the near-overlapped candidates reuse scheme to save the heavy burden of data access. Besides, both on-chip memory requirement and off-chip memory bandwidth can be reduced by the proposed new scheduling. Compared with the hardware requirement for the implementation of full search block-matching algorithm, only 11.5% on-chip SRAM and 3.3% processing elements are needed with a tiny PSNR drop, making it area-efficient while maintaining high stereo video quality and processing capability Li-Fu Ding, Shao-Yi Chien, Liang-Gee Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Analysis and complexity reduction of multiple reference frames motion estimation in H.264/AVCabstractIn the new video coding standard H.264/AVC, motion estimation (ME) is allowed to search multiple reference frames. Therefore, the required computation is highly increased, and it is in proportion to the number of searched reference frames. However, the reduction in prediction residues is mostly dependent on the nature of sequences, not on the number of searched frames. Sometimes the prediction residues can be greatly reduced, but frequently a lot of computation is wasted without achieving any better coding performance. In this paper, we propose a context-based adaptive method to speed up the multiple reference frames ME. Statistical analysis is first applied to the available information for each macroblock (MB) after intra-prediction and inter-prediction from the previous frame. Context-based adaptive criteria are then derived to determine whether it is necessary to search more reference frames. The reference frame selection criteria are related to selected MB modes, inter-prediction residues, intra-prediction residues, motion vectors of subpartitioned blocks, and quantization parameters. Many available standard video sequences are tested as examples. The simulation results show that the proposed algorithm can maintain competitively the same video quality as exhaustive search of multiple reference frames. Meanwhile, 76 %-96 % of computation for searching unnecessary reference frames can be avoided. Moreover, our fast reference frame selection is orthogonal to conventional fast block matching algorithms, and they can be easily combined to achieve further efficient implementations. Yu-Wen Huang, Bing-Yu Hsieh, Shao-Yi Chien, Shyh-Yih Ma, Liang-Gee Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2005 | Reconfigurable Platform for Content Science ResearchabstractThe College of Electrical Engineering and Computer Science at the National Taiwan University has identified the area of content science for media-rich life, broadly construed, as one of core areas for the college's future directions. One major aspect of this project is to develop the enabling technology for reconfigurable platforms for multimedia applications. Specifically, the goal is to develop reconfigurable platforms of system-on-a-chip (SoC) components for the applications to support the needs of multi-modal multimedia contents and to provide rapid system prototyping. The faculties in the College of Electrical Engineering and Computer Science have formed a multi-discipline team to develop such technology. Our team includes seven faculties and more than thirty students from the college. This short report describes the reconfigurable platform for content science research activities currently underway by our team. Our current activities include to develop the technology to analyze the critical path for avoiding hardware contention, to minimize the use of logic components, to design the multimedia IPs, to optimally route the bus and place the logic units, to design energy efficient cache, to evaluate the performance and power consumption, to design the algorithm for temporal floor-planning/placement. Chi-Sheng Shih 0001, Chia-Lin Yang, Mong-Kai Ku, Tei-Wei Kuo, Shao-Yi Chien, Yao-Wen Chang, Liang-Gee Chen |
RTCSA | 5 |
| 2005 | Partial-result-reuse architecture and its design technique for morphological operations with flat structuring elementsabstractMathematical morphology operations are applied in many real-time applications, such as video segmentation. For real-time requirement, efficient hardware implementation is necessary. This paper proposes a new architecture named Partial-Result-Reuse (PRR) architecture for mathematical morphological operations with flat structuring elements. Partial results generated during calculation process are kept and reused in this architecture to reduce hardware cost. With PRR concept and self-affinity property of structuring elements, the proposed architecture is more cost-effective and more general than existing morphology architectures. Moreover, it can be combined with systolic array to give consideration to both flexibility and hardware cost. We also propose a methodology to generate PRR architecture. With graphic method, the PRR architecture can be easily generated, and it can deal with structuring elements of any shape. The very large scale integration implementation of the PRR architecture shows that the area of processing element is small and can be fully piplelined without large overhead. The maximum frequency of the chip is 200 MHz in simulation, while processing speed of 550 morphological operations/s on a 720/spl times/480 frame can be achieved. Shao-Yi Chien, Shyh-Yih Ma, Liang-Gee Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2004 | Global elimination algorithm and architecture design for fast block matching motion estimationabstractThis paper presents a new block matching motion estimation algorithm and its VLSI architecture design. The proposed global elimination algorithm (GEA) was derived from successive elimination algorithm (SEA), which can skip unnecessary sum of absolute difference (SAD) calculation by comparing minimum SAD with subsampled SAD (SSAD). Our basic idea is to separate the decision of early termination and SAD calculation for each candidate block to make data flow more regular and suitable for hardware. In short, we first compare the rough characteristics of all candidate blocks with the current block (SSAD). In turn, we select several best roughly matched candidate blocks to re-compare them with the current block by using detailed characteristics (SAD). Other features of GEA include fixed processing cycles, no initial guess, and high video quality (almost the same as full search). Unlike other fast algorithms, the mapping of GEA to hardware is very simple. We proposed an architecture that is composed of a systolic part to efficiently compute SSAD, an adder tree to support both SSAD and SAD calculations, and a comparator tree to avoid expensive sorting circuits. Simulation results show that our design is much more area efficient than many full-search architectures while maintaining high video quality and processing capability. Yu-Wen Huang, Shao-Yi Chien, Bing-Yu Hsieh, Liang-Gee Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Fast video segmentation algorithm with shadow cancellation, global motion compensation, and adaptive threshold techniquesabstractAutomatic video segmentation plays an important role in real-time MPEG-4 encoding systems. Several video segmentation algorithms have been proposed; however, most of them are not suitable for real-time applications because of high computation load and many parameters needed to be set in advance. This paper presents a fast video segmentation algorithm for MPEG-4 camera systems. With change detection and background registration techniques, this algorithm can give satisfying segmentation results with low computation load. The processing speed of 40 QCIF frames per second can be achieved on a personal computer with an 800 MHz Pentium-III processor. Besides, it has shadow cancellation mode, which can deal with light changing effect and shadow effect. A fast global motion compensation algorithm is also included in this algorithm to make it applicable in slight moving camera situations. Furthermore, the required parameters can be decided automatically, which can enhance the proposed algorithm to have adaptive threshold ability. It can be integrated into MPEG-4 videophone systems and digital cameras. Shao-Yi Chien, Yu-Wen Huang, Bing-Yu Hsieh, Shyh-Yih Ma, Liang-Gee Chen |
IEEE Trans. Multim. | 1 |
| 2003 | Analysis and reduction of reference frames for motion estimation in MPEG-4 AVC/JVT/H.264abstractIn the new video coding standard, MPEG-4 AVC/JVT/H.264, motion estimation is allowed to use multiple reference frames. The reference software adopts a full search scheme, and the increased computation is in proportion to the number of searched reference frames. However, the reduction of prediction residues is highly dependent on the nature of the sequences, not on the number of searched frames. We present a method to speed up the matching process for multiple reference frames. For each macroblock, we analyze the available information after intra prediction and motion estimation from the previous frame to determine whether it is necessary to search more frames. The information we use includes selected mode, inter prediction residues, intra prediction residues, and motion vectors. Simulation results show that the proposed algorithm can save up to 90% of unnecessary frames while keeping the average miss rate of optimal frames less than 4%. Yu-Wen Huang, Bing-Yu Hsieh, Tu-Chih Wang, Shao-Yi Chien, Shyh-Yih Ma, Chun-Fu Shen, Liang-Gee Chen |
ICASSP (3) | 4 |
| 2003 | Efficient stereo video coding system for immersive teleconference with two-stage hybrid disparity estimation algorithmabstractAn efficient coding system is required for immersive teleconference to transmit stereo video. We propose a novel stereo video coding system by exploiting mesh-based disparity estimation and compensation scheme to achieve high coding efficiency and view synthesis ability. Based on the base-layer-enhance-layer structure, this system can provide stereoscopic scalability and compatibility to standards. With considering asymmetric spatial resolution property, good subjective quality can be achieved in ultra low bitrate situation. A novel fast disparity estimation algorithm named as two-stage iterative block and octagonal matching (TS-IBOM) algorithm is also proposed for this system. Experiments show that the proposed disparity estimation algorithm can generate accurate disparity vectors quickly. It is also shown that the proposed coding system has better coding efficiency than MPEG-4 simple profile and temporal scalability tool. The ultra low bitrate of 69 Kbps can be reached to encode 384/spl times/192 60 fps stereo video. Shao-Yi Chien, Shu-Han Yu, Li-Fu Ding, Yun-Nien Huang, Liang-Gee Chen |
ICIP (1) | 1 |
| 2003 | Unsupervised object-based sprite coding system for tennis sportabstractSprite coding is a new objected-based coding technology proposed by MPEG-4 video standard. In this paper, we propose an unsupervised sprite coding system for sport videos, for example, tennis sequences. Our system can provide several important functions. First, the sprite of the background can be generated without any pre-processing masks in our system. Second, it can automatically segment the foreground and background in a video sequence. Third, it can provide the masks of the foreground objects and the tennis ball. The experimental results show that our system has a very good performance and the coding gain of our system compared with MPEG-4 advanced simple profile is 2.5 dB at the low bit rate (270 Kbps) and is 2 dB at the ultra low bit rate (70 Kbps). It can be used in the sport video coding at low bit rate and provides a object-based sport video sequence. Ching-Yeh Chen, Shao-Yi Chien, Yi-Hau Chen, Yu-Wen Huang, Liang-Gee Chen |
ICME | 2 |
| 2003 | Analysis and reduction of reference frames for motion estimation in MPEG-4 AVC/JVT/H.264abstractIn the new video coding standard, MPEG-4 AVC/JVT/H.264, motion estimation is allowed to use multiple reference frames. The reference software adopts full search scheme, and the increased computation is in proportion to the number of searched reference frames. However, the reduction of prediction residues is highly dependent on the nature of sequences, not on the number of searched frames. In this paper, we present a method to speed up the matching process for multiple reference frames. For each macroblock, we analyze the available information after intra prediction and motion estimation from previous one frame to determine whether it is necessary to search more frames. The information we use includes selected mode, inter prediction residues, intra prediction residues, and motion vectors. Simulation results show that the proposed algorithm can save up to 90% of unnecessary frames while keeping the average miss rate of optimal frames less than 4%. Yu-Wen Huang, Bing-Yu Hsieh, Tu-Chih Wang, Shao-Yi Chien, Shyh-Yih Ma, Chun-Fu Shen, Liang-Gee Chen |
ICME | 4 |
| 2003 | Fast disparity estimation algorithm for mesh-based stereo image/video compression with two-stage hybrid approach
Shao-Yi Chien, Shu-Han Yu, Li-Fu Ding, Yun-Nien Huang, Liang-Gee Chen |
VCIP | 1 |
| 2003 | Fast motion estimation algorithm for H.264/MPEG-4 AVC by using multiple reference frame skipping criteria
Bing-Yu Hsieh, Yu-Wen Huang, Tu-Chih Wang, Shao-Yi Chien, Liang-Gee Chen |
VCIP | 4 |
| 2003 | Predictive watershed: a fast watershed algorithm for video segmentationabstractThe watershed transform is a key operator in video segmentation algorithms. However, the computation load of watershed transform is too large for real-time applications. In this paper, a new fast watershed algorithm, named P-watershed, for image sequence segmentation is proposed. By utilizing the temporal coherence property of the video signal, this algorithm updates watersheds instead of searching watersheds in every frame, which can avoid a lot of redundant computation. The watershed process can be accelerated, and the segmentation results are almost the same as those of conventional algorithms. Moreover, an intra-inter watershed scheme (IP-watershed) is also proposed to further improve the results. Experimental results show that this algorithm can save 20%-50% computation without degrading the segmentation results. This algorithm can be combined with any video segmentation algorithm to give more precise segmentation results. An example is also shown by combining a background registration and change-detection-based segmentation algorithm with P-Watershed. This new video segmentation algorithm can give accurate object masks with acceptable computation complexity. Shao-Yi Chien, Yu-Wen Huang, Liang-Gee Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Predictive watershed for image sequences segmentationabstractWatershed transform is a key operator in video segmentation algorithms. In this paper, a new fast watershed algorithm named P-Watershed for image sequences segmentation is proposed. By utilizing the temporal coherence property of video signal, this algorithm updates watersheds instead of searching watersheds in whole image. The watershed process can be accelerated and the segmentation results are almost the same as those of conventional algorithms. Moreover, an intra-inter watershed scheme (IP-Watershed) is also proposed to further improve the results. Experimental results show this algorithm can save 50% computation without degrading the segmentation results. Shao-Yi Chien, Yu-Wen Huang, Shyh-Yih Ma, Liang-Gee Chen |
ICASSP | 1 |
| 2002 | An efficient and low power architecture design for motion estimation using global elimination algorithmabstractThis paper presents a new algorithm and architecture for motion estimation. The proposed global elimination algorithm (GEA) is derived from successive elimination algorithm (SEA). The main idea is to remove the branches of SEA to make data flow more regular and suitable for hardware. Besides, the processing time per motion vector for GEA is fixed, no initial guess is required, and the skipping ratio of search positions can be fixed within frames and is even higher than 99%. The average PSNR of compensated frames is almost the same (within 0.1 dB) as that of full-search block matching algorithm (FBMA). An architecture composed of a systolic part, an adder tree, and a comparator tree is also developed for GEA. Simulation results show our design outperforms many FBMA architectures in normalized processing capability per gate and normalized power at gate level. Yu-Wen Huang, Shao-Yi Chien, Bing-Yu Hsieh, Liang-Gee Chen |
ICASSP | 2 |
| 2002 | A fast and high subjective quality sprite generation algorithm with frame skipping and multiple sprites techniquesabstractSprite coding, which is a new coding tool in MPEG-4, can achieve high coding efficiency with high subjective quality at low bit rate. Many sprite generation algorithms have been proposed; however, the computational intensity is very high and the quality is not good enough because of the limitation of simple motion models. A novel sprite generation algorithm is proposed with several new techniques. A frame skipping technique can generate the sprite using only several important frames to accelerate the process and achieve similar subjective quality. In addition, boundary matching and multiple sprites techniques can overcome the limitation of simple motion models to achieve high subjective quality with little computation overhead. Experiments show the proposed algorithm is 46 times faster than the algorithms in MPEG-4 VM and have high subjective quality. These techniques can be also applied with other sprite generation algorithms. Shao-Yi Chien, Ching-Yeh Chen, Wei-Min Chao, Yu-Wen Huang, Liang-Gee Chen |
ICIP (1) | 1 |
| 2002 | Multiple sprites and frame skipping techniques for sprite generation with high subjective quality and fast speedabstractSprite is an image collecting information of a video object through a video sequence. It can be used for efficient video coding, video summary, browsing, and editing. In this paper, three new techniques for sprite generation are proposed. Boundary matching and multiple sprites techniques can improve the subjective quality with refining the positions of the warped frames and generating more than one sprites. The frame skipping technique can skip redundant frames that contain only little new information when the camera revisits a scene several times to accelerate the sprite generation process. Experimental results show that these techniques can be employed independently and can improve the subjective quality as well as reduce the 47.68% - 17.22% runtime of sprite generation. They can be applied with any sprite generation algorithms. Shao-Yi Chien, Ching-Yeh Chen, Yu-Wen Huang, Liang-Gee Chen |
ICME (1) | 1 |
| 2002 | Simple and effective algorithm for automatic tracking of a single object using a pan-tilt-zoom cameraabstractThis paper presents a simple but effective algorithm for a pan-tilt-zoom camera to automatically track a single moving object. The proposed tracking algorithm is suitable for but not limited to stationary background, and it can tolerate reasonable noise and light change. The main idea is to initially capture the background information and to calculate the difference between incoming frame and background buffer. Block-based processing is adopted to reduce computation and to alleviate noise effects. Skin-color detection is combined with spatial and temporal information to let the camera more likely to focus on human faces. Even more, the tracking algorithm can be integrated with region-of-interest video coding, which allocates more bits for human faces to produce better subject views. Many practical situations have been tested and the simulation results show that the proposed tracking algorithm is useful for surveillance systems. Yu-Wen Huang, Bing-Yu Hsieh, Shao-Yi Chien, Liang-Gee Chen |
ICME (1) | 3 |
| 2002 | Automatic threshold decision of background registration technique for video segmentation
Yu-Wen Huang, Shao-Yi Chien, Bing-Yu Hsieh, Liang-Gee Chen |
VCIP | 2 |
| 2002 | Efficient moving object segmentation algorithm using background registration techniqueabstractAn efficient moving object segmentation algorithm suitable for real-time content-based multimedia communication systems is proposed in this paper. First, a background registration technique is used to construct a reliable background image from the accumulated frame difference information. The moving object region is then separated from the background region by comparing the current frame with the constructed background image. Finally, a post-processing step is applied on the obtained object mask to remove noise regions and to smooth the object boundary. In situations where object shadows appear in the background region, a pre-processing gradient filter is applied on the input image to reduce the shadow effect. In order to meet the real-time requirement, no computationally intensive operation is included in this method. Moreover, the implementation is optimized using parallel processing and a processing speed of 25 QCIF fps can be achieved on a personal computer with a 450-MHz Pentium III processor. Good segmentation performance is demonstrated by the simulation results. Shao-Yi Chien, Shyh-Yih Ma, Liang-Gee Chen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | Partial-result-reuse architecture and its design technique for morphological operationsabstractThis paper proposes a new cost-effective architecture for mathematical morphology named partial-result-reuse (PRR) architecture. For many real-time applications of mathematical morphology, the hardware implementation is necessary; however, the hardware cost of most existing morphology architectures is too high when dealing with large structuring elements. With the partial-result-reuse concept and self-affinity property of general structuring elements, the proposed architecture is more cost-effective and more general than the existing morphology architectures. It can deal with morphological operations with arbitrary structuring elements and can be used for other semi-group operations, and only 2[log/sub 2/n] comparators are needed for n/spl times/n structuring elements. Simulation shows that this architecture can dramatically reduce the hardware cost of morphological operations with all kinds of structuring elements. Shao-Yi Chien, Shyh-Yih Ma, Liang-Gee Chen |
ICASSP | 1 |
| 2001 | Automatic Video Segmentation For MPEG-4 Using PredictivewatershedabstractA fast watershed based video segmentation algorithm is proposed in this paper. Watershed based video segmentation algorithms are the mainstream because of high subjective quality; however, they are too computational intensive for real-time applications. Combining prior fast segmentation algorithm, which is based on change detection and background registration, and predictive watershed, a novel fast watershed algorithm for image sequences segmentation, this algorithm can give precise segmentation results and achieve real-time requirement. Simulation shows processing speed of 9.4 QCIF fps can be achieved on a PC with a Pentium-III 800 MHz processor. Shao-Yi Chien, Yu-Wen Huang, Shyh-Yih Ma, Liang-Gee Chen |
ICME | 1 |
| 2000 | Efficient video segmentation algorithm for real-time MPEG-4 camera system
Shao-Yi Chien, Shyh-Yih Ma, Liang-Gee Chen |
VCIP | 1 |