Hsueh-Ming Hang

dblp:91/3110 · DBLP profile ↗
← Back
79ranked-venue papers
7as first author
8since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 64 · 4 first-author · 5 since 2021Systems, architecture and hardware · 10 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Computer networks · 2 · 2 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2023 Hierarchical B-Frame Video Coding Using Two-Layer CANF Without Motion Coding
abstract
Typical video compression systems consist of two main modules: motion coding and residual coding. This general architecture is adopted by classical coding schemes (such as international standards H.265 and H.266) and deep learning-based coding schemes. We propose a novel B-frame coding architecture based on two-layer Conditional Augmented Normalization Flows (CANF). It has the striking feature of not transmitting any motion information. Our proposed idea of video compression without motion coding offers a new direction for learned video coding. Our base layer is a low-resolution image compressor that replaces the full-resolution motion compressor. The low-resolution coded image is merged with the warped high-resolution images to generate a high-quality image as a conditioning signal for the enhancement-layer image coding in full resolution. One advantage of this architecture is significantly reduced computational complexity due to eliminating the motion information compressor. In addition, we adopt a skip-mode coding technique to reduce the transmitted latent samples. The rate-distortion performance of our scheme is slightly lower than that of the state-of-the-art learned B-frame coding scheme, B-CANF, but outperforms other learned B-frame coding schemes. However, compared to B-CANF, our scheme saves 45% of multiply-accumulate operations (MACs) for encoding and 27% of MACs for decoding. The code is available at https://nycu-clab.github.io.
David Alexandre, Hsueh-Ming Hang, Wen-Hsiao Peng
CVPR2
2023 Fast Vehicle Detection and Tracking on Fisheye Traffic Monitoring Video using Motion Trail
abstract
We develop a vehicle detection and tracking scheme based on the concept of motion trails for fisheye traffic monitoring videos. The motion trail combines the moving object traces in several frames into one image. Because it collects information from multiple frames, the accuracy of detecting a trail is higher than a single-frame object detector. Essentially, it merges the detection and tracking processes into one process. In addition, a lightweight neural net is sufficient to detect the trail, which saves computing time and memory. After detecting the trails, we extract individual car locations at each frame using a multi-head trail extractor. Then, a multi-modal bidirectional LSTM can further improve detection accuracy. We adopt the public ICIP2020 VIP Cup dataset for training and testing. Our approach is 14 percentage points (pp) better than the state-of-the-art single-frame rotated object detector (R3Det) on the challenging nighttime video, and it is 5 FPS faster in inference speed. Our scheme achieves the AP50accuracy comparable with the state-of-the-art video object detector (MEGA), but its speed is 3 times faster, and its model size is only 28% of that of MEGA.
Sandy Ardianto, Hsueh-Ming Hang, Wen-Huang Cheng
ISCAS2
2023 Learned Hierarchical B-frame Coding with Adaptive Feature Modulation for YUV 4: 2: 0 Content
abstract
This paper introduces a learned hierarchical B-frame coding scheme in response to the Grand Challenge on Neural Network-based Video Coding at ISCAS 2023. We address specifically three issues, including (1) B-frame coding, (2) YUV 4:2:0 coding, and (3) content-adaptive variable-rate coding with only one single model. Most learned video codecs operate internally in the RGB domain for P-frame coding. B-frame coding for YUV 4:2:0 content is largely under-explored. In addition, while there have been prior works on variable-rate coding with conditional convolution, most of them fail to consider the content information. We build our scheme on conditional augmented normalized flows (CANF). It features conditional motion and inter-frame codecs for efficient B-frame coding. To cope with YUV 4:2:0 content, two conditional inter-frame codecs are used to process the Y and UV components separately, with the coding of the UV components conditioned additionally on the Y component. Moreover, we introduce adaptive feature modulation in every convolutional layer, taking into account both the content information and the coding levels of B-frames to achieve content-adaptive variable-rate coding. Experimental results show that our model outperforms x265 and the winner of last year's challenge on commonly used datasets in terms of PSNR-YUV.
Mu-Jung Chen, Hong-Sheng Xie, Cheng Chien, Wen-Hsiao Peng, Hsueh-Ming Hang
ISCAS5
2022 Fast Vehicle Detection and Tracking on Fisheye Traffic Monitoring Video Using CNN and Bounding Box Propagation
abstract
We design a fast car detection and tracking algorithm for traffic monitoring fisheye video mounted on crossroads. We use ICIP 2020 VIP Cup dataset and adopt YOLOv5 as the object detection base model. The nighttime video of this dataset is very challenging, and the detection accuracy (AP50) of the base model is about 54%. We design a reliable car detection and tracking algorithm based on the concept of bounding box propagation among frames, which provides 17.9 percentage points (pp) and 7 pp accuracy improvement over the base model for the nighttime and daytime videos, respectively. To speed up, the grayscale frame difference is used for the intermediate frames in a segment, which can double the processing speed.
Sandy Ardianto, Hsueh-Ming Hang, Wen-Huang Cheng
ICIP2
2022 Learned Video Compression for YUV 4: 2: 0 Content Using Flow-based Conditional Inter-frame Coding
abstract
This paper proposes a learning-based video compression framework for variable-rate coding on YUV 4:2:0 content. Most existing learning-based video compression models adopt the traditional hybrid-based coding architecture, which involves temporal prediction followed by residual coding. However, recent studies have shown that residual coding is suboptimal from the information-theoretic perspective. In addition, most existing models are optimized with respect to RGB content. Furthermore, they require separate models for variable-rate coding. To address these issues, this work presents an attempt to incorporate the conditional inter-frame coding for YUV 4:2:0 content. We introduce a conditional flow-based inter-frame coder to improve the inter-frame coding efficiency. To adapt our codec to YUV 4:2:0 content, we adopt a simple strategy of using space-to-depth and depth-to-space conversions. Lastly, we employ a rate-adaption net to achieve variable-rate coding without training multiple models. Experimental results show that our model performs better than x265 on UVG and MCL-JCV datasets in terms of PSNR-YUV. However, on the more challenging datasets from ISCAS’22 GC, there is still ample room for improvement. This insufficient performance is due to the lack of inter-frame coding capability at a large GOP size and can be mitigated by increasing the model capacity and applying an error propagationaware training strategy.
Yung-Han Ho, Chih-Hsuan Lin, Mu-Jung Chen, Chih-Peng Chang, Wen-Hsiao Peng, Hsueh-Ming Hang
ISCAS7
2022 Two-Layer Learning-Based P-Frame Coding with Super-Resolution and Content-Adaptive Conditional ANF
abstract
Deep-learning-based video compression technique has been rapidly growing in recent years. This paper adopts the Conditional Augmented Normalizing Flow video codec (CANF-VC) [8] as our basic system. To improve the quality of the condition signal (image) for CANF, we propose a two-layer structure learning-based video codec. At low cost of extra bit rate, the low-resolution base layer provides side information to improve the quality of motion-compensated reference frame through a super-resolution module with a merge-net. In addition, the base layer also provides information to the skip-mask generator. The skip-mask guides the coding mechanism to reduce the transmitted samples for the high-resolution enhancement layer. The experiment results indicate that the proposed two-layer coding scheme can provide 22.19% PSNR BD-Rate saving and 49.59% MS-SSIM BD-Rate saving over H.265 (HM 16.20) on the UVG test sequences.
David Alexandre, Hsueh-Ming Hang, Wen-Hsiao Peng
MMAsia2
2021 Deep Video Compression for Interframe Coding
abstract
A typical learning-based video compression scheme consists of motion coding and residual coding. In this paper, our deep video compression features a motion predictor and refinement networks for interframe coding. To save the bits for transmitting motion information, our scheme performs local motion prediction and sends only the differential motion vectors to the decoder. In the residual coding, we couple the residual decoder with the refine-net to reduce residual signal bits. The experiments show that our work can produce a very competitive coding performance compared to the other learning-based predictive video codecs.
David Alexandre, Hsueh-Ming Hang, Wen-Hsiao Peng, Marek Domanski
ICIP2
2021 Fined: Fast Inference Network for Edge Detection
abstract
In this paper, we address the design of lightweight deep learning-based edge detection. The deep learning technology offers a significant improvement on the edge detection accuracy. However, typical neural network designs have very high model complexity, which prevents it from practical usage. In contrast, we propose a Fast Inference Network for Edge Detection (FINED), which is a lightweight neural net dedicated to edge detection. By carefully choosing proper components for edge detection purpose, we can achieve the state-of-the-art accuracy in edge detection while significantly reducing its complexity. Another key contribution in increasing the inferencing speed is introducing the training helper concept. The extra subnetworks (training helper) are employed in training but not used in inferencing. It can further reduce the model complexity and yet maintain the same level of accuracy. Our experiments show that our systems outperform all the current edge detectors at about the same model (parameter) size.
Jan Kristanto Wibisono, Hsueh-Ming Hang
ICME2
2020 Traditional Method Inspired Deep Neural Network For Edge Detection
abstract
Recently, Deep-Neural-Network (DNN) based edge prediction is progressing fast. Although the DNN based schemes outperform the traditional edge detectors, they have much higher computational complexity. It could be that the DNN based edge detectors often adopt the neural net structures designed for high-level computer vision tasks, such as image segmentation and object recognition. Edge detection is a rather local and simple job, the over-complicated architecture and massive parameters may be unnecessary. Therefore, we propose a traditional method inspired framework to produce good edges with minimal complexity. We simplify the network architecture to include Feature Extractor, Enrichment, and Summarizer, which roughly correspond to gradient, low pass filter, and pixel connection in the traditional edge detection schemes. The proposed structure can effectively reduce the complexity and retain the edge prediction quality. Our TIN2 (Traditional Inspired Network) model has an accuracy higher than the recent BDCN2 (Bi-Directional Cascade Network) but with a smaller model.
Jan Kristanto Wibisono, Hsueh-Ming Hang
ICIP2
2020 A Hybrid Layered Image Compressor with Deep-Learning Technique
abstract
This paper presents a detailed description of NCTU's proposal for learning-based image compression, in response to the JPEG AI Call for Evidence Challenge. The proposed compression system features a VVC intra codec as the base layer and a learning-based residual codec as the enhancement layer. The latter aims to refine the quality of the base layer via sending a latent residual signal. In particular, a base-layer-guided attention module is employed to focus the residual extraction on critical high-frequency areas. To reconstruct the image, this latent residual signal is combined with the base-layer output in a non-linear fashion by a neural-network-based synthesizer. The proposed method shows comparable rate-distortion performance to single-layer VVC intra in terms of common objective metrics, but presents better subjective quality particularly at high compression ratios in some cases. It consistently outperforms HEVC intra, JPEG 2000, and JPEG. The proposed system incurs 18M network parameters in 16-bit floating-point format. On average, the encoding of an image on Intel Xeon Gold 6154 takes about 13.5 minutes, with the VVC base layer dominating the encoding runtime. On the contrary, the decoding is dominated by the residual decoder and the synthesizer, requiring 31 seconds per image.
Wei-Cheng Lee, Chih-Peng Chang, Wen-Hsiao Peng, Hsueh-Ming Hang
MMSP4
2020 Recent Advances in End-to-End Learned Image and Video Compression
abstract
The DCT-based transform coding technique was adopted by the international standards (ISO JPEG, ITU H.261/264/265, ISO MPEG-2/4/H, and many others) for nearly 30 years. Although researchers are still trying to improve its efficiency by fine-tuning its components and parameters, the basic structure has not changed in the past two decades.The deep learning technology recently developed may provide a new direction for constructing a high-compression image/video coding system. Recent results, particularly from the Challenge on Learned Image Compression (CLIC) at CVPR, indicate that this new type of schemes (often trained end-to-end) may have good potential for further improving compression efficiency.In the first part of this tutorial, we shall (1) summarize briefly the progress of this topic in the past 3 or so years, including an overview of CLIC results and JPEG AI Call-for-Evidence Challenge on Learning-based Image Coding (issued in early 2020). Because Deep Neural Network (DNN)-based image compression is a new area, several techniques and structures have been tested. The recently published autoencoder-based schemes can achieve similar PSNR to BPG (Better Portable Graphics, H.265 still image standard) and has superior subject quality (e.g., MSSSIM), especially at the very low bit rates. In the second part, we shall (2) address the detailed design concepts of image compression algorithms using the autoencoder structure. In the third part, we shall switch gears to (3) explore the emerging area of DNN-based video compression. Recent publications in this area have indicated that end-to-end trained video compression can achieve comparable or superior rate-distortion performance to HEVC/H.265. The CLIC at CVPR 2020 also created for the first time a new track dedicated to P-frame coding.
Wen-Hsiao Peng, Hsueh-Ming Hang
VCIP2
2019 Learned Image Compression with Soft Bit-Based Rate-Distortion Optimization
abstract
This paper introduces the notion of soft bits to address the rate-distortion optimization for learning-based image compression. Recent methods for such compression train an autoencoder end-to-end with an objective to strike a balance between distortion and rate. They are faced with the zero gradient issue due to quantization and the difficulty of estimating the rate accurately. Inspired by soft quantization, we represent quantization indices of feature maps with differentiable soft bits. This allows us to couple tightly the rate estimation with context-adaptive binary arithmetic coding. It also provides a differentiable distortion objective function. Experimental results show that our approach achieves the state-of-the-art compression performance among the learning-based schemes in terms of MS-SSIM and PSNR.
David Alexandre, Chih-Peng Chang, Wen-Hsiao Peng, Hsueh-Ming Hang
ICIP4
2019 Incorporating Luminance, Depth and Color Information by a Fusion-Based Network for Semantic Segmentation
abstract
Semantic segmentation has made encouraging progress due to the success of deep convolutional networks in recent years. Meanwhile, depth sensors become prevalent nowadays; thus, depth maps can be acquired more easily. However, there are few studies that focus on the RGB-D semantic segmentation task. Exploiting the depth information effectiveness to improve performance is a challenge. In this paper, we propose a novel solution named LDFNet, which incorporates Luminance, Depth and Color information by a fusion-based network. It includes a sub-network to process depth maps and employs luminance images to assist the depth information in processes. LDFNet outperforms the other state-of-art systems on the Cityscapes dataset, and its inference speed is faster than most of the existing networks. The experimental results show the effectiveness of the proposed multi-modal fusion network and its potential for practical applications.
Shang-Wei Hung, Shao-Yuan Lo, Hsueh-Ming Hang
ICIP3
2019 Session details: Vision in Multimedia
abstract
No abstract available.
Hsueh-Ming Hang
MMAsia1
2019 Exploring Semantic Segmentation on the DCT Representation
abstract
Typical convolutional networks are trained and conducted on RGB images. However, images are often compressed for memory savings and efficient transmission in real-world applications. In this paper, we explore methods for performing semantic segmentation on the discrete cosine transform (DCT) representation defined by the JPEG standard. We first rearrange the DCT coefficients to form a preferred input type, then we tailor an existing network to the DCT inputs. The proposed method has an accuracy close to the RGB model at about the same network complexity. Moreover, we investigate the impact of selecting different DCT components on segmentation performance. With a proper selection, one can achieve the same level accuracy using only 36% of the DCT coefficients. We further show the robustness of our method under the quantization errors. To our knowledge, this paper is the first to explore semantic segmentation on the DCT representation.
Shao-Yuan Lo, Hsueh-Ming Hang
MMAsia2
2019 Efficient Dense Modules of Asymmetric Convolution for Real-Time Semantic Segmentation
abstract
Real-time semantic segmentation plays an important role in practical applications such as self-driving and robots. Most semantic segmentation research focuses on improving estimation accuracy with little consideration on efficiency. Several previous studies that emphasize high-speed inference often fail to produce high-accuracy segmentation results. In this paper, we propose a novel convolutional network named Efficient Dense modules with Asymmetric convolution (EDANet), which employs an asymmetric convolution structure and incorporates dilated convolution and dense connectivity to achieve high efficiency at low computational cost and model size. EDANet is 2.7 times faster than the existing fast segmentation network, ICNet, while it achieves a similar mIoU score without any additional context module, post-processing scheme, and pretrained model. We evaluate EDANet on Cityscapes and CamVid datasets, and compare it with the other state-of-art systems. Our network can run with the high-resolution inputs at the speed of 108 FPS on one GTX 1080Ti.
Shao-Yuan Lo, Hsueh-Ming Hang, Sheng-Wei Chan, Jing-Jhih Lin
MMAsia2
2019 NCTU-GTAV360: A 360° Action Recognition Video Dataset
abstract
Despite many action recognition video datasets available right now, none of them are in the spherical projection. NCTU-GTAV360 is a new 360° action recognition video dataset captured from a game, Grand Theft Auto V (GTA V). The spherical video is obtained by stitching 24 views from various angles and combining them into a video. The benefit of using 360° cameras is that it can capture the entire surroundings using one single camera. We captured 200 locations within the Los Santos city (city name in the GTA V). This dataset should benefit researchers working on the spherical images, particularly the human action recognition research using machine learning or deep learning technique, which requires a large amount of training data and the associated ground-truth.
Sandy Ardianto, Hsueh-Ming Hang
MMSP2
2019 Multi-Class Lane Semantic Segmentation using Efficient Convolutional Networks
abstract
Lane detection plays an important role in a self-driving vehicle. Several studies leverage a semantic segmentation network to extract robust lane features, but few of them can distinguish different types of lanes. In this paper, we focus on the problem of multi-class lane semantic segmentation. Based on the observation that the lane is a small-size and narrow-width object in a road scene image, we propose two techniques, Feature Size Selection (FSS) and Degressive Dilation Block (DD Block). The FSS allows a network to extract thin lane features using appropriate feature sizes. To acquire fine-grained spatial information, the DD Block is made of a series of dilated convolutions with degressive dilation rates. Experimental results show that the proposed techniques provide obvious improvement in accuracy, while they achieve the same or faster inference speed compared to the baseline system, and can run at real-time on high-resolution images.
Shao-Yuan Lo, Hsueh-Ming Hang, Sheng-Wei Chan, Jing-Jhih Lin
MMSP2
2017 Learning-based human detection applied to RGB-D images
abstract
Accurate human detection is still a challenging topic due to complicated environments in the real world. In addition, the RGB-D cameras are becoming popular at reasonable price, such as Microsoft Kinect sensor, which provides both RGB and depth data. The depth information often helpful for detection. We adopt the R-CNN method in this paper, which combines the Selective Search technique to generate region proposals and the CNNs (Convolutional Neural Networks) to learn features. A depth map encoding technique (HHA) is adopted to match the CNNs format for learning features. The HHA and RGB images are our inputs. We propose several algorithms to combine their information in constructing various human detectors. Our information fusion structures include CNN, SVM together with PCA for features reduction. More accurate human detection results are shown with the aid of depth information.
Patrisia Sherryl Santoso, Hsueh-Ming Hang
ICIP2
2017 Online multiclass passive-aggressive learning on a fixed budget
abstract
This paper presents a budgetary learning algorithm for online multiclass classification. Based on the multiclass passive-aggressive learning with kernels, we introduce a dual perspective that gives rise to the proposed budgetary algorithm. Basically, the proposed algorithm limits the amount of data in use and fully exploits the available data on hand through optimization. The algorithm has both constant time and space complexities and thus can avoid the curse of kernelization. Experimental results with open datasets show that the proposed budgetary algorithm is competitive with state-of-the-art algorithms.
Chung-Hao Wu, Wei-Chen Hsi, Henry Horng-Shing Lu, Hsueh-Ming Hang
ISCAS4
2015 View synthesis for 3D video scene composition
abstract
Scene composition is a method widely used in movie and TV production. Merging two sets of 3D videos into one is a very challenging task. There are several main issues on video composition. Our focus is compositing two sets of videos with different camera motion parameters. The key techniques are the camera motion estimation and view synthesis technique used to produce the synthesized motion-compensated background video. We propose a refined backward warping technique for view synthesis and adopt the ICP algorithm to calculate the camera motion parameters.
Amanda Wang, Chun-Liang Chien, Hsueh-Ming Hang
MMSP3
2013 Virtual view synthesis using backward depth warping algorithm
abstract
The virtual view synthesis reference software offered by the MPEG standard committee adopts the forward warping technique in projecting the depth map from the reference view to the target (virtual) view location. Often, this warping process results in many artifacts, holes and cracks, due to quantization errors and occlusion. In this study, we propose a backward warping process to replace the forward warping process, and the artifacts (particularly the ones produced by quantization) are significantly reduced. The subjective quality of the synthesized virtual view images is thus much improved.
Du-Hsiu Li, Hsueh-Ming Hang, Yu-Lun Liu 0001
PCS2
2013 Quality assessment of 3D synthesized views with depth map distortion
abstract
Most existing 3D image quality metrics use 2D image quality assessment (IQA) models to predict the 3D subjective quality. But in a free viewpoint television (FTV) system, the depth map errors often produce object shifting or ghost artifacts on the synthesized pictures due to the use of Depth Image Based Rendering (DIBR) technique. These artifacts are very different from the ordinary 2D distortions such as blur, Gaussian noise, and compression errors. We thus propose a new 3D quality metric to evaluate the quality of stereo images that may contain artifacts introduced by the rendering process due to depth map errors. We first eliminate the consistent pixel shifts inside an object before the usual 2D metric is applied. The experimental results show that the proposed method enhances the correlation of the objective quality score to the 3D subjective scores.
Chang-Ting Tsai, Hsueh-Ming Hang
VCIP2
2012 Direction alignment algorithm for direction-adaptive discrete wavelet transform
abstract
2-D discrete wavelet transform (2-D DWT) represents non-vertical or non-horizontal edges inefficiently. Direction-adaptive discrete wavelet transform (DA-DWT) solves this problem by filtering along the directions of textures. DA-DWT partitions images into many blocks and finds the most suitable direction of each block. The block partition and direction information need to be transmitted as side information. A conventional DA-DWT often produces inconsistent directions of neighboring blocks and thus results in large amount of side information. In this paper, we propose a bottom-up direction alignment algorithm to align the block directions in local areas. Our algorithm can reduce 50% or more in side-information bits at the cost of negligible prediction error increase.
Chao-Hsiung Hung, Hsueh-Ming Hang
ICASSP2
2012 A reduced-complexity image coding scheme using decision-directed wavelet-based contourlet transform
Chao-Hsiung Hung, Hsueh-Ming Hang
J. Vis. Commun. Image Represent.2
2011 A triangular-warping based view synthesis scheme with enhanced artifact reduction for FTV
abstract
View synthesis is one key enabling technology for free-viewpoint television (FTV). It uses multiple image frames and depth maps to generate the intermediate view at nearly any arbitrary viewpoint. In our proposed 2-stage algorithm, we adopt triangular warping with a new feature point extraction scheme for the first texture mapping stage. The new feature point extraction scheme uses the correlation between the image luminance gradients and the depth map. In the second stage, we use the median filtering and the multi-band blending technique to reduce the artifacts on the image object boundaries caused by the discontinuity and the imperfection of depth maps. Experimental results show that our proposed method produces better visual quality in comparison with a latest view synthesis algorithm.
Chao-Hsuan Li, Jang-Jer Tsai, Hsueh-Ming Hang
ICIP3
2011 Multiview encoder parallelized fast search realization on NVIDIA CUDA
abstract
NVIDIA announced a powerful GPU architecture called Compute Unified Device Architecture (CUDA) in 2007, which is able to provide massive data parallelism under the SIMD architecture constraint. We use NVIDIA GTX-280 GPU system, which has 240 computing cores, as the platform to implement a very complicated video coding scheme, the Multiview Video Coding (MVC) scheme. MVC is an extension of H.264/MPEG-4 Part 10 AVC. It is an efficient video compression scheme; however, its computational complexity is very high. Two of its most time- consuming components are motion estimation (ME) and disparity estimation (DE). In this thesis, we propose a fast search algorithm, called multithreaded one-dimensional search (MODS). It can be used to do both the ME and the DE operations. We implement the integer-pel ME and DE processes with MODS on the GTX-280 platform. The speedup ratio can be 89 times faster than the CPU only configuration. Even when the fast search algorithm of the original JMVC is turned on, the MODS version on CUDA can still be 20 times faster.
Chih-Te Lu, Hsueh-Ming Hang
VCIP2
2011 Fast mode decision algorithm for Residual Quadtree coding in HEVC
abstract
High Efficiency Video Coding (HEVC) is an in- progress next generation video coding standard. It has a similar basic structure to the ITU/MPEG H.264/AVC coder but with further enhancement on each coding tool to increase compression efficiency. Thus, a Residual Quadtree (RQT) coding scheme is adopted in the conventional transform coding procedure. The cost is the additional computation complexity when compared to the fixed-size transform. In this paper, we design a fast algorithm for Residual Quadtree mode decision. Considering the Rate-Distortion efficiency, we replace the original depth-first mode decision process by a Merge-and-Split decision process. Furthermore, because a substantial numbers of zero-blocks are produced after quantization and mode decision, we reduce the unnecessary computation by using the inheritance property of zero-blocks. In addition, for nonzero-blocks, two early termination schemes are developed for both TU Merge and TU Split procedures, respectively. Comparing to HM 2.0, our method saves the RQT encoding time from 42% to 55% for a number of test videos with negligible coding loss.
Su-Wei Teng, Hsueh-Ming Hang, Yi-Fu Chen
VCIP2
2011 One-Sided ρ-GGD Source Modeling and Rate-Distortion Optimization in Scalable Wavelet Video Coder
abstract
We develop an accurate source model, one-sided$\rho$-generalized Gaussian distribution (GGD), for approximating the residual signals in scalable wavelet video coding. An efficient piecewise linear expression is suggested to estimate the shape parameter of the one-sided$\rho$-GGD. We also improve the model accuracy in matching the real data by modifying the$\rho$parameter estimation formula. Continuing our previous work on developing the motion information gain metric to measure the motion information efficiency, we now incorporate the one-sided$\rho$-GGD model in the cost function, which is used for deciding the motion vectors and motion estimation mode in scalable wavelet video coding. Compared with the conventional Lagrangian optimization, our simulation results show that the new mode decision method generally improves the peak signal-to-noise ratio performance in the combined signal-to-noise ratio and temporal scalability cases.
Chia-Yang Tsai, Hsueh-Ming Hang
IEEE Trans. Circuits Syst. Video Technol.2
2011 Fast Bi-Directional Prediction Selection in H.264/MPEG-4 AVC Temporal Scalable Video Coding
abstract
In this paper, we propose a fast algorithm that efficiently selects the temporal prediction type for the dyadic hierarchical-B prediction structure in the H.264/MPEG-4 temporal scalable video coding (SVC). We make use of the strong correlations in prediction type inheritance to eliminate the superfluous computations for the bi-directional (BI) prediction in the finer partitions, 16×8/8×16/8×8 , by referring to the best temporal prediction type of 16 × 16. In addition, we carefully examine the relationship in motion bit-rate costs and distortions between the BI and the uni-directional temporal prediction types. As a result, we construct a set of adaptive thresholds to remove the unnecessary BI calculations. Moreover, for the block partitions smaller than 8 × 8, either the forward prediction (FW) or the backward prediction (BW) is skipped based upon the information of their 8 × 8 partitions. Hence, the proposed schemes can efficiently reduce the extensive computational burden in calculating the BI prediction. As compared to the JSVM 9.11 software, our method saves the encoding time from 48% to 67% for a large variety of test videos over a wide range of coding bit-rates and has only a minor coding performance loss.
Hung-Chih Lin, Hsueh-Ming Hang, Wen-Hsiao Peng
IEEE Trans. Image Process.2
2010 Fast algorithm on selecting bi-directional prediction type in H.264/AVC scalable video coding
abstract
In this paper, we propose a fast algorithm designed for the hierarchical prediction structure in H.264/MPEG scalable video coding (SVC). It can efficiently reduce the heavy computational burden in performing the bi-directional (BI) prediction. We carefully examine the correlation of motion-rate costs and distortion values between BI and FW (forward) and BW (backward) prediction types. Then we construct an adaptive threshold to remove the unnecessary BI calculations. Comparing to JSVM 9.11, our method saves the encoding time from 44% to 54% for a number of test videos with negligible coding loss.
Hung-Chih Lin, Hsueh-Ming Hang
ISCAS2
2010 A fast graph cut algorithm for disparity estimation
abstract
In this paper, we propose a fast graph cut (GC) algorithm for disparity estimation. Two accelerating techniques are suggested: one is the early termination rule, and the other is prioritizing the α-β swap pair search order. Our simulations show that the proposed fast GC algorithm outperforms the original GC scheme by 210% in the average computation time while its disparity estimation quality is almost similar to that of the original GC.
Cheng-Wei Chou, Jang-Jer Tsai, Hsueh-Ming Hang, Hung-Chih Lin
PCS3
2010 A rate-distortion analysis on motion prediction efficiency and mode decision for scalable wavelet video coding
Chia-Yang Tsai, Hsueh-Ming Hang
J. Vis. Commun. Image Represent.2
2010 Fast Context-Adaptive Mode Decision Algorithm for Scalable Video Coding With Combined Coarse-Grain Quality Scalability (CGS) and Temporal Scalability
abstract
To speed up the H.264/MPEG scalable video coding (SVC) encoder, we propose a layer-adaptive intra/inter mode decision algorithm and a motion search scheme for the hierarchical B-frames in SVC with combined coarse-grain quality scalability (CGS) and temporal scalability. To reduce computation but maintain the same level of coding efficiency, we examine the rate-distortion (R-D) performance contributed by different coding modes at the enhancement layers (EL) and the mode conditional probabilities at different temporal layers. For the intra prediction on inter frames, we can reduce the number of Intra4×4/Intra 8×8 prediction modes by 50% or more, based on the reference/base layer intra prediction directions. For the EL inter prediction, the look-up tables containing inter prediction candidate modes are designed to use the macroblock (MB) coding mode dependence and the reference/base layer quantization parameters (Qp). In addition, to avoid checking all motion estimation (ME) reference frames, the base layer (BL) reference frame index is selectively reused. And according to the EL MB partition, the BL motion vector can be used as the initial search point for the EL ME. Compared with Joint Scalable Video Model 9.11, our proposed algorithm provides a 20× speedup on encoding the EL and an 85% time saving on the entire encoding process with negligible loss in coding efficiency. Moreover, compared with other fast mode decision algorithms, our scheme can demonstrate a 7-41% complexity reduction on the overall encoding process.
Hung-Chih Lin, Wen-Hsiao Peng, Hsueh-Ming Hang
IEEE Trans. Circuits Syst. Video Technol.3
2010 On the Design of Pattern-Based Block Motion Estimation Algorithms
abstract
Pattern-based block motion estimation (PBME) is a critical element in the contemporary video coding system because it typically dominates the coding efficiency and the computing power. Therefore, many proposals have been suggested to reduce its computational complexity, but most of them are devised based on experimental data or heuristic ideas. In this letter, we look into every component of a typical PBME algorithm and fine tune the major components systematically to achieve the optimal or nearly optimal results. Our methodology is developed based on our proposed analytical model together with statistical tools. First, we use the analytic model to analyze and design effective genetic-algorithm-based search patterns. Moreover, we propose an adaptive switching strategy that dynamically switches between two search patterns. Second, we extend our PBME model to evaluate the efficiency of starting (initial search) points. A near optimal set of starting points is progressively identified. Last, we study the early termination threshold technique and suggest a metric in selecting an effective threshold. An accurate threshold mechanism is thus constructed. Combining all these techniques, we develop a PBME algorithm that outperforms most popular algorithms.
Jang-Jer Tsai, Hsueh-Ming Hang
IEEE Trans. Circuits Syst. Video Technol.2
2009 Fast temporal prediction selection for H.264/AVC scalable video coding
abstract
In this paper, we propose a fast algorithm that selects the temporal prediction type for the dyadic hierarchical prediction structure in scalable video coding (SVC). We make use of the strong correlations in the large block partitions to eliminate the unnecessary computations for bi-directional prediction. Moreover, based upon the information of an 8 × 8 partition, either forward or backward prediction is skipped for its smaller block partitions. Comparing to the JSVM 9.11, our method saves the encoding time from 50% to 60% for a number of test videos over a typical range of coding bit-rates and its coding penalty is negligible.
Hung-Chih Lin, Hsueh-Ming Hang, Wen-Hsiao Peng
ICIP2
2009 Decision-directed Adaptive Wavelet Image Coding with Diretional Decomposition
abstract
In this paper, we propose an adaptive wavelet coding scheme. A decision method is designed to identify whether the LH, HL, and HH wavelet subbands are suitable for directional decomposition. The basic concept behind this decision method is to detect the large impulses in the frequency domain. The transformed data are coded by the bit-plane arithmetic code to achieve scalability. Experimental results show that this adaptive scheme has better results than the 1-D wavelet transform and the original wavelet-based contourlet scheme that applies the directional transform to all the LH, HL and HH subbands.
Chao-Hsiung Hung, Hsueh-Ming Hang
ISCAS2
2009 Modeling of Pattern-Based Block Motion Estimation and Its Application
abstract
Pattern-based block motion estimation (PBME) is one of the most widely adopted compression tools in the contemporary video coding systems. However, despite that many researches have studied PBME, few have yet attempted to construct an analytical model that can explain the underneath principle and mechanism of various PBME algorithms. In this paper, we propose a statistical PBME model that consists of two components: 1) a statistical probability distribution for motion vectors and 2) the minimal number of search points (so-called weighting function) achieved by a search algorithm. We first verify the accuracy of the proposed model by checking the experimental data. Then, an application example using this model is shown. Starting from an ideal weighting function, we devise a novel genetic rhombus pattern search (GRPS) to match the design target. Simulations show that, comparing to the other popular search algorithms, GRPS reduces the average search points for more than 20% and, in the meanwhile, it maintains a similar level of coded image quality.
Jang-Jer Tsai, Hsueh-Ming Hang
IEEE Trans. Circuits Syst. Video Technol.2
2009 Consistent Picture Quality Control Strategy for Dependent Video Coding
abstract
Typically, a video rate control algorithm minimizes the average distortion (denoted as MINAVE) at the cost of large temporal quality variation, especially for videos with high motion and frequent scene changes. To alleviate the negative effect on subjective video quality, another criterion that restricts a small amount of quality variation among adjacent frames is preferred for practical applications. As pointed out by [20], although some existing proposals can produce consistent quality videos, they often fail to fully utilize the available bits to minimize the global total distortion. In this paper, we would like to achieve the triple goal of consistent quality video, minimizing the total distortion, and meeting the bit budget strictly all at the same time on the interframe dependent coding structure. Two approaches are taken to accomplish this goal. In the first algorithm, a trellis-based framework is proposed. One of our contributions is to derive an equivalent condition between the distortion minimization problem and the budget minimization problem. Second, our trellis state (tree node) is defined in terms of distortion, which facilitates the consistent quality control. Third, by adjusting one key parameter in our algorithm, a solution in between the MINAVE and the constant quality criteria can be obtained. The second approach is to combine the Lagrange multipliers method together with the consistent quality control. The PSNR performance is degraded slightly but the computational complexity is significantly reduced.Simulation results show that both our approaches produce a much smaller PSNR variation at a slight average PSNR loss as compared to the MPEG JM rate control. When they are compared to the other consistent quality proposals, only the proposed algorithms can strictly meet the target bit budget requirement (no more, no less) and produce the largest average PSNR at a small PSNR variation.
Kao-Lung Huang, Hsueh-Ming Hang
IEEE Trans. Image Process.2
2008 Image coding using short wavelet-based contourlet transform
abstract
In this paper, we propose two enhanced elements on the wavelet-based contourlet transform (WBCT) image coding scheme. A short-length 2-D non-separable directional filter bank is designed for significantly reducing the calculation. New zero coding context models are designed to improve coding efficiency. Experiments show that the proposed compression scheme produces better visual quality images, particularly, with fine and regular texture at low rates. The short-length directional filter bank performs nearly as well as their long-length counterpart.
Chao-Hsiung Hung, Hsueh-Ming Hang
ICIP2
2008 On modeling genetic pattern search for block motion estimation
abstract
Pattern search algorithms, such as diamond search, hexagonal search and their variations, have been widely adopted by the block matching motion estimations in the modern video encoding systems. Recently we propose a weighting function (WF) to model the number of search points of a pattern search. Yet, WF fails to properly describe the behavior of the genetic pattern search algorithms due to some over-simplifications in their models. Therefore, we propose a refined weighting function (RWF) to more accurately describe both genetic and non-genetic pattern searches. In addition, we propose a new search algorithm, namely, the momentum directed genetic rhombus pattern search (MD-GRPS). It can accelerate the previous genetic rhombus pattern search by 8% on the average and this concept can be applied to the other genetic pattern searches.
Jang-Jer Tsai, Hsueh-Ming Hang
ICIP2
2008 H.264/AVC motion estimation implmentation on Compute Unified Device Architecture (CUDA)
abstract
Due to the rapid growth of graphics processing unit (GPU) processing capability, using GPU as a coprocessor to assist the central processing unit (CPU) in computing massive data becomes essential. In this paper, we present an efficient block-level parallel algorithm for the variable block size motion estimation (ME) in H.264/AVC with fractional pixel refinement on a computer unified device architecture (CUDA) platform, developed by NVIDIA in 2007. The CUDA enhances the programmability and flexibility for general-purpose computation on GPU. We decompose the H.264 ME algorithm into 5 steps so that we can achieve highly parallel computation with low external memory transfer rate. Experimental results show that, with the assistance of GPU, the processing time is 12 times faster than that of using CPU only.
Wei-Nien Chen, Hsueh-Ming Hang
ICME2
2008 A fast mode decision algorithm with macroblock-adaptive rate-distortion estimation for intra-only scalable video coding
abstract
In this paper, we propose a fast mode decision algorithm with macroblock-adaptive rate-distortion (R-D) estimation for intra-only scalable video coding (SVC). We make use of the log-linear R-D relationship of inter-dependent layers to predict the better performer among the Intra4times4 and Intra8times8 prediction types at the enhancement layers. Based upon the base-layer chosen prediction type, we can further reduce the number of candidate modes. In addition, to ensure the best trade-off between complexity and coding efficiency, the Intra16times16 prediction is retained and enabled only for coding high-resolution videos with smooth image contents. Comparing to the joint scalable video model v.8 (JSVM 8), an encoder time saving from 49% to 64% depending on the encoder configurations, is achieved with negligible penalty in coding efficiency.
Hung-Chih Lin, Wen-Hsiao Peng, Hsueh-Ming Hang
ICME3
2008 rho-GGD source modeling for wavelet coefficients in image/video coding
abstract
ρ-GGD source model is proposed in this paper. The probability distribution of wavelet coefficients has been previously modeled as Laplacian or generalized Gaussian distribution (GGD). The Laplacian model is simple in calculation but not accurate; on the other hand, GGD is accurate but requires a very complicated modeling procedure. In this paper, we introduce a new parameter ρ into the GGD source model, where ρ is the probability of zero-value coefficients. The shape parameter of the GGD model can then be easily estimated from ρ and the standard deviation. Moreover, a piecewise linear approximation method is proposed to further reduce complexity. Our experiments show that the ρ-GGD model has high accuracy and consistent performance for modeling the wavelet coefficient pdfs for both spatial 2-D DWT and interframe wavelet video cases.
Chia-Yang Tsai, Hsueh-Ming Hang
ICME2
2007 On Adaptive Pattern Selection for Block Motion Estimation Algorithms
abstract
Pattern-based block motion estimation (PBME) algorithm has been widely adopted in digital video coding systems. Due to the large characteristics variations among video sequences, adaptive PBME algorithms that switch search patterns have been proposed. However, most adaptive search algorithms are heuristically designed based on the experimental data. In this paper, we like to construct an analytical model and explore the problem systematically. Also, we propose an adaptive genetic pattern search algorithm (AGPS). Simulations show that the proposed AGPS in average outperforms the existing popular search algorithms quite significantly in speed, while the peak signal noise ratio (PSNR) quality is maintained at the same level.
Jang-Jer Tsai, Hsueh-Ming Hang
ICASSP (1)2
2007 Layer-Adaptive Mode Decision and Motion Search for Scalable Video Coding with Combined Coarse Granular Scalability (CGS) and Temporal Scalability
abstract
In this paper, we propose a layer-adaptive mode decision algorithm and a motion search scheme for the scalable video coding (SVC) with combined coarse granular scalability (CGS) and temporal scalability. To speed up the encoder while minimizing the loss in coding efficiency, our layer-adaptive mode decision recursively refers to the prediction modes and quantization parameter of the reference/base layer to minimize the number of modes tested at the enhancement layer. Moreover, our motion search scheme adaptively reuses the reference frame indices of the base layer and determines the initial search point using the motion vector at the base layer or the motion vector predictor at the enhancement layer. As compared with JSVM 8, the proposed algorithms provide up to 75% overall time saving and more than 85% time reduction for encoding enhancement layers with negligible loss in coding efficiency.
Hung-Chih Lin, Wen-Hsiao Peng, Hsueh-Ming Hang, Wen-Jen Ho
ICIP (2)3
2007 Acceleration and Implementation of JPEG2000 Encoder on TI DSP Platform
abstract
JPEG2000 provides excellent compression performance and fine granularity scalability but at the cost of high computational complexity. We propose two speed-up techniques and use the TI DSP optimization tools to accelerate the Tierl module. We eliminate the unnecessary checking cycles by recording the NBC (need-to-be-coded) samples on a list. Furthermore, the sample index is reordered to facilitate fast execution. In the DSP implementation of the proposed methods, we use code acceleration techniques, cache memory allocation, and TI DSP compiler-level optimization tools. Even when the original program is compiled with the same DSP optimization tools and proper cache assignment, our fast algorithm can still reduce the computation by 45%.
Chien-Chih Liu, Hsueh-Ming Hang
ICIP (3)2
2007 A Genetic Rhombus Pattern Search for Block Motion Estimation
abstract
Pattern-based block motion estimation (PBME) is one of the most effective yet computational intensive tools in digital video coding standards. Many PBME algorithms have been proposed but few papers have investigated in depth the fundamental characteristics of various PBME algorithms. Why one search algorithm outperforms the others? In this paper, we propose a systematic approach to examine this problem. The minimal numbers of search points achievable by a PBME algorithm form a discrete function in the search area. We analyze this so-called weighting function and suggest an ideal target. Then, we design a genetic rhombus pattern search (GRPS) to match this ideal weighting function. Simulations show that, comparing to the other popular search algorithms, GRPS reduces the average search points for more than 20% while it maintains a similar level of coded image PSNR quality.
Jang-Jer Tsai, Hsueh-Ming Hang
ISCAS2
2006 Algorithms and DSP implementation of H.264/AVC
abstract
This survey paper intends to provide a comprehensive coverage of the techniques that are pertinent to the processor-based implementation of H.264/AVC video codec, particularly on DSP. Most of this paper is devoted to the computationally efficient algorithms, or the fast algorithms. Fast algorithms for motion estimation, intra-prediction and mode decision are described to reduce the computational complexity. In addition, in order to port the H.264/AVC codec to DSP, we also outline the basic principles of DSP code optimization
Hung-Chih Lin, Yu-Jen Wang, Kai-Ting Cheng, Shang-Yu Yeh, Wei-Nien Chen, Chia-Yang Tsai, Tian-Sheuan Chang, Hsueh-Ming Hang
ASP-DAC8
2006 Adding selective enhancement in scalable video coding for region-of-interest functionality
abstract
In this paper, we propose a lossless, graceful, and arbitrary-shaped region-of-interest (ROI) functionality for the scalable video coding in MPEG-4 Part 10 Amd. 1. Specifically, we propose a prioritized block coding scheme and a layer remapping technique. Within an enhancement layer, the prioritized block coding reshuffles the transform coefficients for ROI. Moreover, the layer remapping technique prioritizes different ROI among enhancement layers. For graceful and arbitrary-shaped properties, additional syntax are coded at the levels of slice, layer, and macroblock. To minimize overhead, an efficient representation for priority information is proposed. Experimental results show that our schemes can offer the ROI functionality over a wide range of bit rates while maintaining coding efficiency.
Wen-Hsiao Peng, Tihao Chiang, Hsueh-Ming Hang
ISCAS3
2006 Cascaded trellis-based rate-distortion control algorithm for MPEG-4 advanced audio coding
abstract
In this paper, a few low-complexity and high-performance rate-distortion control algorithms for MPEG-4 Advanced Audio Coding (AAC) are proposed. One key element in producing good quality compressed audio particularly at medium and low rates is a high performance rate-distortion controller in the audio encoder. Although the trellis-based rate-distortion control algorithms previously proposed can achieve a praiseworthy performance, their computational complexity is extremely high. Therefore, for practical applications, it is very desirable to achieve a similar performance at a much lower complexity. Two types of techniques are proposed in this paper to reduce the computational burden of the trellis-based algorithms. One is splitting a very heavy calculation stage into two sequential steps with much less computation. The other is reducing the candidates in the trellis for parameter search. Together, when applicable, our approach achieves a similar coding performance (audio quality) but requires less than 1/1000 complexity in computation.
Cheng-Han Yang, Hsueh-Ming Hang
IEEE Trans. Speech Audio Process.2
2006 A Context Adaptive Bit-Plane Coder With Maximum-Likelihood-Based Stochastic Bit-Reshuffling Technique for Scalable Video Coding
abstract
In this paper, we propose a context adaptive bit-plane coding (CABIC) with a stochastic bit reshuffling (SBR) scheme to deliver higher coding efficiency and better subjective quality for fine granular scalable (FGS) video coding. Traditional bit-plane coding in FGS algorithm suffers from poor coding efficiency and subjective quality. To improve coding efficiency, our CABIC constructs context models based on both the energy distribution in a block and the spatial correlations in the adjacent blocks. Moreover, it exploits the context across bit-planes to save side information. To improve subjective quality, our SBR reorders the coefficient bits by their estimated rate-distortion performance. Particularly, we model transform coefficients with Laplacian distributions and incorporate them into the context probability models for content-aware parameter estimation. Moreover, our SBR is implemented with a dynamic priority management that uses a low-complexity dynamic memory organization. Experimental results show that our CABIC improves the PSNR by 0.5/spl sim/1.0 dB at medium and high bit rates. While maintaining similar or even higher coding efficiency, our SBR improves the subjective quality.
Wen-Hsiao Peng, Tihao Chiang, Hsueh-Ming Hang, Chen-Yi Lee
IEEE Trans. Multim.3
2005 Advances of MPEG Scalable Video Coding Standard
Wen-Hsiao Peng, Chia-Yang Tsai, Tihao Chiang, Hsueh-Ming Hang
KES (4)4
2004 Enhanced motion estimation for interframe wavelet video coding
abstract
An enhanced motion estimation scheme is incorporated into the interframe wavelet coding architecture in this paper. Interframe wavelet coding has the advantage of SNR, temporal, and spatial scalability and is a potential candidate for the on-going MPEG-21 scalable video coding standard. Motion-compensated temporal filtering (MCTF) is one of its essential components. Therefore, motion estimation plays an important role in deciding the coding performance. In this paper, we modified the motion estimation syntax/scheme originally specified in the MPEG Advanced Video Coding (AVC) and use it in the interframe wavelet structure. Besides, the techniques of l-block, bidirectional motion estimation, /spl lambda/-value adjustment and motion information partitioning are employed. Simulation results show very promising performance particularly on subjective quality.
Chia-Yang Tsai, Han-Kuang Hsu, Hsiang-Cheh Huang, Hsueh-Ming Hang, Guo-Zua Wu
ICIP4
2004 Content-based image retrieval using both positive and negative feedback
abstract
Satisfactory content-based search has long been considered a difficult task. One critical step in the content-based search is to estimate the user intention (perception) based on the query images. Our proposal is developed based on the combined weighted low-level image features. One distinct concept of our algorithm is that a sparse (scattered) feature is considered to be less important (which is not necessarily perceptually dissimilar). The other concept is that we define the image feature stability and include it in calculating the similarity measure. Yet, the third concept is using negative feedback as a pruning criterion to improve searching accuracy. Finally, quantitative simulation results are used to show the effectiveness of these concepts.
Feng-Cheng Chang, Hsueh-Ming Hang
ICME2
2004 Motion information scalability for MC-EZBC
Sam S. Tsai, Hsueh-Ming Hang
Signal Process. Image Commun.2
2003 Efficient bit assignment strategy for perceptual audio coding
abstract
For the purpose of efficient audio coding at low rates, a new bit allocation strategy is proposed in this paper. The basic idea behind this approach is "Give bits to the band with the maximum NMR-Gain/bit" or "Retrieve bits from the band with the maximum bits/NMR-Loss". The notion of "bit-use efficiency" is suggested and it can be employed to construct a bit assignment algorithm operated at band-level as compared to the traditional frame-level bit assignment methods. Based on this strategy a new bit assignment scheme, called Max-BNLR, is designed for the MPEG-4 AAC. Simulation results show that the performance of the Max-BNLR scheme is significantly better than that of the MPEG-4 AAC Verification Model (VM) and is close to that of TB-ANMR, which is the (nearly) optimal solution. Moreover, the Max-BNLR scheme has the advantages of low computational complexity compared to TB-ANMR.
Cheng-Han Yang, Hsueh-Ming Hang
ICASSP (5)2
2001 Vector quantization based on genetic simulated annealing
Hsiang-Cheh Huang, Jeng-Shyang Pan 0001, Zheming Lu 0001, Sheng-He Sun, Hsueh-Ming Hang
Signal Process.5
1998 A Novel Approach for Edge Orientation Determination based on Pixel Pair Matching
abstract
A novel edge orientation determination algorithm is proposed. It can detect edges and decide very precisely their orientations based on a pixel pair matching technique. Three image edge structure measures, profile diversity, edge convexity and edge continuity, are designed to facilitate the determination process. The basic algorithm outputs an integer-pixel orientation vector for each detected edge and this orientation vector can be refined to subpixel accuracy by a polynomial fitting method. Simulation results are included to show the advantages of this approach.
Hou-Chun Ting, Hsueh-Ming Hang
ICIP (2)2
1997 The Impact of Rate Control Algorithms on Video Codec Hardware Design
abstract
This paper presents an evaluation of rate control algorithms from a system-level VLSI design viewpoint. Rate control in video coding has a significant influence on the coded bits and image quality. Many rate control algorithms have been proposed mainly focusing on the optimal rate-distortion performance without considering their overall performance on the VLSI implementation. However, a system-level designer should design an algorithm not only good in performance but also good in implementation. Three different types of popular rate control algorithms have been analyzed based on their picture quality, the internal buffer size and the hardware cost. The methodology and results presented should provide useful guidelines for selecting an appropriate rate control algorithm for system-level VLSI design.
Sheu-Chih Cheng, Hsueh-Ming Hang
ICIP (2)2
1997 Adaptive Piecewise Linear Bits Estimation Model for MPEG Based Video Coding
Jia-Bao Cheng, Hsueh-Ming Hang
J. Vis. Commun. Image Represent.2
1997 A New Motion Estimation Method Using Frequency Components
Yung-Ming Chou, Hsueh-Ming Hang
J. Vis. Commun. Image Represent.2
1997 Edge Preserving Interpolation of Digital Images Using Fuzzy Inference
Hou-Chun Ting, Hsueh-Ming Hang
J. Vis. Commun. Image Represent.2
1997 Source model for transform video coder and its application. II. Variable frame rate coding
abstract
In the first part of this paper, we derived a source model describing the relationship between bits, distortion, and quantization step size for transform coders. Based on this source model, a variable frame rate coding algorithm is developed. The basic idea is to select a proper picture frame rate to ensure a minimum picture quality for every frame. Because our source model can predict approximately the number of coded bits when a certain quantization step size is used, we could predict the quality and bits of coded images without going through the entire real-coding process. Therefore, we could skip the right number of picture frames to accomplish the goal of constant image quality. Our proposed variable frame rate coding schemes are simple but quite effective as demonstrated by simulation results. The results of using another variable frame rate scheme, Test Model for H.263 (TMN-5), and the results of using a fixed frame rate coding scheme, Reference Model 8 for H.261 (RM8), are also provided for comparison.
Jiann-Jone Chen, Hsueh-Ming Hang
IEEE Trans. Circuits Syst. Video Technol.2
1997 A comparison of block-matching algorithms mapped to systolic-array implementation
abstract
This paper presents an evaluation of several well-known block-matching motion estimation algorithms from a system-level very large scale integration (VLSI) design viewpoint. Because a straightforward block-matching algorithm (BMA) demands a very large amount of computing power, many fast algorithms have been developed. However, these fast algorithms are often designed to merely reduce arithmetic operations without considering their overall performance in VLSI implementation. Three criteria are used to compare various block-matching algorithms: (1) silicon area, (2) input/output requirement, and (3) image quality. A basic systolic array architecture is chosen to implement all the selected algorithms. The purpose of this study is to compare these representative BMAs using the aforementioned criteria. The advantages/disadvantages of these algorithms in terms of their hardware tradeoff are discussed. The methodology and results presented provide useful guidelines to system designers in selecting a BMA for VLSI implementation.
Sheu-Chih Cheng, Hsueh-Ming Hang
IEEE Trans. Circuits Syst. Video Technol.2
1997 Source model for transform video coder and its application. I. Fundamental theory
abstract
A source model describing the relationship between bits, distortion, and quantization step sizes of a large class of block-transform video coders is proposed. This model is initially derived from the rate-distortion theory and then modified to match the practical coders and real image data. The realistic constraints such as quantizer dead-zone and threshold coefficient selection are included in our formulation. The most attractive feature of this model is its simplicity in its final form. It enables us to predict the bits needed to encode a picture at a given distortion or to predict the quantization step size at a given bit rate. There are two aspects of our contribution: one, we extend the existing results of rate-distortion theory to the practical video coders, and two, the nonideal factors in real signals and systems are identified, and their mathematical expressions are derived from empirical data. One application of this model, as shown in the second part of this paper, is the buffer/quantizer control on a CCITT P/spl times/64 k coder with the advantage that the picture quality is nearly constant over the entire picture sequence.
Hsueh-Ming Hang, Jiann-Jone Chen
IEEE Trans. Circuits Syst. Video Technol.1
1997 An adaptive inverse halftoning algorithm
abstract
A class of inverse halftoning algorithms that recovers grayscale (continuous-tone) images from halftone images is proposed. The basic structure is an optimized linear filter. Then, a properly designed adaptive postprocessor is employed to enhance the recovered image quality. Finally, a multistage space-varying algorithm is developed that uses the basic linear filter structure as before but with spatially adaptive parameters.
Li-Ming Chen 0004, Hsueh-Ming Hang
IEEE Trans. Image Process.2
1996 Combined quantizer and linear error control code design for noisy channels
abstract
Most optimal quantizer design algorithms do not take into account the changes of the channel characteristics due to the inserted channel coder. The overall channel characteristics including the channel coder is examined and an approximation model is proposed. Based on this model, a method for designing a quantizer (source coder) and the error control code together to achieve the best overall performance is proposed. Preliminary, simulation results reinforce the speculation that the error control codes would be useful only when the raw error rate is below a certain value.
Chi-Hsi Su, Hsueh-Ming Hang
ICIP (3)2
1995 Inverse halftoning of scanned images
abstract
Our goal in this research is to find a good inverse halftoning algorithm that recovers gray-scale images from the scanned images. To this aim, we develop the printer and the scanner models and two types of reconstruction methods. In the first method, the reconstruction filter is derived directly from the scanned data and the ideal original gray-scale image. In the second method, the forward halftoning, the printing and the scanning processes are reversed one by one. Either approach seems to produce reasonably good results.
Tsi-Yi Chao, Hsueh-Ming Hang
ICIP (3)2
1995 Adaptive piecewise linear bits estimation model for MPEG based video coding
abstract
We propose an adaptive piecewise linear bits estimation model whose structure is similar to a tree-structured piecewise linear filter. Each node in the tree is associated with a linear relationship between macroblock bits and activity/stepsize. The parameters in this tree structure are adjusted by a modified LMS algorithm. Computer simulation results indicate that the adaptive bits model is able to precisely estimate the compressed bits regardless how the image contents vary along time. Also, when compared to the table-look-up bits model derived based on cluster analysis, the adaptive piecewise linear bits model has a much lower complexity to achieve about the same high performance.
Jia-Bao Cheng, Hsueh-Ming Hang
ICIP2
1994 A Transform Video Coder Source Model and Its Application
abstract
A source model describes the relationship between the bits, distortion, and quantization step sizes of a large class of block-transform video coder is proposed. This model is derived from rate-distortion theory, and verified by real images. It enables us to predict the bits needed to encode a picture for a given distortion or to adjust the quantization scales of a coder for a given bit rate. Based on this derived model, a variable frame rate coding algorithm is developed. It can be used to control the frame rate of a coder to ensure a minimum picture quality of every frame. Simulation results indicate that improved performance is obtained by using this variable frame rate coding scheme when compared to the simple approach of varying the quantization scale linearly proportional to the encoder buffer level.>
Jiann-Jone Chen, Hsueh-Ming Hang
ICIP (2)2
1994 Inverse Halftoning for Monochrome Pictures
abstract
A class of inverse halftoning methods together with post-processing is proposed in this paper. In order to optimize the inverse operation of halftoning, we adopt the inverse modeling concept to find the least-squares solution of this problem. And then, the knowledge of image characteristics helps us in designing improved algorithms that produce the most promising pictures to our eyes. They are conceptually rather different from the traditional inverse halftoning methods that use simple low-pass filters. Furthermore, we develop a general space-varying sliding-window filter (SV-SWF) scheme. The goal is to be able to handle various types of pictures using the same basic inverse halftoning structure together with spatially adaptive parameters. This universal inverse halftoning operator is achieved by using an adequate classification scheme to separate image data into different groups, each corresponding to a set of pre-trained parameters. Due to the additional information in the inputs and the outputs utilized in our processing, better reconstructed images are obtained.>
Li-Ming Chen 0004, Hsueh-Ming Hang
ICIP (2)2
1994 Improved Two-Layer Coding Schemes for Motion Picture Sequences
abstract
This paper proposes several improved schemes on the two-layer codec. In a two-layer coder, the packets produced by the base layer are set to the high priority and those produced by the second layer are low. Three improved schemes on the base layer are proposed. The base layer bits and/or the total bits have been significant reduced by these improved algorithms. In addition, the base images have been also remarkably enhanced.>
Shang-Pin Chang, Tsorng-Yang Mei, Hsueh-Ming Hang
ISCAS3
1993 Transform-domain postprocessing of DCT-coded images
abstract
Data compression algorithms are developed to transmit massive image data under limited channel capacity. When a channel rate is not sufficient to transmit good quality compressed images, a degraded image after compression is reconstructed at the decoder. In this situation, a postprocessor can be used to improve the receiver image quality. Ideally, the objective of postprocessing is to restore the original pictures from the received distorted pictures. However, when the received pictures are heavily distorted, there may not exist enough information to restore the original images. Then, what a postprocessor can do is to reduce the subjective artifact rather than to minimize the differences between the received and the original images. In this paper, we propose two post processing techniques, namely, error pattern compensation and inter-block transform coefficient adjustment. Since Discrete Cosine Transform (DCT) coding is widely adopted by the international image transmission standards, our postprocessing schemes are proposed in the DCT domain. When the above schemes are applied to highly distorted images, quite noticeable subjective improvement can be observed.
Chung-Nan Tien, Hsueh-Ming Hang
VCIP2
1991 Digital HDTV compression using parallel motion-compensated transform coders
abstract
The authors suggest a parallel processing structure using the proposed international standard for visual telephony (CCITT P*64 kbs standard) as processing elements, to compress digital high definition television (HDTV) pictures. The basic idea is to partition an HDTV picture, in space or in frequency, into smaller sub-pictures and then compress each sub-picture using a CCITT P*64 kbs coder. This seems to be a cost-effective solution to the HDTV hardware. Since each sub-picture is processed by an independent coder, without coordination these coded sub-pictures may have unequal picture quality. To maintain a uniform quality HDTV picture, the following two issues are studied: sub-channel control strategy (bits allocated to each sub-picture); and quantization and buffer control strategy for individual sub-picture coders. Algorithms to resolve these problems and their computer simulations are presented.>
Hsueh-Ming Hang, Riccardo Leonardi, Barry G. Haskell, Robert L. Schmidt, Hemant Bheda, Joseph H. Othmer
IEEE Trans. Circuits Syst. Video Technol.1
1990 Digital HDTV compression at 44 Mbps using parallel motion-compensated transform coders
abstract
High Definition Television (HDTV) promises to offer wide-screen, much better quality pictures as compared to the today’s television. However, without compression a digital HDTV channel may cost up to one Gbits/sec transmission bandwidth. We suggest a parallel processing structure using the proposed international standard for visual telephony (CCITT Px64 kbs standard) as processing elements, to compress the digital HDTV pictures. The basic idea is to partition an HDTV picture into smaller sub-pictures and then compress each sub-picture using a CCITT Px64kbs coder, which is cost-effective, by today’s technology, only on small size pictures. Since each sub-picture is processed by an independent coder, without coordination these coded sub-pictures may have unequal picture quality. To maintain a uniform quality HDTV picture, the following two issues are studied: (l) sub-channel control strategy (bits allocated to each sub-picture), and (2) quantization and buffer control strategy for individual sub-picture coder. Algorithms to resolve the above problems and their computer simulations are presented.
Hsueh-Ming Hang, Riccardo Leonardi, Barry G. Haskell, Robert L. Schmidt, Hemant Bheda, Joseph H. Othmer
VCIP1
1988 Interpolative vector quantization of color images
abstract
Interpolative vector quantization has been devised to alleviate the visible block structure of coded images plus the sensitive codebook problems produced by a simple vector quantizer. In addition, the problem of selecting color components for color picture vector quantization is discussed. Computer simulations demonstrate the success of this coding technique for color image compression at approximately 0.3 b/pel. Some background information on vector quantization is provided.>
Hsueh-Ming Hang, Barry G. Haskell
IEEE Trans. Commun.1
1985 Predictive Vector Quantization of Images
abstract
The purpose of this paper is to present new image coding schemes based on a predictive vector quantization (PVQ) approach. The predictive part of the encoder is used to partially remove redundancy, and the VQ part further removes the residual redundancy and selects good quantization levels for the global waveform. Two implementations of this coding approach have been devised, namely, sliding block PVQ and block tree PVQ. Simulations on real images show significant improvement over the conventional DPCM and tree codes using these new techniques. The strong robustness property of these coding schemes is also experimentally demonstrated.
Hsueh-Ming Hang, John W. Woods
IEEE Trans. Commun.1
1984 Near merging of paths in suboptimal tree searching
abstract
The near merging of paths in tree searching is explored. This near merging can degrade coding performance when a suboptimal search algorithm is used. The problem is first identified and then a feasible solution is presented. Some examples from image source coding using the(M, L)algorithm are given.
Hsueh-Ming Hang, John W. Woods
IEEE Trans. Inf. Theory1