VLDB 2026 Research / reviewers in the wild / expert
Oscar C. Au
dblp:12/3939
· DBLP profile ↗
330ranked-venue papers
7as first author
0since 2021 · last 2020
0000-0002-3235-6517ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 277 · 7 first-authorSystems, architecture and hardware · 43Artificial intelligence and machine learning · 9 · 1 first-authorSecurity and privacy · 5Databases, data management, data science and information retrieval · 3Computer networks · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
23 papers |
Image and video processing · 50% Image and video coding · 42% Rendering · 3% | |
| Network and information security
2 papers |
Digital forensics and information hiding · 100% | |
| Theoretical computer science
4 papers |
Coding theory · 45% Mathematical optimization · 23% Algorithms and data structures · 22% | |
| Artificial intelligence
3 papers |
3D vision · 100% | |
| Computer networks
1 paper |
Physical-layer communications · 100% |
Topics — the 30 heaviest of 56, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › image restoration
demosaicing |
0.7 | 3 | 2017 | Maximum a Posterior and Perceptually Motivated Reconstruction Algorithm: A Generic Framework · IEEE Trans. Multim. 2017 Adaptive Multispectral Demosaicking Based on Frequency-Domain Analysis of Spectral Correlation · IEEE Trans. Image Process. 2017 Joint Demosaicing and Subpixel-Based Down-Sampling for Bayer Images: A Fast Frequency-Domain Analysis Approach · IEEE Trans. Multim. 2012 |
Image and video coding
rate-distortion optimization |
0.5 | 2 | 2019 | A New Rate-Complexity-Distortion Model for Fast Motion Estimation Algorithm in HEVC · IEEE Trans. Multim. 2019 Merge Frame Design for Video Stream Switching Using Piecewise Constant Functions · IEEE Trans. Image Process. 2016 |
Image and video processing › image resampling › image rescaling
image downscaling |
0.5 | 3 | 2013 | Luma-Chroma Space Filter Design for Subpixel-Based Monochrome Image Downsampling · IEEE Trans. Image Process. 2013 Joint Demosaicing and Subpixel-Based Down-Sampling for Bayer Images: A Fast Frequency-Domain Analysis Approach · IEEE Trans. Multim. 2012 Antialiasing Filter Design for Subpixel Downsampling via Frequency-Domain Analysis · IEEE Trans. Image Process. 2012 |
Image and video coding › image compression
encrypted image compression |
0.4 | 2 | 2014 | Designing an Efficient Image Encryption-Then-Compression System via Prediction Error Clustering and Random Permutation · IEEE Trans. Inf. Forensics Secur. 2014 Scalable Compression of Stream Cipher Encrypted Images Through Context-Adaptive Sampling · IEEE Trans. Inf. Forensics Secur. 2014 |
Image and video processing
motion estimation |
0.4 | 1 | 2019 | A New Rate-Complexity-Distortion Model for Fast Motion Estimation Algorithm in HEVC · IEEE Trans. Multim. 2019 |
Image and video coding
video compression |
0.4 | 1 | 2019 | A New Rate-Complexity-Distortion Model for Fast Motion Estimation Algorithm in HEVC · IEEE Trans. Multim. 2019 |
Digital forensics and information hiding › watermarking › image watermarking
halftone image watermarking |
0.4 | 2 | 2018 | Halftone Image Watermarking by Content Aware Double-Sided Embedding Error Diffusion · IEEE Trans. Image Process. 2018 Data hiding watermarking for halftone images · IEEE Trans. Image Process. 2002 |
Digital forensics and information hiding
watermarking |
0.4 | 2 | 2018 | Halftone Image Watermarking by Content Aware Double-Sided Embedding Error Diffusion · IEEE Trans. Image Process. 2018 Data hiding watermarking for halftone images · IEEE Trans. Image Process. 2002 |
Image and video processing
image reconstruction |
0.3 | 2 | 2017 | Maximum a Posterior and Perceptually Motivated Reconstruction Algorithm: A Generic Framework · IEEE Trans. Multim. 2017 Scalable Compression of Stream Cipher Encrypted Images Through Context-Adaptive Sampling · IEEE Trans. Inf. Forensics Secur. 2014 |
Image and video processing
image enhancement |
0.3 | 2 | 2016 | Image Bit-Depth Enhancement via Maximum A Posteriori Estimation of AC Signal · IEEE Trans. Image Process. 2016 Texture optimization for seamless view synthesis through energy minimization · ACM Multimedia 2012 |
Image and video processing › color image processing
color demosaicking |
0.3 | 1 | 2017 | Adaptive Multispectral Demosaicking Based on Frequency-Domain Analysis of Spectral Correlation · IEEE Trans. Image Process. 2017 |
Image and video processing › image restoration
image denoising |
0.3 | 1 | 2017 | Maximum a Posterior and Perceptually Motivated Reconstruction Algorithm: A Generic Framework · IEEE Trans. Multim. 2017 |
Image and video processing › video frame interpolation › interpolation
image interpolation |
0.3 | 1 | 2017 | Maximum a Posterior and Perceptually Motivated Reconstruction Algorithm: A Generic Framework · IEEE Trans. Multim. 2017 |
Computational photography and imaging › spectral imaging › multispectral imaging
multispectral demosaicing |
0.3 | 1 | 2017 | Adaptive Multispectral Demosaicking Based on Frequency-Domain Analysis of Spectral Correlation · IEEE Trans. Image Process. 2017 |
Image and video processing
image segmentation |
0.3 | 2 | 2012 | Bag of textons for image segmentation via soft clustering and convex shift · CVPR 2012 Nonparametric density estimation on a graph: Learning framework, fast approximation and application in image segmentation · CVPR 2011 |
Image and video coding
image compression |
0.3 | 2 | 2015 | Scalable Compression of Stream Cipher Encrypted Images Through Context-Adaptive Sampling · IEEE Trans. Inf. Forensics Secur. 2014 Multiresolution Graph Fourier Transform for Compression of Piecewise Smooth Images · IEEE Trans. Image Process. 2015 |
Image and video processing › image enhancement
bit-depth enhancement |
0.2 | 1 | 2016 | Image Bit-Depth Enhancement via Maximum A Posteriori Estimation of AC Signal · IEEE Trans. Image Process. 2016 |
Image and video coding
distributed source coding |
0.2 | 1 | 2016 | Merge Frame Design for Video Stream Switching Using Piecewise Constant Functions · IEEE Trans. Image Process. 2016 |
Image and video coding
transform coding |
0.2 | 1 | 2015 | Multiresolution Graph Fourier Transform for Compression of Piecewise Smooth Images · IEEE Trans. Image Process. 2015 |
Computer vision › 3D vision
3d shape reconstruction |
0.2 | 1 | 2014 | Rate-Constrained 3D Surface Estimation From Noise-Corrupted Multiview Depth Videos · IEEE Trans. Image Process. 2014 |
Computer vision › 3D vision
depth estimation |
0.2 | 1 | 2014 | Seamless View Synthesis Through Texture Optimization · IEEE Trans. Image Process. 2014 |
Image and video coding › video compression
3d video coding |
0.2 | 1 | 2014 | An Analytical Model for Synthesis Distortion Estimation in 3D Video · IEEE Trans. Image Process. 2014 |
Image and video coding › video compression › intra prediction
chroma intra-prediction |
0.2 | 1 | 2014 | Chroma Intra Prediction Based on Inter-Channel Correlation for HEVC · IEEE Trans. Image Process. 2014 |
Image and video coding › video compression › 3d video coding
depth map coding |
0.2 | 1 | 2014 | Rate-Constrained 3D Surface Estimation From Noise-Corrupted Multiview Depth Videos · IEEE Trans. Image Process. 2014 |
Rendering
novel view synthesis |
0.2 | 1 | 2014 | Seamless View Synthesis Through Texture Optimization · IEEE Trans. Image Process. 2014 |
Physical-layer communications › signal detection
maximum likelihood detection |
0.2 | 1 | 2013 | Universal Binary Semidefinite Relaxation for ML Signal Detection · IEEE Trans. Commun. 2013 |
Physical-layer communications
signal detection |
0.2 | 1 | 2013 | Universal Binary Semidefinite Relaxation for ML Signal Detection · IEEE Trans. Commun. 2013 |
Mathematical optimization › convex relaxation
semidefinite relaxation |
0.2 | 1 | 2013 | Universal Binary Semidefinite Relaxation for ML Signal Detection · IEEE Trans. Commun. 2013 |
Computer vision › 3D vision
novel view synthesis |
0.1 | 1 | 2012 | Texture optimization for seamless view synthesis through energy minimization · ACM Multimedia 2012 |
Image and video processing › image segmentation
superpixel segmentation |
0.1 | 1 | 2012 | Bag of textons for image segmentation via soft clustering and convex shift · CVPR 2012 |
Methods — techniques the papers use, named apart from their topics
frequency-domain analysis · 0.9energy minimization · 0.7noise visibility function · 0.7error diffusion · 0.7rate-complexity-distortion modeling · 0.4quadtree coding · 0.4semidefinite relaxation · 0.3dual barrier method · 0.3decision feedback · 0.3maximum a posteriori estimation · 0.3intra-prediction · 0.3gradient magnitude similarity · 0.3anti-aliasing filter · 0.3node shifting · 0.2mean shift · 0.2rate-constrained maximum a posteriori estimation · 0.2gauss-seidel iteration · 0.2depth-image-based rendering · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | A Low-Power Motion Estimation Architecture for HEVC Based on a New Sum of Absolute Difference ComputationabstractHigh-efficiency video coding (HEVC) poses a considerable challenge to hardware implementation due to its complexity. Mobile devices are powered by batteries that are limited in capacity. Therefore, reducing the power consumption arising from the implementation of sophisticated coding tools in HEVC is an especially important issue for mobile devices. In particular, motion estimation (ME) is the major contributor to the power consumption of the encoder and the calculation of the sum of absolute difference (SAD) for ME consumes more than 50% of the total ME power. In this paper, a low-power motion estimation VLSI architecture is proposed based on a novel method of calculating the SAD. By reusing the calculation, the computation complexity and, hence, the power consumption are reduced. A low-power systolic processing elements array and a novel memory hierarchy are developed, which enable real-time processing of 8K resolution video with only half of the power consumption when compared with the state-of-the-art design. Luheng Jia, Chi-Ying Tsui, Oscar C. Au, Kebin Jia |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | A New Rate-Complexity-Distortion Model for Fast Motion Estimation Algorithm in HEVCabstractIn the high efficiency video coding (HEVC) standard, motion estimation (ME) adopts a quadtree coding structure and a larger search range to improve the coding performance. These advanced coding tools, however, dramatically increase the computational complexity. To accelerate ME, fast methods have been proposed that reduce ME complexity at the expense of rate-distortion (R-D) performance. However, none of these methods can claim their tradeoff to be optimal. In this paper, we propose an optimal fast motion estimation (FME) algorithm based on an analytical model of rate-complexity-distortion (R-C-D). We extend the traditional R-D model by introducing the ME complexity into it, which enables us to explicitly express the R-D performance under different complexity budgets. Based on the R-C-D model, the proposed FME finds the R-C-D optimized search ranges for some representative prediction units (PUs), which are then extended or refined dynamically according to motion characteristics for neighboring PUs. The proposed FME enables an R-D performance close to that of full search. When compared with the default FME method in the reference software of HEVC, the proposed fast algorithm can reduce the complexity by over 80% while improving the R-D performance. Furthermore, our proposed FME algorithm is hardware friendly, as regular data flow enables high data reuse efficiency. Luheng Jia, Chi-Ying Tsui, Oscar C. Au, Kebin Jia |
IEEE Trans. Multim. | 3 |
| 2018 | Halftone Image Watermarking by Content Aware Double-Sided Embedding Error DiffusionabstractIn this paper, we carry out a performance analysis from a probabilistic perspective to introduce the error diffusion-based halftone visual watermarking (EDHVW) methods' expected performances and limitations. Then, we propose a new general EDHVW method, content aware double-sided embedding error diffusion (CaDEED), via considering the expected watermark decoding performance with specific content of the cover images and watermark, different noise tolerance abilities of various cover image content, and the different importance levels of every pixel (when being perceived) in the secret pattern (watermark). To demonstrate the effectiveness of CaDEED, we propose CaDEED with expectation constraint (CaDEED-EC) and CaDEED-noise visibility function (NVF) and importance factor (IF) (CaDEED-N&I). Specifically, we build CaDEED-EC by only considering the expected performances of specific cover images and watermark. By adopting the NVF and proposing the IF to assign weights to every embedding location and watermark pixel, respectively, we build the specific method CaDEED-N&I. In the experiments, we select the optimal parameters for NVF and IF via extensive experiments. In both the numerical and visual comparisons, the experimental results demonstrate the superiority of our proposed work. Yuanfang Guo, Oscar C. Au, Rui Wang 0032, Lu Fang 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2017 | Adaptive Multispectral Demosaicking Based on Frequency-Domain Analysis of Spectral CorrelationabstractColor filter array (CFA) interpolation, or three-band demosaicking, is a process of interpolating the missing color samples in each band to reconstruct a full color image. In this paper, we are concerned with the challenging problem of multispectral demosaicking, where each band is significantly undersampled due to the increment in the number of bands. Specifically, we demonstrate a frequency-domain analysis of the subsampled color-difference signal and observe that the conventional assumption of highly correlated spectral bands for estimating undersampled components is not precise. Instead, such a spectral correlation assumption is image dependent and rests on the aliasing interferences among the various color-difference spectra. To address this problem, we propose an adaptive spectral-correlation-based demosaicking (ASCD) algorithm that uses a novel anti-aliasing filter to suppress these interferences, and we then integrate it with an intra-prediction scheme to generate a more accurate prediction for the reconstructed image. Our ASCD is computationally very simple, and exploits the spectral correlation property much more effectively than the existing algorithms. Experimental results conducted on two data sets for multispectral demosaicking and one data set for CFA demosaicking demonstrate that the proposed ASCD outperforms the state-of-the-art algorithms. Sunil Prasad Jaiswal, Lu Fang 0001, Vinit Jakhetiya, Jiahao Pang, Klaus Mueller 0001, Oscar C. Au |
IEEE Trans. Image Process. | 6 |
| 2017 | Maximum a Posterior and Perceptually Motivated Reconstruction Algorithm: A Generic FrameworkabstractMost of the existing image reconstruction algorithms are application specific, and have generalization issues due to the need for parameter tuning and an unknown level of signal distortion. Addressing these problems, in this paper, we propose an efficient perceptually motivated and maximum a posterior (MAP)-based generic framework for image reconstruction. This can be applied to several image/video processing applications, where there is a necessity to improve reconstruction accuracy and suppress visible artifacts, such as denoising, deinterlacing, interpolation, de-blocking of Jpeg/Jpeg-2000, and demosaicing. The gradient magnitudes are noise insensitive to a moderate levels of noise and we propose to utilize this property for finding pixels with similar edge semantics in the neighborhood when neighboring pixels are noisy. With this view, we incorporate the gradient magnitude similarity based image quality assessment metric with the MAP estimation and, in turn, it can better approximate the variance of the MAP, as compared to nonlinear filters. The proposed generic algorithm (without manually tuning any parameters) is shown to produce a better quality of reconstruction when compared to the state-of-the-art application-specific algorithms, for most of the image processing applications. Vinit Jakhetiya, Weisi Lin, Sunil Prasad Jaiswal, Sharath Chandra Guntuku, Oscar C. Au |
IEEE Trans. Multim. | 5 |
| 2016 | Optimized high-frequency based interpolation for multispectral demosaickingabstractMultispectral demosaicking, which is an extension of color demosaicking, is a challenging problem because each band is significantly undersampled and thus precise reconstruction is needed for the restoration of high-frequency components, such as edges, textures etc. In general, existing algorithms borrow high-frequency information either from different bands via inter-color correlation or from within the bands, and produces artifact in the reconstructed image. To meet this inherent shortcoming, we propose to incorporate two different high-frequency components and integrate them optimally in the linear minimum mean square sense (LMMSE) for the precise reconstruction of undersampled components. Experimental results demonstrate that the proposed algorithm based on the optimized high-frequency achieves superior performance compared to existing algorithms both in terms of objective and subjective quality. Sunil Prasad Jaiswal, Lu Fang 0001, Vinit Jakhetiya, Manohar Kuse, Oscar C. Au |
ICIP | 5 |
| 2016 | Halftone image watermarking via optimization
Yuanfang Guo, Oscar C. Au, Jiantao Zhou 0001, Ketan Tang, Xiaopeng Fan 0001 |
Signal Process. Image Commun. | 2 |
| 2016 | Adaptive Block Coding Order for Intra Prediction in HEVCabstractIn this paper, an adaptive block coding order for intra prediction is proposed. Modern video coding standards, including the most recent High Efficiency Video Coding (HEVC), utilize fixed scan orders in processing blocks during intra coding. However, the fixed scan orders typically result in residual blocks with noticeable edge patterns. That means the fixed scan orders cannot fully exploit the content-adaptive spatial correlations between adjacent blocks, thus the bitrate after compression tends to be large. To reduce the bitrate induced by inaccurate intra prediction, the proposed approach adaptively chooses both the block and subblock coding orders by minimizing the coding cost. Specifically, determining the block coding order is formulated as a traveling salesman problem that is solved using dynamic programming. Besides the block coding order, we also design the subblock coding order in each block with an adaptive manner. The experimental results demonstrate a Bjøntegaard-Delta-rate reduction of up to 4.4% compared with HEVC anchor. Amin Zheng, Yuan Yuan 0002, Jiantao Zhou 0001, Yuanfang Guo, Haitao Yang 0001, Oscar C. Au |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2016 | Secure Reversible Image Data Hiding Over Encrypted Domain via Key ModulationabstractThis paper proposes a novel reversible image data hiding scheme over encrypted domain. Data embedding is achieved through a public key modulation mechanism, in which access to the secret encryption key is not needed. At the decoder side, a powerful two-class SVM classifier is designed to distinguish encrypted and nonencrypted image patches, allowing us to jointly decode the embedded message and the original image signal. Compared with the state-of-the-art methods, the proposed approach provides higher embedding capacity and is able to perfectly reconstruct the original image as well as the embedded message. Extensive experimental results are provided to validate the superior performance of our scheme. Jiantao Zhou 0001, Weiwei Sun 0009, Li Dong 0006, Xianming Liu 0005, Oscar C. Au, Yuan Yan Tang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Merge Frame Design for Video Stream Switching Using Piecewise Constant FunctionsabstractThe ability to efficiently switch from one pre-encoded video stream to another (e.g., for bitrate adaptation or view switching) is important for many interactive streaming applications. Recently, stream-switching mechanisms based on distributed source coding (DSC) have been proposed. In order to reduce the overall transmission rate, these approaches provide a merge mechanism, where information is sent to the decoder, such that the exact same frame can be reconstructed given that any one of a known set of side information (SI) frames is available at the decoder (e.g., each SI frame may correspond to a different stream from which we are switching). However, the use of bit-plane coding and channel coding in many DSC approaches leads to complex coding and decoding. In this paper, we propose an alternative approach for merging multiple SI frames, using a piecewise constant (PWC) function as the merge operator. In our approach, for each block to be reconstructed, a series of parameters of these PWC merge functions are transmitted in order to guarantee identical reconstruction given the known SI blocks. We consider two different scenarios. In the first case, a target frame is first given, and then merge parameters are chosen, so that this frame can be reconstructed exactly at the decoder. In contrast, in the second scenario, the reconstructed frame and the merge parameters are jointly optimized to meet a rate-distortion criteria. Experiments show that for both scenarios, our proposed merge techniques can outperform both a recent approach based on DSC and the SP-frame approach in H.264, in terms of compression efficiency and decoder complexity. Wei Dai 0002, Gene Cheung, Ngai-Man Cheung, Antonio Ortega, Oscar C. Au |
IEEE Trans. Image Process. | 5 |
| 2016 | Image Bit-Depth Enhancement via Maximum A Posteriori Estimation of AC SignalabstractWhen images at low bit-depth are rendered at high bit-depth displays, missing least significant bits needs to be estimated. We study the image bit-depth enhancement problem: estimating an original image from its quantized version from a minimum mean squared error (MMSE) perspective. We first argue that a graph-signal smoothness prior-one defined on a graph embedding the image structure-is an appropriate prior for the bit-depth enhancement problem. We next show that directly solving for the MMSE solution is, in general, too computationally expensive to be practical. We then propose an efficient approximation strategy. In particular, we first estimate the ac component of the desired signal in a maximum a posteriori formulation, efficiently computed via convex programming. We then compute the dc component with an MMSE criterion in a closed form given the computed ac component. Experiments show that our proposed two-step approach has improved performance over the conventional bit-depth enhancement schemes in both objective and subjective comparisons. Pengfei Wan 0001, Gene Cheung, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Trans. Image Process. | 5 |
| 2015 | Optimal graph laplacian regularization for natural image denoisingabstractImage denoising is an under-determined problem, and hence it is important to define appropriate image priors for regularization. One recent popular prior is the graph Laplacian regularizer, where a given pixel patch is assumed to be smooth in the graph-signal domain. The strength and direction of the resulting graph-based filter are computed from the graph's edge weights. In this paper, we derive the optimal edge weights for local graph-based filtering using gradient estimates from non-local pixel patches that are self-similar. To analyze the effects of the gradient estimates on the graph Laplacian regularizer, we first show theoretically that, given graph-signal hDis a set of discrete samples on continuous function h(x; y) in a closed region Ω, graph Laplacian regularizer (hD)TLhDconverges to a continuous functional SΩintegrating gradient norm of h in metric space G-i.e., (∇h)TG-1(∇h)-over Ω. We then derive the optimal metric space G*: one that leads to a graph Laplacian regularizer that is discriminant when the gradient estimates are accurate, and robust when the gradient estimates are noisy. Finally, having derived G* we compute the corresponding edge weights to define the Laplacian L used for filtering. Experimental results show that our image denoising algorithm using the per-patch optimal metric space G* outperforms non-local means (NLM) by up to 1.5 dB in PSNR. Jiahao Pang, Gene Cheung, Antonio Ortega, Oscar C. Au |
ICASSP | 4 |
| 2015 | Image compression via dense descriptors assisted synthesisabstractIn this paper, we propose a novel image compression approach towards visual quality rather than pixel fidelity. We intentionally remove several blocks at the encoder and reconstruct them at the decoder to get bits reduction. The removal blocks are wisely and adaptively selected based on blocks clustering, patch similarity and removal priority. A well-suited similarity measurement is defined to capture the common pattern between patches as well as tell their substitutability based on boundary consistency. To assist the removal blocks reconstruction at the decoder, we extract some dense descriptors as the side information to the decoder. Encouraging experimental results show that our compression scheme achieves up to 20.26%bits reduction with a comparable visual quality compared to the most recent standard High Efficiency Video Coding (HEVC). Yuan Yuan 0002, Amin Zheng, Haitao Yang 0001, Oscar C. Au |
ICASSP | 4 |
| 2015 | Image colorization via color propagation and rank minimizationabstractImage colorization aims to add colors to grayscale images, which used to be a time-consuming and tedious task that requires lots of human efforts. In this paper, we present a novel colorization method based on color propagation and rank minimization. Given a small portion of chrominance values and a grayscale image, we firstly propagate the known color values to other pixels to be colorized. As the colorized image after color propagation is not accurate, we then define a confidence matrix to measure the propagation fidelity. Finally, pixels that have propagated chrominance values with confidence are colorized by rank minimization, which exploits the redundancy of natural images. Experimental results on real data set show that our proposed method achieves state-of-the-art colorization quality. Yonggen Ling, Oscar C. Au, Jiahao Pang, Jin Zeng 0004, Yuan Yuan 0002, Amin Zheng |
ICIP | 2 |
| 2015 | Motion estimation via hierarchical block matching and graph cutabstractBlock matching based motion estimation algorithms are adopted in numerous practical video processing applications due to their low complexity. However, conventional block matching based methods process each block independently to minimize the energy function, which results in a local minimum. It fails to preserve the motion details. In this paper, we formulate the motion estimation as a labeling problem. The candidate labels are initialized by adopting a hierarchical block matching method. Then, we employ a graph cut algorithm to efficiently solve the global labeling problem with candidate labels. Experimental results show that the proposed approach can well preserve the motion details and outperforms all other block based motion estimation methods in terms of endpoint error and angle error on the Middleburry optical flow benchmark. Amin Zheng, Yuan Yuan 0002, Sunil Prasad Jaiswal, Oscar C. Au |
ICIP | 4 |
| 2015 | Motion vector fields based video codingabstractMotion vector fields (MVFs) are able to produce a more accurate prediction image than conventional block based motion compensation. However, MVFs are not used in conventional video coding standards due to the difficulty of efficient estimation and compression. In this work, we propose an MVF based video coding framework. We formulate the estimation of the MVF as a discrete optimization problem by both optimizing the residual energy and MVF smoothness, which can be efficiently solved by a graph cut algorithm with initialized motion vectors for each pixel. We then propose a modified rate distortion optimization approach for the MVF compression. Experimental results show that the proposed method has comparable performance in terms of object quality compared to the state-of-art of HEVC, while it has a better subjective performance by overcoming the block artifacts problem. Amin Zheng, Yuan Yuan 0002, Hong Zhang 0024, Haitao Yang 0001, Pengfei Wan 0001, Oscar C. Au |
ICIP | 6 |
| 2015 | A fast variable block size motion estimation algorithm with refined search range for a two-layer data reuse schemeabstractMotion estimation (ME) serves as a key tool in a variety of video coding standards. With the increasing need for higher resolution video format, the limited memory bandwidth becomes a bottleneck for ME implementation. The huge data loading from external memory to the on-chip memory and the frequent data fetching from the on-chip memory to the ME engine are two major problems. To reduce both off-chip and on-chip memory bandwidth, we propose a two-layer data reuse scheme. On the macroblock (MB) layer, an advanced Level C data reuse scheme is presented. It employs two cooperating on-chip caches which load data in a novel local-snake scanning manner. On the block layer, we propose a fast variable block size motion estimation with a refined search window (RSW-VBSME). A new approach for hardware implementation of VBSME is then employed based on the fast algorithm. Instead of obtain the SADs of all the modes at the same time, the ME of different block sizes are performed separately. This enables higher data reusability within an MB. The two-layer data reuse scheme archives a more than 90% reduction of off-chip memory bandwidth with a slight increase of on-chip memory size. Moreover, the on-chip memory bandwidth is also greatly reduced compared with other reuse methods with different VBSME implementations. Luheng Jia, Chi-Ying Tsui, Oscar C. Au, Amin Zheng |
ISCAS | 3 |
| 2015 | Precision Enhancement of 3-D Surfaces from Compressed Multiview Depth MapsabstractTransmitting depth maps captured from multiple viewpoints of a 3-D scene enables a wide range of receiver-side 3-D applications, including virtual view synthesis via depth-image-based rendering (DIBR). Observing that compressed depth maps from different viewpoints constitute multiple descriptions (MD) of the same signal, we propose to reconstruct 3-D surfaces of the scene by considering multiple compressed depth maps jointly. Specifically, we propose an alternating projection algorithm, inspired by the theory of projection onto convex sets (POCS), which at convergence returns a 3-D surface that satisfies three sets of conditions: spatial smoothness prior, quantization bin constraints in the block transform domain, and inter-view consistency. We present a theoretical proof that shows convergence of our algorithm under benign conditions. Compared to existing multiview depth map denoising schemes and single image de-quantization schemes, our proposed solution achieves higher objective quality for both reconstructed depth maps and synthesized virtual views. Pengfei Wan 0001, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Signal Process. Lett. | 6 |
| 2015 | Stereo Matching with Optimal Local Adaptive Radiometric CompensationabstractA common assumption in stereo matching is that the corresponding pixels in stereo images have similar pixel values. Unfortunately, such an assumption may not be true due to radiometric variations in different views, leading to severely degraded matching results. In this letter, we propose a radiometrically invariant stereo matching algorithm called Optimal Local Adaptive Radiometric Compensation (LARAC). In LARAC, we approximate the spatially varying Pixel Value Correspondence Function (PVCF) between a corresponding pixel pair as a locally consistent polynomial within an optimal local adaptive window. The optimal polynomial coefficients are obtained for each candidate disparity value and are used to compute the matching cost. Meanwhile, a self-correction property is achieved by the proposed LARAC, leading to reduced matching errors for the outlier pixels. Experimental results suggest that the proposed LARAC outperforms other state-of-the-art stereo matching algorithms. Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Lu Fang 0001, Feng Zou 0006 |
IEEE Signal Process. Lett. | 2 |
| 2015 | Multiresolution Graph Fourier Transform for Compression of Piecewise Smooth ImagesabstractPiecewise smooth (PWS) images (e.g., depth maps or animation images) contain unique signal characteristics such as sharp object boundaries and slowly varying interior surfaces. Leveraging on recent advances in graph signal processing, in this paper, we propose to compress the PWS images using suitable graph Fourier transforms (GFTs) to minimize the total signal representation cost of each pixel block, considering both the sparsity of the signal's transform coefficients and the compactness of transform description. Unlike fixed transforms, such as the discrete cosine transform, we can adapt GFT to a particular class of pixel blocks. In particular, we select one among a defined search space of GFTs to minimize total representation cost via our proposed algorithms, leveraging on graph optimization techniques, such as spectral clustering and minimum graph cuts. Furthermore, for practical implementation of GFT, we introduce two techniques to reduce computation complexity. First, at the encoder, we low-pass filter and downsample a high-resolution (HR) pixel block to obtain a low-resolution (LR) one, so that a LR-GFT can be employed. At the decoder, upsampling and interpolation are performed adaptively along HR boundaries coded using arithmetic edge coding, so that sharp object boundaries can be well preserved. Second, instead of computing GFT from a graph in real-time via eigen-decomposition, the most popular LR-GFTs are pre-computed and stored in a table for lookup during encoding and decoding. Using depth maps and computer-graphics images as examples of the PWS images, experimental results show that our proposed multiresolution-GFT scheme outperforms H.264 intra by 6.8 dB on average in peak signal-to-noise ratio at the same bit rate. Wei Hu 0003, Gene Cheung, Antonio Ortega, Oscar C. Au |
IEEE Trans. Image Process. | 4 |
| 2014 | SSIM-based rate-distortion optimization in H.264abstractIn the current video coding standards, rate-distortion optimization (RDO) plays an important role in achieving best tradeoff between the perceived distortion and transmission rate. It is widely used in all kinds of encoder decisions, including block mode decision, motion vector selection and so on. Generally, the sum of absolute difference (SAD) or the sum of square difference (SSD) is used as the distortion measurement. However, it is well known that both of them cannot always reflect the perceptual quality of the encoded video. In this paper, an objective quality measurement structural similarity (SSIM) index is proposed as the distortion measurement in the RDO framework for video coding standards. By fully exploiting the relationship between SSIM and mean square error (MSE), the SSIM-based RDO framework can be approximated by the original SSD-based RDO framework with only a scaling of the Lagrange multiplier. Experimental results show that the proposed method outperforms the latest H.264 codec and also the state-of-the-art SSIM-based RDO video codec. Wei Dai 0002, Oscar C. Au, Pengfei Wan 0001, Wei Hu 0003, Jiantao Zhou 0001 |
ICASSP | 2 |
| 2014 | Fast and efficient intra-frame deinterlacing using observation model based bilateral filterabstractRecently, a few bilateral filter based interpolation and intraframe deinterlacing algorithms have been proposed, but these algorithms only use prior information (bilateral filter). In this paper, we propose an efficient and fast intra-frame deinterlacing algorithm using an observation model based bilateral filter (using both likelihood and prior information). Our proposed algorithm is also able to use approximated horizontal pixels for the deinterlacing, which results into the better prediction of the edges. From extensive experiments, it is observed that the proposed algorithm has the capability of provide satisfactory results in terms of both objective and subjective quality. Vinit Jakhetiya, Oscar C. Au, Sunil Prasad Jaiswal, Luheng Jia, Hong Zhang 0024 |
ICASSP | 2 |
| 2014 | Image compression via sparse reconstructionabstractThe traditional compression system only considers the statistical redundancy of images. Recent compression works exploit the visual redundancy of images to further improve the coding efficiency. However, the existing works only provide suboptimal visual redundancy removal schemes. In this paper, we propose an efficient image compression scheme based on the selection and reconstruction of the visual redundancy. The visual redundancy in an image is defined by some images blocks, named redundant blocks, which can be well reconstructed by the others in the image. At the encoder, we design an effective optimization strategy to elaborately select redundant blocks and intentionally remove them. At the decoder, we propose an image restoration method to reconstruct the removed redundant blocks with minimum reconstructed error. Encouraging experimental results show that our compression scheme achieves up to 13.67% bit rate reduction with a comparable visual quality compared to traditional High Efficiency Video Coding (HEVC). Yuan Yuan 0002, Oscar C. Au, Amin Zheng, Haitao Yang 0001, Ketan Tang, Wenxiu Sun |
ICASSP | 2 |
| 2014 | Analysis of sampling pattern and Luma-Chroma filter design for subpixel-based image downsamplingabstractSubpixel-based image downsampling is attractive in that it produces higher apparent resolution of down-sampled images on LCD displays. However increased luminance resolution is achieved at the price of color fringing artifacts. In this paper, we propose an algorithm to find a pleasing balance between increased resolution and color fidelity. We separate the subpixel-based downsampling into two stages, shifting followed by downsampling with anti-aliasing filtering. In stage one, we find special characteristics of the luminance and chrominance spectra of the shifted image, based on which the optimal sampling pattern is found. In stage two, anti-aliasing filters for luminance and chrominance are designed respectively. Experimental results verify that the proposed method manages to suppress color artifacts while maintaining high luminance sharpness. Jin Zeng 0004, Oscar C. Au, Yuanfang Guo, Jiahao Pang, Ketan Tang, Yonggen Ling |
ICASSP | 2 |
| 2014 | Joint Denoising and demosaicking of noisy CFA images based on inter-color correlationabstractMost digital cameras use a single sensor coupled with a Color Filter Array (CFA) to capture images, and apply demosaicking to interpolate the full color images. In reality, the CFA image is noisy, which causes problems in the demosaicking process. This paper proposes a Joint Denoising and Demo-saicking based on inter-Color correlation (JDDC) scheme. We propose a new framework that linearly combines an extracted luminance image and a low-passed RGB images to get a full color image. Given the noise in the extracted luminance image and the low-passed RGB images are non-stationary and partially correlated, we modify the classical Non-Local Means (NLM) filter to denoise the extracted luminance image and the low-passed RGB images before the combination. Experimental results verify the effectiveness of the proposed scheme both objectively and subjectively. Ming-Ting Sun, Lu Fang 0001, Oscar C. Au |
ICASSP | 4 |
| 2014 | Palette-based compound image compression in HEVC by exploiting non-local spatial correlationabstractNon-camera captured images (also known as compound image) contain a mixture of camera-captured natural images and computer-generated graphics and texts. Nowadays, there are more and more applications calling for non-camera captured image/video compression scheme. However, current video coding standards, which are designed for natural video, treat non-camera captured video less carefully. For example, the state-of-the-art video coding standard High Efficiency Video Coding (HEVC) may blur or even remove edges in text/graphic region. A lot of schemes are proposed to preserve direction property of texts and graphics, such as palette-based intra coding. In this paper, a novel palette coding scheme is proposed for palette-based intra coding in HEVC. The palette in a block is predicted from an adaptive palette template, which records the statistical non-local spatial correlation of an image. Every block chooses its own palette using the palette template as the prediction in a rate-distortion optimized manner. Experimental results show that the proposed scheme can achieve up to 5.2% bit-rate saving compared to the state-of-the-art palette-based coding scheme in HEVC. Oscar C. Au, Wei Dai 0002, Haitao Yang 0001, Luheng Jia, Jin Zeng 0004, Pengfei Wan 0001 |
ICASSP | 2 |
| 2014 | Graph-based joint denoising and super-resolution of generalized piecewise smooth imagesabstractImages are often decoded with noise at receiver due to capturing errors and/or signal quantization during compression. Further, it is often necessary to display a decoded image at a higher resolution than the captured one, given available high-resolution (HR) display or a need to zoom-in for detailed examination. In this paper, we address the problems of image denoising and super-resolution (SR) jointly in one unified graph-based framework, focusing on a special class of signals called generalized piecewise smooth (GPWS) images. GPWS images are composed mostly of smooth regions connected by transition regions, and represent an important subclass of images, including cartoon, sub-regions of video frames with captions, graphics images in video games, etc. Like our previous work on piecewise smooth (PWS) images, GPWS images also imply simple-enough graph representations in the pixel domain, so that suitable graph-based filtering techniques can be readily applied. Specifically, leveraging on previous work on graph spectral analysis, for a given pixel block in low-resolution (LR) we first use the second eigenvector of a computed graph Laplacian matrix to identify a hard boundary, and then use the third eigenvector to identify two piecewise smooth regions and a transition region that separates them. The LR hard boundary is then super-resolved into HR via a procedure based on local self-similarity, while graph weights of the LR transition region is mapped to those of the HR transition region via polynomial fitting. Using the computed HR boundary and weights in the transition region, we construct a suitable HR graph corresponding to the LR counterpart, and perform joint denoising / SR using a graph smoothness prior. Experimental results show that our proposed algorithm outperforms two representative separable denoising / SR schemes in both subjective and objective quality. Wei Hu 0003, Gene Cheung, Xin Li 0005, Oscar C. Au |
ICIP | 4 |
| 2014 | Exploitation of inter-color correlation for color image demosaickingabstractImage demosaicking or color filter array interpolation is a process of interpolating missing color samples to reconstruct a full color image. In general, existing algorithms assume that the high frequency components such as edges, texture etc. of different color channels are similar and thus take an advantage of it to estimate the missing samples. In this paper, we efficiently analyze the relationships of intra and inter-color correlation among the channels and observe that such assumption fails in most cases. In view of this observation, we propose a scheme that exploits the correlation between different color channels much effectively than the existing algorithms. Experimental results demonstrate that the proposed algorithm outperforms the existing methods both in terms of Peak Signal to Noise Ratio (PSNR) and visual perception. Sunil Prasad Jaiswal, Oscar C. Au, Vinit Jakhetiya, Yuan Yuan 0002 |
ICIP | 2 |
| 2014 | Self-similarity-based image colorizationabstractIn this work, we tackle the problem of coloring black-and-white images, which is image colorization. Existing image colorization algorithms can be categorized into two types: scribble-based colorization algorithms and example-based colorization algorithms. Differently, we propose a hybrid scheme that combines the advantages of both categories. Given the grayscale image to be colorized and a few color scribbles (or scattered color labels) as input, the proposed method manages to colorize the grayscale image with high quality. Similar to the mechanisms in example-based colorization methods, our algorithm firstly propagates chrominance information based on the assumption that similar image patches should have similar colors. Therefore colors of some pixels can be transferred from similar patches with known colors. After that, we apply scribble-based colorization algorithm to fully colorize the grayscale image, with different confidences assigned onto the transferred color labels. Experimental results show that, the proposed method effectively utilizes the known chrominance, and provides pleasant colorizations with very few user interventions. Jiahao Pang, Oscar C. Au, Yukihiko Yamashita, Yonggen Ling, Yuanfang Guo, Jin Zeng 0004 |
ICIP | 2 |
| 2014 | DCT coefficients generation model for film grain noise and its application in super-resolutionabstractFilm grain noise (FGN) is generated by the procedure of capturing pictures using photographic film. Images with FGN are subjectively pleasing. However, FGN is difficult to compress and its pleasant features are difficult to preserve when the images are resized. So in literature, FGN is extracted first, then regenerated for the processed noise-free images. In this paper a new method is proposed to generate FGN. In contrast to some other models which generate FGN in spatial domain, our method captures the statistic feature of FGN in frequency domain. FGN is signal dependent and the signal independent scaled noise image (SNI) is obtained by scaling FGN by corresponding noise-free image. We model each discrete cosine transform (DCT) coefficient of SNI as a Gaussian random variable. The Gaussian model parameters can be estimated from the stack of all the blocks in SNI based on stationary assumption. Experimental results show that proposed model recovers FGN with similar visual and spectrum properties to the original FGN. We also apply proposed model in superresolution and the quality of resultant images are improved. Ting Sun 0001, Luhong Liang, King Hung Chiu, Pengfei Wan 0001, Oscar C. Au |
ICIP | 5 |
| 2014 | High bit-precision image acquisition and reconstruction by planned sensor distortionabstractWe present a novel framework for high bit-precision image acquisition and reconstruction. This framework is designed based on the inherent Markov property of image signals. In acquisition stage, we add planned sensor distortion (PSD) to the analog image signal before feeding it to A/D converters (or quantizers) in camera sensor. In reconstruction stage, the acquired quantized pixel values are jointly combined to get the reconstructed signal with reduced uncertainty range. Advantages of proposed PSD framework include 1) simplicity: it does not require any change to the core hardware of existing A/D converters; 2) effectiveness: experiment results demonstrate significant PSNR gain over traditional methods (up to 10 dB when quantizer bit-depth is relatively low); and 3) generality: this framework can also be applied for acquisition of other analog signals, including audio, video, etc. Pengfei Wan 0001, Oscar C. Au, Jiahao Pang, Ketan Tang |
ICIP | 2 |
| 2014 | Image bit-depth enhancement via maximum-a-posteriori estimation of graph AC componentabstractWhile modern displays offer high dynamic range (HDR) with large bit-depth for each rendered pixel, the bulk of legacy image and video contents were captured using cameras with shallower bit-depth. In this paper, we study the bit-depth enhancement problem for images, so that a high bit-depth (HBD) image can be reconstructed from an input low bit-depth (LBD) image. The key idea is to apply appropriate smoothing given the constraints that reconstructed signal must lie within the per-pixel quantization bins. Specifically, we first define smoothness via a signal-dependent graph Laplacian, so that natural image gradients can nonetheless be interpreted as low frequencies. Given defined smoothness prior and observed LBD image, we then demonstrate that computing the most probable signal via maximum a posteriori (MAP) estimation can lead to large expected distortion. However, we argue that MAP can still be used to efficiently estimate the AC component of the desired HBD signal, which along with a distortion-minimizing DC component, can result in a good approximate solution that minimizes the expected distortion. Experimental results show that our proposed method outperforms existing bit-depth enhancement methods in terms of reconstruction error. Pengfei Wan 0001, Gene Cheung, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
ICIP | 5 |
| 2014 | Natural image matting via adaptive local and nonlocal sample clusteringabstractDigital image matting is the determination of foreground color, background color, and an opacity value of each pixel for an input image. Inherently, matting is a highly ill-posed and under-constrained problem. Thus, some assumptions need to be made to resolve it. Inspired by closed-form matting and color clustering matting, in this work, we first develop an adaptive sample clustering criterion to automatically assign either local or nonlocal neighborhood to each pixel. After that, in order to enhance matting accuracy, we improve the nonlocal clustering performance by introducing a new feature selection parameter to choose preferred feature space for different images in a fully automatic way. And finally we solve the problem using a closed form solution. Experimental results show that our algorithm achieves equal or even better performance among many state-of-the-art matting techniques. Oscar C. Au, Yuan Yuan 0002, Wenxiu Sun, Yonggen Ling, Jiahao Pang |
ICIP | 2 |
| 2014 | Intra prediction with adaptive CU processing order in HEVCabstractThe High Efficiency Video Coding (HEVC) utilizes Z-scan order to process coding units (CUs). For intra prediction, this order cannot fully exploit the spatial correlation between adjacent CUs. After transform and quantization, the residue still contains lots of energy along edges which consumes many bits for compression. To effectively reduce the residue energy along edges, a novel intra prediction approach is proposed, where the CU processing order is changed adaptively. Two additional orders are introduced in this paper besides traditional Z-scan order. Up to 1.9% bit saving is achieved in our experiments on HEVC test model. We also propose two fast order selection algorithms and the observed gains are obtained with 27% and 2% encoding time increase compared to HEVC, respectively. Amin Zheng, Oscar C. Au, Yuan Yuan 0002, Haitao Yang 0001, Jiahao Pang, Yonggen Ling |
ICIP | 2 |
| 2014 | Scalable coding of stream cipher encrypted images via adaptive samplingabstractThis work proposes a novel scalable image compression method for stream cipher encrypted images. The bit stream in the base layer is produced by coding a series of non-overlapping patches of the uniformly down-sampled version of the encrypted image. An off-line learning approach can be exploited to model the reconstruction error of original image patch based on the intrinsic relationship between the local complexity and the length of the compressed bit stream. This error model leads to a greedy strategy of adaptively selecting pixels to be coded in the enhancement layer. At the decoder side, an iterative, multi-scale technique is developed to reconstruct the image from available pixel samples. Experimental results demonstrate that the proposed scheme outperforms the state-of-the-art in terms of rate-distortion (RD) performance at low and medium rate regions. Jiantao Zhou 0001, Oscar C. Au |
ICIP | 2 |
| 2014 | Fast algorithm of arbitrary factor subpixel downsampling based on frequency analysisabstractSubpixel-based downsampling has shown its advantages over pixel-based downsampling in terms of preserving more spatial details along edges and generating sharper images, at the cost of certain amount of color-fringing artifacts in the downsampled image. To balance the sharpness and color-fringing artifacts, some algorithms are proposed to design optimal anti-aliasing (AA) filters, which are either image independent, or computationally too expensive. And all of the existing AA filters are designed for fixed downsampling factor, which makes them impractical for real applications. In this paper we propose two fast algorithms to design AA filter for arbitrary factor subpixel downsampling based on frequency analysis of the input image. The proposed algorithms generate image dependent AA filter which is as good as the state-of-the-art algorithm, but much faster. Ketan Tang, Oscar C. Au, Lu Fang 0001, Jiahao Pang, Yuanfang Guo |
ICME | 2 |
| 2014 | A low latency cloud gaming system using edge preserved image homographyabstractThe emerging cloud gaming technology has been growing fast, driving up huge mobile consumer demands. The video streaming based cloud gaming scenario renders the game scenes in the cloud servers, and streams the encoded sequences to the thin clints where the game scenes are decoded and displayed to the players. However, current existing clouding gaming services have some problems, such as the latency and bandwidth limitation. The size of the video stream is usually quite large which requires heavy transmission. Worse still, the frame data rate will burst when the game scenes contain fast translation or rotation, resulting in strong latency problem. In this paper, we propose a novel video streaming based cloud gaming algorithm which reduces the burst of the frame rate significantly. There are mainly two innovations in this paper. Firstly, based on the analysis of the motion estimation strategy in the video codec, we introduce image homography technique for better motion prediction. Meanwhile, according to the rasterization rules of the game engine, we present a special designed interpolation algorithm named Edge Preserved Interpolation (EPI), for more accurate edge interpolation and further reduce the residues in the edge regions. The proposed algorithm is implemented on the x264 platform. Experimental results show that our algorithm has 18.0% BD-rate reduction compared with x264. Lingfeng Xu, Xun Guo 0002, Yan Lu 0001, Shipeng Li 0001, Oscar C. Au, Lu Fang 0001 |
ICME | 5 |
| 2014 | Symmetrical predictor structure based integrated lossy, near lossless/lossless coding of imagesabstractPrediction based algorithms reported in the literature are not able to integrate lossy and near-lossless/lossless coding and uses only causal pixels (non-symmetrical predictor structure) for prediction. A non-symmetrical predictor structure, however, is not able to efficiently adapt near the intensity varying areas, which results into poor prediction. Hence, we propose a novel two-stage algorithm for lossy, near lossless/lossless compression using a symmetrical predictor structure is proposed. In the first stage, the proposed algorithm encodes and decodes the given image using the JPEG-2000 standard algorithm (lossy coding). This JPEG-2000 decoded image in the first stage, enables us to use the symmetrical predictor (using both causal and non-causal pixels) for prediction in the second stage. A performance evaluation shows that our algorithm is significantly better in terms of compression performance as compared to some of the computationally complex methods. Vinit Jakhetiya, Oscar C. Au, Sunil Prasad Jaiswal, Luheng Jia, Gaurav Mittal |
ISCAS | 2 |
| 2014 | Photo album compression By leveraging temporal-spatial correlations and HEVCabstractThe advancing digital photography technology has resulted in a large number of photos stored in personal computers. Photo album compression algorithms aim to save storage space and efficiently manage photos. In this paper, a general forest structure model involving depth constrain for photo album compression is proposed, which further exploits the correlations between images in the photo album. We firstly represent the images as nodes in a graph and directed edges between them as predictive coding relationship. Affinity propagation is then applied to compute for a depth-constrained forest. Finally, we adopt depth-first search algorithm to generate the compression order according to forest structure and HEVC to compress the images with adaptive GOPs and reference list. Experimental results show that the proposed compression method provides much better rate-distortion performance compared to JPEG and significantly reduce the storage space. Yonggen Ling, Oscar C. Au, Ruobing Zou, Jiahao Pang, Amin Zheng |
ISCAS | 2 |
| 2014 | Rate distortion modeling and adaptive rate control scheme for high efficiency video coding (HEVC)abstractThis paper explores a novel rate control scheme for HEVC, based on Pre-analysis Sum of Absolute Transformed Differences (pSATD). The needs for improving the rate control method of HEVC have been observed since the latest method proposed in JCT-VC H213 is still using the classical quadratic model which is implemented in H.264/AVC. The same rate control scheme on different coding structures will definitely lead to a poor performance. The proposed method aims at selecting accurate quantization parameters (QP) for inter frames according to the target bit rate. After encoding video sequences exhaustively, a strong linear relationship between the pSATD and bpp (bit per pixel) is revealed. Thus, a rate control scheme for HEVC with parameters updating adaptively is proposed accordingly. Apart from that, a new rate and distortion model are established using pSATD and Quantization Step Size (QStep). By minimizing the cost function built based on the proposed rate-distortion (R-D) model, the optimization problem can be solved properly. According to the optimal solution, a bit allocation scheme with consideration of pSATD and buffer status is formed. The experimental results illustrate that the proposed method outperforms JCT-VC H213 implemented in HEVC reference software for all test sequences, up to 19% coding gain can be achieved. Lin Sun 0004, Oscar C. Au, Fiona H. Huang |
ISCAS | 2 |
| 2014 | Adaptive Predictor Structure Based Interpolation for Reversible Data Hiding
Sunil Prasad Jaiswal, Oscar C. Au, Vinit Jakhetiya, Yuanfang Guo, Anil Kumar Tiwari |
IWDW | 2 |
| 2014 | A fast intermode decision algorithm based on analysis of inter prediction residualabstractRate-distortion-optimized (RDO) intermode decision is one of the most effective tools that greatly improves the coding performance of modern coding standards, for example, H.264/AVC and HEVC. However, RDO intermode decision also leads to extremely intense computation. To reduce the complexity, a fast intermode decision algorithm is presented in this paper. Mathematical analysis of inter prediction residual is performed which explicitly shows the impact of edge information and motion characteristics to the prediction accuracy. Moreover, It is shown that with fixed quantization step that minimizing of R-D costs over different partition types is equivalent to minimizing the variance of transformed residual which can be expressed by motion vector and edge gradient components. In consequence, the complex calculation of rate and distortion is replaced by simple pre-analysis of video content. The repetition of motion estimation (ME) and entropy coding process are avoided. Inspired by the theoretical analysis, a fast inter mode decision algorithm is proposed. Experimental results show that the fast method achieves considerable complexity reduction with negligible coding performance degradation. Luheng Jia, Oscar C. Au, Chi-Ying Tsui, Wei Dai 0002, Pengfei Wan 0001 |
MMSP | 2 |
| 2014 | Solving dense stereo matching via quadratic programmingabstractWe study the problem of formulating the discrete dense stereo matching using continuous convex optimization. One of the previous work derived a relaxed convex formulation by establishing the relationship between the disparity vector and a warping matrix. However it suffers from high computational complexity. In this paper, the previous convex formulation is translated into an equivalent quadratic programming (QP). Then redundant variables and constraints are eliminated by exploiting the internal sparse property of the warping matrix. The resulting QP can be efficiently tackled using interior point solvers. Moreover, enhanced smoothness term and effective post-processing procedures are also incorporated to further improve the disparity accuracy. Experimental results show that the proposed method is much faster and better than the previous convex formulation, and provides competitive results against existing convex approaches. Oscar C. Au, Pengfei Wan 0001, Wenxiu Sun, Lingfeng Xu, Luheng Jia |
VCIP | 2 |
| 2014 | Early detection of all-zero 4×4 blocks in High Efficiency Video Coding
Hanli Wang, Weiyao Lin, Sam Kwong, Oscar C. Au, Jun Wu 0006, Zhihua Wei 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2014 | View Synthesis Prediction in the 3-D Video Coding Extensions of AVC and HEVCabstractAdvanced multiview video systems are able to generate intermediate viewpoints of a 3-D scene. To enable low-complexity free view generation, texture and its associated depth are used as input data for each viewpoint. To improve the coding efficiency of such content, view synthesis prediction (VSP) is proposed to further reduce interview redundancy in addition to traditional disparity compensated prediction. This paper describes and analyzes rate-distortion optimized VSP designs, which were adopted in the 3-D extensions of both Advanced Video Coding (AVC) and High Efficiency Video Coding (HEVC). In particular, we propose a novel backward-VSP scheme using a derived disparity vector, as well as efficient signalling methods in the context of AVC and HEVC. In addition, we put forward a novel depth-assisted motion vector prediction method to optimize the coding efficiency. A thorough analysis of coding performance is provided using different VSP schemes and configurations. Experimental results demonstrate average bit rate reductions of 2.5% and 1.2% in AVC and HEVC coding frameworks, respectively, with up to 23.1% bit rate reduction for dependent views. Feng Zou 0006, Dong Tian, Anthony Vetro, Huifang Sun, Oscar C. Au, Shinya Shimizu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2014 | Scalable Compression of Stream Cipher Encrypted Images Through Context-Adaptive SamplingabstractThis paper proposes a novel scalable compression method for stream cipher encrypted images, where stream cipher is used in the standard format. The bit stream in the base layer is produced by coding a series of nonoverlapping patches of the uniformly down-sampled version of the encrypted image. An off-line learning approach can be exploited to model the reconstruction error from pixel samples of the original image patch, based on the intrinsic relationship between the local complexity and the length of the compressed bit stream. This error model leads to a greedy strategy of adaptively selecting pixels to be coded in the enhancement layer. At the decoder side, an iterative, multiscale technique is developed to reconstruct the image from all the available pixel samples. Experimental results demonstrate that the proposed scheme outperforms the state-of-the-arts in terms of both rate-distortion performance and visual quality of the reconstructed images at low and medium rate regions. Jiantao Zhou 0001, Oscar C. Au, Guangtao Zhai, Yuan Yan Tang, Xianming Liu 0005 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2014 | Designing an Efficient Image Encryption-Then-Compression System via Prediction Error Clustering and Random PermutationabstractIn many practical scenarios, image encryption has to be conducted prior to image compression. This has led to the problem of how to design a pair of image encryption and compression algorithms such that compressing the encrypted images can still be efficiently performed. In this paper, we design a highly efficient image encryption-then-compression (ETC) system, where both lossless and lossy compression are considered. The proposed image encryption scheme operated in the prediction error domain is shown to be able to provide a reasonably high level of security. We also demonstrate that an arithmetic coding-based approach can be exploited to efficiently compress the encrypted images. More notably, the proposed compression approach applied to encrypted images is only slightly worse, in terms of compression efficiency, than the state-of-the-art lossless/lossy image coders, which take original, unencrypted images as inputs. In contrast, most of the existing ETC solutions induce significant penalty on the compression efficiency. Jiantao Zhou 0001, Xianming Liu 0005, Oscar C. Au, Yuan Yan Tang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | An Analytical Model for Synthesis Distortion Estimation in 3D VideoabstractWe propose an analytical model to estimate the synthesized view quality in 3D video. The model relates errors in the depth images to the synthesis quality, taking into account texture image characteristics, texture image quality, and the rendering process. Especially, we decompose the synthesis distortion into texture-error induced distortion and depth-error induced distortion. We analyze the depth-error induced distortion using an approach combining frequency and spatial domain techniques. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that our analytical model can accurately estimate the synthesis noise power. Thus, the model can be used to estimate the rendering quality for different system designs. Lu Fang 0001, Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun, Oscar C. Au |
IEEE Trans. Image Process. | 6 |
| 2014 | Seamless View Synthesis Through Texture OptimizationabstractIn this paper, we present a novel view synthesis method named Visto, which uses a reference input view to generate synthesized views in nearby viewpoints. We formulate the problem as a joint optimization of inter-view texture and depth map similarity, a framework that is significantly different from other traditional approaches. As such, Visto tends to implicitly inherit the image characteristics from the reference view without the explicit use of image priors or texture modeling. Visto assumes that each patch is available in both the synthesized and reference views and thus can be applied to the common area between the two views but not the out-of-region area at the border of the synthesized view. Visto uses a Gauss–Seidel-like iterative approach to minimize the energy function. Simulation results suggest that Visto can generate seamless virtual views and outperform other state-of-the-art methods. Wenxiu Sun, Oscar C. Au, Lingfeng Xu, Wei Hu 0003 |
IEEE Trans. Image Process. | 2 |
| 2014 | Rate-Constrained 3D Surface Estimation From Noise-Corrupted Multiview Depth VideosabstractTransmitting compactly represented geometry of a dynamic 3D scene from a sender can enable a multitude of imaging functionalities at a receiver, such as synthesis of virtual images at freely chosen viewpoints via depth-image-based rendering. While depth maps—projections of 3D geometry onto 2D image planes at chosen camera viewpoints-can nowadays be readily captured by inexpensive depth sensors, they are often corrupted by non-negligible acquisition noise. Given depth maps need to be denoised and compressed at the encoder for efficient network transmission to the decoder, in this paper, we consider the denoising and compression problems jointly, arguing that doing so will result in a better overall performance than the alternative of solving the two problems separately in two stages. Specifically, we formulate a rate-constrained estimation problem, where given a set of observed noise-corrupted depth maps, the most probable (maximum a posteriori (MAP)) 3D surface is sought within a search space of surfaces with representation size no larger than a prespecified rate constraint. Our rate-constrained MAP solution reduces to the conventional unconstrained MAP 3D surface reconstruction solution if the rate constraint is loose. To solve our posed rate-constrained estimation problem, we propose an iterative algorithm, where in each iteration the structure (object boundaries) and the texture (surfaces within the object boundaries) of the depth maps are optimized alternately. Using the MVC codec for compression of multiview depth video and MPEG free viewpoint video sequences as input, experimental results show that rate-constrained estimated 3D surfaces computed by our algorithm can reduce coding rate of depth maps by up to 32% compared with unconstrained estimated surfaces for the same quality of synthesized virtual views at the decoder. Wenxiu Sun, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
IEEE Trans. Image Process. | 6 |
| 2014 | Chroma Intra Prediction Based on Inter-Channel Correlation for HEVCabstractIn this paper, we investigate a new inter-channel coding mode called LM mode proposed for the next generation video coding standard called high efficiency video coding. This mode exploits inter-channel correlation using reconstructed luma to predict chroma linearly with parameters derived from neighboring reconstructed luma and chroma pixels at both encoder and decoder to avoid overhead signaling. In this paper, we analyze the LM mode and prove that the LM parameters for predicting original chroma and reconstructed chroma are statistically the same. We also analyze the error sensitivity of the LM parameters. We identify some LM mode problematic situations and propose three novel LM-like modes called LMA, LML, and LMO to address the situations. To limit the increase in complexity due to the LM-like modes, we propose some fast algorithms with the help of some new cost functions. We further identify some potentially-problematic conditions in the parameter estimation (including regression dilution problem) and introduce a novel model correction technique to detect and correct those conditions. Simulation results suggest that considerable BD-rate reduction can be achieved by the proposed LM-like modes and model correction technique. In addition, the performance gain of the two techniques appears to be essentially additive when combined. Christophe Gisquet, Edouard François, Feng Zou 0006, Oscar C. Au |
IEEE Trans. Image Process. | 5 |
| 2013 | Predicting YouTube content popularity via Facebook data: A network spread model for optimizing multimedia deliveryabstractThe recent popularity of social networking websites have resulted in a greater usage of internet bandwidth for sharing multimedia content through websites such as Facebook and YouTube. Moving large volumes of multi-media data through limited network resources remains a technical challenge to this day. The current state-of-art solution in optimizing cache server utilization depends heavily on efficient caching policies to determine content priority. This paper proposes a Fast Threshold Spread Model (FTSM) to predict the future access pattern of multi-media content based on the social information of its past viewers. The prediction results are compared and evaluated against ground truth statistics of the respective YouTube video. A complexity analysis on the proposed algorithm for large datasets along with the correlation between Facebook social sharing and YouTube global hit count are explored. Dinuka Soysa, Denis Guangyin Chen, Oscar C. Au, Amine Bermak |
CIDM | 3 |
| 2013 | Low Bit-Rate Subpixel-Based Color Image CompressionabstractWe propose a novel low bit-rate compression scheme with sub pixel-based down-sampling and reconstruction (SPDR) for full color images. In the encoder stage, a decoder-dependent multi-channel sub pixel-based down-sampling is proposed, which is more effective in retaining high frequency detail than conventional pixel-based process. The decoder first decompresses the low-resolution image and then up-converts it to the original resolution using encoder dependent sub pixel-based reconstruction scheme by jointly considering the sub pixel-based down-sampling effect and the compression degradation. Compared to existing algorithms with comparable encoder and decoder complexity, the proposed SPDR offers complete standard compliance, competitive rate-distortion performance, and superior subjective quality. Lu Fang 0001, Ngai-Man Cheung, Oscar C. Au, Houqiang Li, Ketan Tang |
DCC | 3 |
| 2013 | Image colorization using sparse representationabstractImage colorization is the task to color a grayscale image with limited color cues. In this work, we present a novel method to perform image colorization using sparse representation. Our method first trains an over-complete dictionary in YUV color space. Then taking a grayscale image and a small subset of color pixels as inputs, our method colorizes overlapping image patches via sparse representation; it is achieved by seeking sparse representations of patches that are consistent with both the grayscale image and the color pixels. After that, we aggregate the colorized patches with weights to get an intermediate result. This process iterates until the image is properly colorized. Experimental results show that our method leads to high-quality colorizations with small number of given color pixels. To demonstrate one of the applications of the proposed method, we apply it to transfer the color of one image onto another to obtain a visually pleasing image. Jiahao Pang, Oscar C. Au, Ketan Tang, Yuanfang Guo |
ICASSP | 2 |
| 2013 | Arbitrary factor image interpolation using geodesic distance weighted 2D autoregressive modelingabstractLeast square regression has been widely used in image interpolation. Some existing regression-based interpolation methods used ordinary least squares (OLS) to formulate cost functions. These methods usually have difficulties at object boundaries because OLS is sensitive to outliers. Weighted least squares (WLS) is then adopted to solve the outlier problem. Some weighting schemes have been proposed in the literature. In this paper we propose to use geodesic distance weighting in that geodesic distance can simultaneously measure both the spatial distance and color difference. Another contribution of this paper is that we propose an optimization scheme that can handle arbitrary factor interpolation. The idea is to separate the problem into two parts, an adaptive pixel correlation model and a convolution based image degradation model. Geodesic distance weighted 2D autoregressive model is used to model the pixel correlation which preserves local geometry. The convolution based image degradation model provides the flexibility to handle arbitrary interpolation factor. The entire problem is formulated as a WLS problem constrained by a linear equality. Ketan Tang, Oscar C. Au, Yuanfang Guo, Jiahao Pang |
ICASSP | 2 |
| 2013 | 3D motion in visual saliency modelingabstractVisual saliency is a probabilistic estimate of how likely a given spatial area in an image or video is to attract human visual attention relative to other areas. Bottom-up saliency models aggregate low-level image features like luminance and color contrast, flicker, 2D motion, etc. to construct a plausible saliency map. In this paper, we introduce 3D motion (object movements towards or away from the observer) into bottom-up video saliency modeling. Given availability of per-pixel depth maps, we first propose a novel algorithm to estimate 3D motion vectors (3DMVs) for arbitrarily shaped sub-blocks in texture-plus-depth videos. We then derive two feature channels from 3DMVs to be incorporated into a widely accepted bottom-up saliency model. Experiments on subjective quality of Region-of-Interest (ROI) based video coding show that our enriched saliency model with 3DMV channels is more accurate in estimating human visual attention. Pengfei Wan 0001, Yunlong Feng, Gene Cheung, Ivan V. Bajic, Oscar C. Au, Yusheng Ji |
ICASSP | 5 |
| 2013 | Ray-space based camera spacing correction via convex optimizationabstract3D technologies such like three-dimensional television and free viewpoint television have caught enormous attentions in the consumer market recently. However, because of the inaccurate camera configuration and environmental constraint, there are errors in the assumed equally-spaced camera intervals. In this paper, we propose a novel camera spacing correction algorithm to detect the spacing errors among the multiple cameras by making the corresponding points co-linear in the epipolar plane images. Experimental results show that the proposed algorithms are robust and can achieve good performance even if the corresponding pixels are not well detected. Meanwhile, our algorithm can be solved by convex optimization with an extremely low complexity. Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Wei Hu 0003 |
ICASSP | 2 |
| 2013 | On the design of an efficient encryption-then-compression systemabstractIn many practical scenarios, image encryption has to be conducted prior to image compression. This has led to the problem of how to design a pair of encryption and compression algorithms such that compressing the encrypted image can still be efficiently performed. In this work, we propose a permutation-based image encryption method conducted over the prediction error domain. We also design an arithmetic coding (AC)-based approach to efficiently compress the encrypted image. It can be shown that the proposed scheme can provide reasonably high level of security. More notably, the compression performance on the encrypted image is only slightly degraded, compared with that of compressing the original, un-encrypted one. In contrast, most of the existing approaches induce significant penalty on the compression performance. Jiantao Zhou 0001, Xianming Liu 0005, Oscar C. Au |
ICASSP | 3 |
| 2013 | A robust interpolation-free approach for sub-pixel accuracy motion estimationabstractMotion estimation (ME) is one of the key elements in video coding standard which eliminates the temporal redundancy by using a motion vector (MV) to indicate the best match between the current frame and reference frame. A coarse to fine process is taken to find the best MV. First of all, integer-pixel ME finds a coarse MV and followed by the sub-pixel ME around the best integer-pixel point. The sub-pixel ME plays an important role in improving the coding efficiency. However, the computational complexity of searching one sub-pixel point is much higher than the integer-pixel point searching because of the interpolation and Hadamard transform operation. In this paper, an accurate optimal sub-pixel position prediction algorithm is presented. With the information of the 8 neighboring integer-pixel points, the optimal sub-pixel position is predicted directly without explicitly solving model parameters. Moreover, an outlier rejection scheme is applied to improve the robustness of the proposed algorithm. Experimental results show that the proposed algorithm outperforms the state of the art interpolation-freesub-pixel ME algorithms. Wei Dai 0002, Oscar C. Au, Wei Hu 0003, Pengfei Wan 0001 |
ICIP | 2 |
| 2013 | Rate-distortion optimized merge frame using piecewise constant functionsabstractThe ability to efficiently switch from one pre-encoded video stream to another is a valuable attribute for a variety of interactive streaming applications, such as switching among streams of the same video encoded in different bit-rates for real-time bandwidth adaptation, or view-switching among videos capturing the same dynamic 3D scene but from different viewpoints. It is well known that intra-coded I-frames can be used at switch boundaries to facilitate stream-switching. However, the size of an I-frame is large, making frequent insertion impractical. A recent proposal towards a more efficient stream-switching mechanism is distributed source coding (D-SC), which exploits worst-case correlation between a set of potential predictor frames in the decoder buffer (called side information (SI) frames) and a target frame to lower encoding rate. However, the conventional use of bit-plane and channel coding means the encoding and decoding complexity of DSC frames is large. In this paper, we pursue a novel approach to the stream-switching problem based on the concept of “signal merging”, using piecewise constant (p-wc) function as the merge operator. Specifically, we propose a new merge mode for a code block, where for each k-th transform coefficient in the block, we encode appropriate step size and horizontal shift parameters at the encoder, so that the resulting floor function at the decoder can map corresponding coefficients from any SI frame to the same reconstructed value, resulting in an identically merged signal. The selection of shift parameter per coefficient, as well as coding modes between intra and merge per block, are optimized in a rate-distortion (RD) optimal manner. Experiments show encouraging coding gain over a previous implementation of DSC frame at low-to mid-bitrates at reduced computation complexity. Wei Dai 0002, Gene Cheung, Ngai-Man Cheung, Antonio Ortega, Oscar C. Au |
ICIP | 5 |
| 2013 | Efficient adaptive prediction based reversible image watermarkingabstractIn this paper, we propose a new reversible watermarking algorithm based on additive prediction-error expansion which can recover original image after extracting the hidden data. Embedding capacity of such algorithms depend on the prediction accuracy of the predictor. We observed that the performance of a predictor based on full context prediction is preciser as compared to that of partial context prediction. In view of this observation, we propose an efficient adaptive prediction (EAP) method based on full context, that exploits local characteristics of neighboring pixels much effectively than other prediction methods reported in literature. Experimental results demonstrate that the proposed algorithm has a better embedding capacity and also gives better Peak Signal to Noise Ratio (PSNR) as compared to state-of-the-art reversible watermarking schemes. Sunil Prasad Jaiswal, Oscar C. Au, Vinit Jakhetiya, Yuanfang Guo, Anil Kumar Tiwari, Yue Kong |
ICIP | 2 |
| 2013 | Novel distortion metric for depth coding of 3D videoabstractIn state-of-the-art HEVC-based 3D video codec, multiview video plus associated depth maps are used. In order to achieve better coding performance, instead of the conventional sum of squared errors (SSE), view synthesis optimization (VSO) is proposed and included in the anchor encoder software to calculate view synthesis distortion in rate-distortion optimization (RDO) of depth coding. The anchor VSO achieves high rate-distortion (RD) performance. However, it requires partial rendering and is quite complex and time-consuming. On the other hand, simple SSE metric is fast but RD performance is low. In this paper, we propose a new distortion metric to be used in RDO for depth coding. The complexity of the proposed method is slightly higher than SSE, while its RD performance remains competitive. With a good trade-off between complexity and performance, the proposed method can replace the conventional SSE metric in RDO for depth coding, and can be used as a low-complexity alternative to the anchor VSO. Ngai-Man Cheung, Oscar C. Au, Dong Tian |
ICIP | 3 |
| 2013 | Optimal dependent bit allocation for AVS intra-frame coding via successive convex approximationabstractWe consider the optimal dependent bit allocation strategy for AVS intra-frame coding. Due to the block-based predictive coding, the rate-distortion (R-D) characteristics of neighboring blocks are dependent with each other. However, the interblock coding dependency is neglected in most of the existing bit allocation methods. Different from the conventional methods, the proposed method fully exploit the interblock coding dependency and carefully leverage it in the problem formulation. Then successive convex optimization techniques are employed to convert the original nonconvex optimization problem into a series of convex optimization problems which can be solved efficiently and optimally. Experimental results have proved the superiority of the proposed method in terms of significant R-D performance improvement. Oscar C. Au, Feng Zou 0006, Wei Hu 0003, Pengfei Wan 0001 |
ICIP | 2 |
| 2013 | Arbitrary factor image interpolation by convolution kernel constrained 2-D autoregressive modelingabstractAmong existing interpolation methods, convolution-based methods are able to perform arbitrary factor interpolation but the results are usually blurry or jaggy, adaptive interpolation methods usually can reduce the blurry and jaggy artifacts but cannot handle arbitrary factor interpolation. In this paper we propose an arbitrary factor adaptive interpolation algorithm by combining 2-D piecewise autoregressive (PAR) modeling and convolution kernel constraint. PAR model ensures local geometries are well preserved thus the resultant image is not blurry or jaggy. Convolution kernel constraint ensures the recovered high resolution image consistent with the low resolution image, and also provides the flexibility to handle arbitrary interpolation factor. Experiment results show that our algorithm achieves state-of-the-art performance for any interpolation factor. Ketan Tang, Oscar C. Au, Yuanfang Guo, Jiahao Pang, Lu Fang 0001 |
ICIP | 2 |
| 2013 | 2-SiMDoM: A 2-Sieve model for detection of mitosis in multispectral breast cancer imageryabstractIn this paper, we propose a 2-Sieve model for the detection of mitosis in breast cancer multispectral images. Multiresolution wavelet features & Gray Level Entropy Matrix (GLEM) features have been computed for each candidate on all the spectral bands. A novel dimensionality selection algorithm has been introduced and its performance compared with other existing algorithms. Data imbalance and data cleaning have been taken care of using classical data mining techniques. Furthermore, a Second Sieve classification is performed to increase the Positive Predictive Value (PPV) with minimal loss in Sensitivity. A final Sensitivity and PPV of 82.35% & 73.04% respectively was achieved over the testing set using the proposed scheme. Ardhendu Shekhar Tripathi, Atin Mathur, Mohit Daga, Manohar Kuse, Oscar C. Au |
ICIP | 5 |
| 2013 | BDCT compressed image deblocking usingweighted adaptive total variationabstractImages encoded at low bit rate usually exhibit visually annoying coding artifacts, which are commonly referred as blocking artifacts. In this paper, a novel weighted adaptive total variation method is proposed to remove the blocking artifacts. Based on the analysis of image coding process and block-based discrete cosine transform (BDCT) image properties, image deblocking is formulated as an optimization problem which is solved through approximating the objective function by a set of convex functions. Experimental results show that the proposed method can achieve better objective and subjective quality performance compared to other deblocking algorithms. Wei Dai 0002, Oscar C. Au, Feng Zou 0006 |
ICME | 2 |
| 2013 | Color clustering mattingabstractNatural image matting refers to the problem of extracting regions of interest such as foreground object from an image based on user inputs like scribbles or trimap. More specifically, we need to estimate the color information of background, foreground and the corresponding opacity, which is an ill-posed problem inherently. Inspired by closed-form matting and KNN matting, in this paper, we extend the local color line model which is based on the assumption of linear color clustering within a small local window, to nonlocal feature space neighborhood. New affinity matrix is defined to achieve better clustering. Further, we demonstrate that good clustering ensures better prediction of alpha matte. Experimental evaluations on benchmark datasets and comparisons show that our matting algorithm is of higher accuracy and better visual quality than some state-of-the-art matting algorithms. Yongfang Shi, Oscar C. Au, Jiahao Pang, Ketan Tang, Wenxiu Sun, Hong Zhang 0024, Luheng Jia |
ICME | 2 |
| 2013 | Rate-distortion optimized 3D reconstruction from noise-corrupted multiview depth videosabstractTransmitting compactly represented geometry of a dynamic scene from a sender can enable a multitude of 3D imaging functionalities at a receiver, such as synthesis of virtual images from freely chosen viewpoints via depth-image-based rendering (DIBR). While depth maps can now be readily captured using inexpensive depth sensors, they are often corrupted by non-negligible acquisition noise. In this paper, we derive 3D surfaces of a dynamic scene from noise-corrupted depth maps in a rate-distortion (RD) optimal manner. Specifically, unlike previous work that finds the most likely (e.g., maximum likelihood) 3D surface from noisy observations regardless of representation size, we judiciously search for the best fitting (i.e., minimum distortion) 3D surface subject to a bitrate constraint. Our RD-optimal solution reduces to the maximum likelihood solution as the rate constraint is loosened. Using the MVC codec for compression of multiview depth video and MPEG free viewpoint test sequences as input, experimental results show that RD-optimized 3D reconstructions computed by our algorithm outperform unprocessed depth maps by up to 2:42dB in PSNR of synthesized virtual views at the decoder for the same bitrate. Wenxiu Sun, Gene Cheung, Philip A. Chou, Dinei A. F. Florêncio, Cha Zhang, Oscar C. Au |
ICME | 6 |
| 2013 | Data hiding in error diffused color halftone imagesabstractHalftone image watermarking has been explored and developed rapidly over the past decade. However, there are still issues to be studied. This paper presents a data hiding method called Data Hiding by Dual Color Conjugate Error Diffusion (DHDCCED) to hide a binary secret pattern into two error diffused color halftone images, such that when the two color halftone images are overlaid, the secret pattern will be revealed. The experimental results show that DHDCCED can significantly improve the performances when comparing both the correct decoding rate and the visual quality of the revealed secret pattern to the existing method Color Conjugate Error Diffusion (CCED). Yuanfang Guo, Oscar C. Au, Ketan Tang, Jiahao Pang, Wenxiu Sun, Lingfeng Xu |
ISCAS | 2 |
| 2013 | A parallel deblocking filter based on H.264/AVC video coding standardabstractThe deblocking filter in H.264/AVC is one of the most time consuming part of video decoder as its high content adaptation and data dependency lead to lots of computation. In this paper, we propose a novel parallel deblocking filter design based on the H.264/AVC video coding standard, taking the advantage that the data dependency of the deblocking filter are “periodic” in one dimension. Our proposed architecture successfully reduces the dependency between horizontal and vertical filters and utilizes the “periodic” property to achieve pixel-level parallelism. Algorithm analysis and experiment results on JM and GPU show that the proposed deblocking filter keeps as good a coding efficiency as that in H.264/AVC, and its high parallelism is suitable and promising in multi-core/multi-thread computing. Oscar C. Au, Lu Fang 0001, Lin Sun 0004, Wenxiu Sun, Dinuka Soysa |
ISCAS | 2 |
| 2013 | Rate-distortion optimized block classification and bit allocation in screen video compressionabstractDue to the divergent characteristics of image contents and text contents in screen videos, how to make the joint optimization leveraging rate-distortion (R-D) optimized block classification and bit allocation is critical to the compression performance. In this paper, a general model-based solution is proposed as an attempt to solve this problem. The contributions of this paper are twofold: First, the rate and distortion characteristics of image blocks and text blocks in block-based content-adaptive screen video encoder (BASC) are carefully studied, and the rate and distortion models are proposed. Second, with the proposed rate and distortion models, the R-D optimized block classification and bit allocation are derived using bisection searched Lagrange multiplier method. Experimental results demonstrate that the proposed R-D optimized block classification and bit allocation algorithms are able to adapt to diverse screen contents, which results in a significant gain of up to 4.5dB in PSNR. Oscar C. Au, Jingjing Fu, Yan Lu 0001, Shipeng Li 0001 |
ISCAS | 2 |
| 2013 | Content based fast prediction unit quadtree depth decision algorithm for HEVCabstractThe nested quadtree based partitioning scheme of HEVC contributes a lot to the coding efficiency improvement, however, it adds significant complexity to the encoder. This paper introduces a fast prediction unit (PU) level quadtree depth decision (FPDD) algorithm. It is achieved by making use of the inherited correlation of PU quadtree structure between current largest coding unit (LCU) and its spatial and temporal neighbors. To reduce error propagation, we also propose a confidence grading scheme to prevent LCUs with bad prediction from being referred to by others. Results show that our proposed algorithm provides averagely 20.0% (up to 39.3%) encoding time reduction whilst causing negligible RD performance loss (0.2% BD-Rate increase on average) compared with HM 7.0. Yongfang Shi, Oscar C. Au, Hong Zhang 0024, Luheng Jia |
ISCAS | 2 |
| 2013 | Stereo matching by adaptive weighting selection based cost aggregationabstractCost aggregation is the most essential step for dense stereo correspondence searching, which measures the similarity between pixels in the stereo images. In this paper, based on the analysis of the optimal adaptive weight, we propose a novel support aggregation strategy by adaptive weighting selection. The proposed method calculates the aggregation cost by the joint optimization of both left and right matching cost. By assigning more reasonable weighting coefficients, we exclude the occlusion pixels while preserving sufficient support region for accurate matching. The proposed optimal strategy can be integrated by any other adaptive weighting based cost aggregation method to generate more reasonable similarity measurement. Experimental results show that, compare with traditional methods, our algorithm can reduce the foreground fatten phenomenon while increasing the accuracy in the high texture regions. Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Lu Fang 0001, Ketan Tang, Yuanfang Guo |
ISCAS | 2 |
| 2013 | HEVC-based adaptive quantization for screen content by detecting low contrast edge regionsabstractHigh-Efficiency Video Coding (HEVC) is the newest video coding standard which can significantly reduce the bit rate by 50% compared with existing standards. The key features and new tools in HEVC are designed for natural video sequences captured by a real camera. Different from natural videos, screen content contain much more edges in text and icon regions. The current video coding standards may blur or even remove low contrast edges, which are very important in screen content for human eyes to recognize the character and the icon. Therefore, this paper proposes an effective modification on HEVC to preserve the low contrast edges in screen content. First, discrete laplacian filter is adopted for edge detection, and then we adaptively adjust QPs for low contrast edge regions, which can be detected based on our designed measurement for edge contrast. Experimental results show that nearly all the regions containing low contrast edges can be detected, and the adjustment of QPs for these regions can greatly protect the edges with no RD performance reduction. Hong Zhang 0024, Oscar C. Au, Yongfang Shi, Ketan Tang, Yuanfang Guo |
ISCAS | 2 |
| 2013 | Personal photo album compression and managementabstractThe advance in multimedia technologies have resulted in an explosive growth of pictures in personal computers and in cloud. Typically many pictures taken in the same occasion are similar. The cost to store and transmit them can be very significant. Thus it is important to find an efficient method to store these pictures. This paper proposed a compression scheme for similar images. Our approach is to arrange all the similar images into tree structure then apply video coding technique along each branch. To maximize the inter-image correlation between adjacent photos, we consider the minimum spanning tree (MST) subjecting to a maximum depth limit to ensure fast access to all images. This structure is encoded by the latest video coding technique High Efficiency Video Coding (HEVC), which is reported to has advantage in high definition video/image compression. It also supports deleting, adding and modifying images. Experiments show that the proposed method saved 75% space comparing to JPEG format. Ruobing Zou, Oscar C. Au, Guyue Zhou, Wei Dai 0002, Wei Hu 0003, Pengfei Wan 0001 |
ISCAS | 2 |
| 2013 | Hiding a Secret Pattern into Color Halftone Images
Yuanfang Guo, Oscar C. Au, Ketan Tang, Jiahao Pang |
IWDW | 2 |
| 2013 | Inferring Depth from a Pair of Images Captured Using Different Aperture Settings
Oscar C. Au, Lingfeng Xu, Wenxiu Sun, Wei Hu 0003 |
MMM (2) | 2 |
| 2013 | Reconfigurable hardware-friendly CU-group based merge/skip mode for high efficient video codingabstractMerge/skip mode is one of the most important inter prediction tools adopted in the High Efficiency Video Coding (HEVC) standard which is the state-of-the-art video coding standard. It is very efficient in reducing the side information for the blocks within the same object. However, it is difficult for parallel encoding and decoding due to the data dependency problem between neighboring prediction units (PU). Furthermore, different shapes and positions of PUs would result in different definition of the merge/skip candidate list (MCL), which would lead to potentially extra hardware cost and is not easy to be efficiently implemented by the hardware. To deal with this problem, two reconfigurable hardware-friendly MCL construction schemes are proposed in this paper. The first scheme which is called unified MCL (UMCL) uses one candidate list for all PUs inside the motion estimation region (MER), which is regarded as the basic parallel processing unit for the hardware realization. The second scheme which is named boundary MCL (BMCL) allows different candidate lists for the PUs on the boundary of MER. Both of the two schemes can have flexible parallel degree based on the requirement specification. Experimental results show that UMCL reduces the hardware complexity significantly with little coding performance degradation and BMCL achieves significant coding gain while maintaining the hardware complexity. Wei Dai 0002, Oscar C. Au, Feng Zou 0006, Vinit Jakhetiya |
MMSP | 2 |
| 2013 | Depth map denoising using graph-based transform and group sparsityabstractDepth maps, characterizing per-pixel physical distance between objects in a 3D scene and a capturing camera, can now be readily acquired using inexpensive active sensors such as Microsoft Kinect. However, the acquired depth maps are often corrupted due to surface reflection or sensor noise. In this paper, we build on two previously developed works in the image denoising literature to restore single depth maps-i.e., to jointly exploit local smoothness and nonlocal self-similarity of a depth map. Specifically, we propose to first cluster similar patches in a depth image and compute an average patch, from which we deduce a graph describing correlations among adjacent pixels. Then we transform similar patches to the same graph-based transform (GBT) domain, where the GBT basis vectors are learned from the derived correlation graph. Finally, we perform an iterative thresholding procedure in the GBT domain to enforce group sparsity. Experimental results show that for single depth maps corrupted with additive white Gaussian noise (AWGN), our proposed NLGBT denoising algorithm can outperform state-of-the-art image denoising methods such as BM3D by up to 2.37dB in terms of PSNR. Wei Hu 0003, Xin Li 0005, Gene Cheung, Oscar C. Au |
MMSP | 4 |
| 2013 | Local saliency detection based fast mode decision for HEVC intra codingabstractThe High Efficiency Video Coding (HEVC) is the next generation video coding standard beyond H.264/AVC. Compared with only up to 9 modes for intra prediction in H.264/AVC, HEVC provides 35 intra prediction modes (IPM) to improve coding efficiency, which inevitably poses a huge complexity burden to the encoder. To speed up the HEVC encoder, a novel fast mode decision (FMD) algorithm for HEVC intra prediction is proposed. In the proposed algorithm, we analyzed the costs generated by rough mode decision (RMD), which has already been incorporated in the HM software. We found that the RMD costs listed by mode number generally follow the same trend with the rate-distortion optimization (RDO) costs. Further, the local salient modes, whose RMD costs have a significant drop compared with adjacent modes, tend to be promising competitors for the optimal mode. Based on these observations, we further reduced the number of the candidates for the RDO process. Experimental results show that our proposed algorithm achieves averagely 19.0% (up to 33.6%) encoding time saving whilst causing negligible RD performance loss (0.4% BD-Rate increase on average) compared with HM 7.0 anchor. Yongfang Shi, Oscar C. Au, Hong Zhang 0024, Luheng Jia, Wei Dai 0002 |
MMSP | 2 |
| 2013 | Chroma Replacing and adaptive Chroma Blending for subpixel-based downsamplingabstractSubpixel-based downsampling generates images with higher apparent resolution with the expense of annoying color-fringing artifacts near strong edges. In this paper we propose two methods that find a balance in the tradeoff of apparent resolution and color-fringing artifacts. The first method is called Chroma Replacing in which the color-fringing artifacts are completely removed but the subpixel rendering effect is also removed. The second one is called Chroma Blending in which only the color-fringing artifacts that are strong enough to be noticed are removed, and also the subpixel rendering effect is retained. We also propose two objective measures for measuring the similarity of downsampled image to the original image. Experiment results show that the proposed methods are effective in removing color-fringing artifacts, without harming the high apparent resolution. Ketan Tang, Oscar C. Au, Lu Fang 0001, Yuanfang Guo, Jiahao Pang |
MMSP | 2 |
| 2013 | An analytical study of subpixel-based image down-sampling patterns in frequency domainabstractSubpixel-based image down-sampling is a class of methods that can provide improved apparent resolution of the down-scaled image compared to the pixel-based methods. The frequency characteristics of all possible subpixel-based down-sampling patterns for RGB vertical stripes are analytically studied in this paper. Our proposed algorithm reveals that there are merely seven equivalent energy distributions in the luminance frequency spectrum. To achieve higher luminance resolution, we then calculate and choose the optimal down-sampling pattern with anti-aliasing low-pass filter designed for it so as to maximize the energy of the luminance component within the cut-off shape. Experimental results show that the proposed method provides sharper images compared to the state-of-art subpixel-based methods, with little color distortion. Yonggen Ling, Oscar C. Au, Ketan Tang, Jiahao Pang, Jin Zeng 0004, Lu Fang 0001 |
VCIP | 2 |
| 2013 | Simplified generalized residual prediction in scalable extension of HEVCabstractScalable video coding (SVC), which is an extension of H.264/AVC video coding standard, was introduced to provide scalability in different dimensions for adaptation to heterogeneous network and terminals. After the finalization of the new video coding standard called High Efficiency Video Coding (HEVC), the effort of the standardization committee has been redirected to the investigation of the scalable extension of HEVC. In addition to the basic inter-layer texture prediction mechanism, several other coding tools were proposed for coding performance improvement. Among those coding tools, the one called generalized residual prediction (GRP) scheme achieves most significant coding gain while the consumption of computational power is also huge. In this paper, the GRP mechanism is formulated and analyzed in detail. In addition, the combination of GRP mechanism with merge mode in HEVC is proposed for simplification of the existing GRP mechanism. Results show a better trade-off could be achieved by greatly reducing computational complexity while maintaining most of the coding gain. Oscar C. Au, Haitao Yang 0003, Wei Dai 0002, Hong Zhang 0024 |
VCIP | 2 |
| 2013 | Generalized multihypothesis motion compensated filter for grayscale and color video denoising
Jingjing Dai, Oscar C. Au, Feng Zou 0006 |
Signal Process. | 2 |
| 2013 | 3-D Motion Estimation for Visual Saliency ModelingabstractVisual saliency is a probabilistic estimate of how likely a spatial area in an image or video frame is to attract human visual attention relative to other areas. When existing bottom-up saliency models aggregate low-level features to construct a plausible saliency map, only 2-D motion cues are used as motion features, even though videos typically capture dynamic 3-D scenes. In this paper, we introduce 3-D motion into bottom-up saliency modeling for texture-plus-depth videos. We first propose an efficient 3-D motion estimation algorithm, which computes a 3-D motion vector (3DMV) for each sub-block in the frame. Using the computed 3DMVs, we then derive several saliency channels (called 3DMV channels), which are incorporated into a bottom-up saliency model to obtain enhanced saliency maps. Experiments tracking human gaze show that incorporating our 3DMV channels into bottom-up saliency model significantly improves the accuracy of derived saliency maps. Pengfei Wan 0001, Yunlong Feng, Gene Cheung, Ivan V. Bajic, Oscar C. Au |
IEEE Signal Process. Lett. | 5 |
| 2013 | Universal Binary Semidefinite Relaxation for ML Signal DetectionabstractSemidefinite relaxation (SDR) provides a computationally efficient polynomial-time approximation of the maximum likelihood detector. However, most of the existing works mainly focus on particular signal constellations. In this paper, we propose a universal binary semidefinite relaxation scheme that can handle arbitrary signal constellations in polynomial time. The proposed scheme first binarizes the original signal space to a linearly constrained binary space, and then solves the detection problem through SDR.colorblack{{} A specialized dual barrier method is provided to solve the SDR more efficiently. In addition, we propose to apply on-the-fly decision feedback to further reduce the computational complexity and improve the detection performance. The} proposed binary SDR, together with on-the-fly decision feedback scheme, can provide comparable or better solutions compared to existing SDR methods specialized to specific constellations such as 16-QAM and 8-PSK in terms of computational complexity and symbol error rate. Furthermore, the proposed scheme is universal and can solve any other constellations such as 12-QAM, 32-QAM, or M-PSK. Xiaopeng Fan 0001, Junxiao Song, Daniel Pérez Palomar, Oscar C. Au |
IEEE Trans. Commun. | 4 |
| 2013 | Multichannel Nonlocal Means Fusion for Color Image DenoisingabstractIn this paper, we propose an advanced color image denoising scheme called multichannel nonlocal means fusion (MNLF), where noise reduction is formulated as the minimization of a penalty function. An inherent feature of color images is the strong interchannel correlation, which is introduced into the penalty function as additional prior constraints to expect a better performance. The optimal solution of the minimization problem is derived, consisting of constructing and fusing multiple nonlocal means (NLM) spanning all three channels. The weights in the fusion are optimized to minimize the overall mean squared denoising error, with the help of the extended and adapted Stein's unbiased risk estimator (SURE). Simulations on representative test images under various noise levels verify the improvement brought by the multichannel NLM, compared to the traditional single-channel NLM. In the meantime, MNLF provides competitive performance both in terms of the color peak signal-to-noise ratio and in perceptual quality when compared with other state-of-the-art benchmarks. Jingjing Dai, Oscar C. Au, Lu Fang 0001, Feng Zou 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Color Video Denoising Based on Combined Interframe and Intercolor PredictionabstractAn advanced color video denoising scheme which we call CIFIC based on combined interframe and intercolor prediction is proposed in this paper. CIFIC performs the denoising filtering in the RGB color space, and exploits both the interframe and intercolor correlation in color video signal directly by forming multiple predictors for each color component using all three color components in the current frame as well as the motion-compensated neighboring reference frames. The temporal correspondence is established through the joint-RGB motion estimation (ME) which acquires a single motion trajectory for the red, green, and blue components. Then the current noisy observation as well as the interframe and intercolor predictors are combined by a linear minimum mean squared error (LMMSE) filter to obtain the denoised estimate for every color component. The ill condition in the weight determination of the LMMSE filter is detected and remedied by gradually removing the “least contributing” predictor. Furthermore, our previous work on the LMMSE filter applied in the adaptive luminance-chrominance space (LAYUV for short) is revisited. By reformulating LAYUV and comparing it with CIFIC, we deduce that LAYUV is a restricted version of CIFIC, and thus CIFIC can theoretically achieve lower denoising error. Experimental results verify the improvement brought by the joint-RGB ME and the integration of the intercolor prediction, as well as the superiority of CIFIC over LAYUV. Meanwhile, when compared with other state-of-the-art algorithms, CIFIC provides competitive performance both in terms of the color peak signal-to-noise ratio and in perceptual quality. Jingjing Dai, Oscar C. Au, Feng Zou 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Distributed Wireless Visual Communication With Power Distortion OptimizationabstractThis paper proposes a novel framework called DCast for distributed video coding and transmission over wireless networks, which is different from existing distributed schemes in three aspects. First, coset quantized DCT coefficients and motion data are directly delivered to the channel coding layer without syndrome or entropy coding. Second, transmission power is directly allocated to coset data and motion data according to their distributions and magnitudes without forward error correction. Third, these data are transformed by Hadamard and then directly mapped using a dense constellation (64K-QAM) for transmission without Gray coding. One of the most important properties in this framework is that the coding and transmission rate is fixed and distortion is minimized by allocating the transmission power. Thus, we further propose a power distortion optimization algorithm to replace the traditional rate distortion optimization. This framework avoids the annoying cliff effect caused by the mismatch between transmission rate and channel condition. In multicast, each user can get approximately the best quality matching its channel condition. Our experiment results show that the proposed DCast outperforms the typical solution using H.264 over 802.11 up to 8 dB in video PSNR in video broadcast. Even in video unicast, the proposed DCast is still comparable to the typical solution. Xiaopeng Fan 0001, Feng Wu 0001, Debin Zhao, Oscar C. Au |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2013 | Dependent Joint Bit Allocation for H.264/AVC Statistical Multiplexing Using Convex RelaxationabstractIn this paper, we address the dependent joint bit allocation problem in H.264/AVC statistical multiplexing. In most existing methods, to improve the overall visual quality, the bit allocation is based upon the instantaneous relative frame complexity of different video programs. However, due to the temporal prediction employed in H.264, the influence of a frame on the rate-distortion characteristics of the future frames should be taken into account as well. In this paper, we use our previously proposed simple, but accurate, inter-frame dependency model (IFDM) to quantitatively measure the coding dependency between the current frame and its reference frame. Based on the IFDM, we formulate the dependent joint bit allocation problem, considering both the inter-program relative frame complexity and intra-program coding dependency. We prove that the dependent joint bit allocation (DeJoBA) problem can actually be formulated and relaxed into a convex optimization problem, which can be optimally and efficiently solved. Experimental results suggest that the proposed DeJoBA method can achieve 36.81% and 13.11% bitrate reduction, on average, compared with the equal bit allocation and optimal independent joint bit allocation methods, respectively. Oscar C. Au, Jingjing Dai, Feng Zou 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | An Analytic Framework for Frame-Level Dependent Bit Allocation in Hybrid Video CodingabstractIn this paper, we address the frame-level dependent bit allocation (DBA) problem in hybrid video coding. In most existing methods, the DBA solution is achieved at the expense of high, sometimes even unbearable, computational complexity because of the multipass coding involved. Motivated by this, we propose a model-based approach as an attempt to solve this problem analytically. Leveraging the predictive nature in hybrid video coding, we develop a novel interframe dependency model (IFDM), which enables a quantitative measure of the coding dependency between the current frame and its reference frame. Based on the IFDM, the buffer-constrained frame-level DBA problem is carefully formulated. Finally, the model-based DBA method called IFDM-DBA is derived, in which successive convex approximation techniques are employed to convert the original optimization problem into a series of convex optimization problem s of which the optimal solutions can be obtained efficiently. Experimental results suggest that the proposed IFDM-DBA method can achieve up to a 23% bitrate reduction over the JM reference software of H.264. Oscar C. Au, Feng Zou 0006, Jingjing Dai, Wei Dai 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Luma-Chroma Space Filter Design for Subpixel-Based Monochrome Image DownsamplingabstractIn general, subpixel-based downsampling can achieve higher apparent resolution of the down-sampled images on LCD or OLED displays than pixel-based downsampling. With the frequency domain analysis of subpixel-based downsampling, we discover special characteristics of the luma-chroma color transform choice for monochrome images. With these, we model the anti-aliasing filter design for subpixel-based monochrome image downsampling as a human visual system-based optimization problem with a two-term cost function and obtain a closed-form solution. One cost term measures the luminance distortion and the other term measures the chrominance aliasing in our chosen luma-chroma space. Simulation results suggest that the proposed method can achieve sharper down-sampled gray/font images compared with conventional pixel and subpixel-based methods, without noticeable color fringing artifacts. Lu Fang 0001, Oscar C. Au, Ngai-Man Cheung, Aggelos K. Katsaggelos, Houqiang Li, Feng Zou 0006 |
IEEE Trans. Image Process. | 2 |
| 2012 | Bag of textons for image segmentation via soft clustering and convex shiftabstractWe propose an unsupervised image segmentation method based on texton similarity and mode seeking. The input image is first convolved with a filter-bank, followed by soft clustering on its filter response to generate textons. The input image is then superpixelized where each belonging pixel is regarded as a voter and a soft voting histogram is constructed for each superpixel by averaging its voters' posterior texton probabilities. We further propose a modified mode seeking method — called convex shift — to group superpixels and generate segments. The distribution of superpixel histograms is modeled nonparametrically in the histogram space, using Kullback-Leibler divergence (K-L divergence) and kernel density estimation. We show that each kernel shift step can be formulated as a convex optimization problem with linear constraints. Experiment on image segmentation shows that convex shift performs mode seeking effectively on an enforced histogram structure, grouping visually similar superpixels. With the incorporation of texton and soft voting, our method generates reasonably good segmentation results on natural images with relatively complex contents, showing significant superiority over traditional mode seeking based segmentation methods, while outperforming or being comparable to state of the art methods. Zhiding Yu, Oscar C. Au, Chunjing Xu |
CVPR | 3 |
| 2012 | Distributed Soft Video Broadcast (DCAST) with Explicit MotionabstractVideo broadcasting is a popular application of wireless network. However, the existing layered approaches can hardly accommodate users with diverse channel conditions as analog communication can do. The newly emerged `soft cast' approach, utilizing soft broadcast, provides smooth multicast performance but is not very efficient in inter frame compression. In this work, we propose a motion-aligned wireless video multicast scheme DCAST. Instead of using conventional close loop prediction (CLP), DCAST is based on distributed source coding (DSC) theory. This helps DCAST to avoid error propagation but still achieve high compression efficiency in inter frame coding. DCAST outperforms soft cast 5dB in video PSNR while maintaining the similar graceful degradation feature as soft cast. Xiaopeng Fan 0001, Feng Wu 0001, Debin Zhao, Oscar C. Au, Wen Gao 0001 |
DCC | 4 |
| 2012 | A novel fast two step sub-pixel motion estimation algorithm in HEVCabstractMotion estimation (ME) is one of the most time consuming parts in video coding standard. As fast integer-pixel ME algorithm becoming more and more powerful, it is important to develop fast sub-pixel ME algorithm since the computational complexity of sub-pixel ME compared to integer-pixel ME has become relatively significant. In this paper, a novel fast sub-pixel ME algorithm is proposed. This algorithm first approximates the error surface of the sub-pixel position by a second order function and predicts the minimum point by minimizing the function at half-pixel accuracy. Then another second order approximation within a smaller area which is determined by the previous step is modeled to predict the best sub-pixel position. Experimental results show that the proposed method can reduce the sub-pixel search points significantly with negligible quality degradation. Wei Dai 0002, Oscar C. Au, Lin Sun 0004, Ruobing Zou |
ICASSP | 2 |
| 2012 | Model-based optimal dependent joint bit allocation of H.264/AVC statistical multiplexingabstractIn this paper, we address the dependent joint bit allocation problem in H.264/AVC statistical multiplexing. For most of the existing methods, in order to improve the overall visual quality, the bit allocation is performed upon the relative complexities of different video programs. However, due to the temporal prediction employed in H.264, the influence of current frame to the rate-distortion (R-D) performances of the future frames should be taken in to account as well. The contributions of the paper are two-fold. First, a simple but accurate inter-frame dependency model (IFDM) is introduced which can quantitatively measure the coding dependency between the current coding frame and its reference frame. Second, based on the IFDM, the dependent joint bit allocation problem is revisited, and both the frame complexity and inter-frame dependency are considered in the bit allocation process. Then, it is proved that the dependent joint bit allocation problem can actually be relaxed into a convex optimization problem which can be optimally and efficiently solved. Experimental results demonstrate that the proposed dependent bit allocation method achieves 33.45% and 11.10% bitrate reduction on average compared with the equivalent bit allocation (EBA) and the optimal independent joint bit allocation (OIJBA) methods respectively. Oscar C. Au, Feng Zou 0006, Jingjing Dai |
ICASSP | 2 |
| 2012 | On the determination of capacity parameters in PEE-based reversible image watermarkingabstractIn the existing prediction-error expansion (PEE)-based reversible image watermarking schemes, the capacity parameters are determined in a recursive manner by gradually turning these parameters to fit the payload. This class of method needs multiple rounds of embedding iterations, and hence, it is computationally inefficient. In addition, when multiple capacity parameters need to be handled, the previous methods are generally not capacity-distortion optimized. In this work, we formulate the task of determining the capacity parameters as a capacity-distortion optimization problem, which can be shown to be convex. We also prove that under some conditions, even simple analytical solutions exist. Jiantao Zhou 0001, Oscar C. Au |
ICASSP | 2 |
| 2012 | Color image denoising based on multichannel non-local means fusionabstractIn this paper, we investigate the problem of color image denoising, and propose a novel algorithm called multichannel non-local means fusion (MNLMF), building on the grayscale denoiser non-local means filter. By analyzing and modeling the inter-channel correlation in color images, we formulate the color noise reduction as a minimization problem with a specifically-designed penalty function which fully takes advantages of the inter-channel prior information. The optimal solution is derived consisting of constructing multiple non-local means spanning all three channels and fusing them together. The weights in the fusion are optimized to minimize the overall denoising error. Simulation results under various noise levels demonstrate that when compared to other state-of-the-art algorithms, the proposed MNLMF achieves competitive performance both in terms of the color peak signal-to-noise ratio (cPSNR) and in perceptual quality. Jingjing Dai, Oscar C. Au, Feng Zou 0006, Lu Fang 0001 |
ICIP | 2 |
| 2012 | Analytical study of RGB vertical stripe and RGBX square-shaped subpixel arrangementsabstractThe frequency characteristics of subpixel-based decimation with RGB vertical stripe and RGBX square-shaped subpixel arrangements are studied. To achieve higher apparent resolution than pixel-based decimation, the sampling locations are specially chosen for each of two subpixel arrangements, resulting in relatively small magnitudes of horizontal and vertical aliasing spectra in frequency domain. Thanks to 2-D RGBX square-shaped subpixel arrangement, all the horizontal, vertical, diagonal and anti-diagonal aliasing spectra merely contain low-frequency information, indicating that subpixel-based decimation with RGBX square-shaped panel is more effective in retaining original high frequency details than RGB vertical stripe subpixel arrangement. Lu Fang 0001, Oscar C. Au, Jingjing Dai, Hanli Wang, Ngai-Man Cheung |
ICIP | 2 |
| 2012 | Depth map compression using multi-resolution graph-based transform for depth-image-based renderingabstractDepth map compression is important for efficient network transmission of 3D visual data in texture-plus-depth format, where the observer can synthesize an image of a freely chosen viewpoint via depth-image-based rendering (DIBR) using received neighboring texture and depth maps as anchors. Unlike texture maps, depth maps exhibit unique characteristics like smooth interior surfaces and sharp edges that can be exploited for coding gain. In this paper, we propose a multi-resolution approach to depth map compression using previously proposed graph-based transform (GBT). The key idea is to treat smooth surfaces and sharp edges of large code blocks separately and encode them in different resolutions: encode edges in original high resolution (HR) to preserve sharpness, and encode smooth surfaces in low-pass-filtered and down-sampled low resolution (LR) to save coding bits. Because GBT does not filter across edges, it produces small or zero high-frequency components when coding smooth-surface depth maps and leads to a compact representation in the transform domain. By encoding down-sampled surface regions in LR GBT, we achieve representation compactness for a large block without the high computation complexity associated with an adaptive large-block GBT. At the decoder, encoded LR surfaces are up-sampled and interpolated while preserving encoded HR edges. Experimental results show that our proposed multi-resolution approach using GBT reduced bitrate by 68% compared to native H.264 intra with DCT encoding original HR depth maps, and by 55% compared to single-resolution GBT encoding small blocks. Wei Hu 0003, Gene Cheung, Xin Li 0005, Oscar C. Au |
ICIP | 4 |
| 2012 | Novel temporal domain hole filling based on background modeling for view synthesisabstractView synthesis is a technique to generate images/videos in a virtual viewpoint. In this paper, the dis-occlusion/hole problem in view synthesis is resolved from the temporal domain. By the fact that dis-occlusions belong to the background, firstly we build an online background under a newly designed Switchable Gaussian Model (SGM), owning to its computationally simplicity and scene adaptivity. Then, real textures in the dis-occlusions are able to be recovered with the built background. Experimental results have verified the improvements in rendering quality and computation complexity by comparing to the conventional spatial filling methods and other temporal filling methods. Wenxiu Sun, Oscar C. Au, Lingfeng Xu, Wei Hu 0003 |
ICIP | 2 |
| 2012 | Image de-quantization via spatially varying sparsity priorabstractWe address the problem of image de-quantization, which is also known as bit-depth expansion if the reconstructed 2D signal is re-quantized into higher bit-precision. In this paper, a novel image de-quantization method based on convex optimization theory is proposed, which exploits the spatially varying characteristics of image surface. We test our method on image bit-depth expansion problems, and the experimental results show that proposed method can achieve superior PSNR and SSIM performance. Pengfei Wan 0001, Oscar C. Au, Ketan Tang, Yuanfang Guo |
ICIP | 2 |
| 2012 | New chroma intra prediction modes based on linear model for HEVCabstractRGB-to-YUV conversion reduces the inter-channel redundancy in video compression. However, inter-channel correlation is still observed locally. In HEVC HM4.0, LM mode uses a linear model to predict chroma from luma as a chroma intra prediction mode. The linearity is derived from the reconstructed causal pixels. Specifically, for a NxN chroma block, the above N neighbors and the left N neighbors are used for derivation. Considering the variety of the content inside a block, the linearity derived in LM mode does not always match the linearity of the current block. In this paper, we propose two new modes LML and LMA for chroma intra prediction. Those two modes are quite similar to LM mode, except that they use different neighbors for linearity derivation. By adding the proposed modes, about an average of 0.2%, 5.9%, 6.7% BD-rate reduction is achieved under all intra configuration for Y, Cb and Cr components respectively. Oscar C. Au, Jingjing Dai, Feng Zou 0006 |
ICIP | 2 |
| 2012 | Combined Inter-frame and Inter-color Prediction for Color Video DenoisingabstractAn advanced color video denoising scheme which we call CIFIC based on combined inter-frame and inter-color prediction is presented in this paper. CIFIC performs the denoising filtering in the RGB color space, and exploits both the inter-frame and inter-color correlation in color video signal directly by forming multiple predictors for each color component using all three color components in the current frame as well as the motion-compensated neighboring reference frames. The temporal correspondence is established through the joint-RGB noise-robust motion estimation (ME) which acquires a single motion trajectory for the RGB components. Then the current noisy observation as well as the inter-frame and inter-color predictors are combined by a linear minimum mean squared error (LMMSE) filter to obtain the denoised estimate for every color component. The experimental results verify that CIFIC provides competitive performance both in terms of the objective metric and in perceptual quality when compared with other state-of-the-art algorithms. Jingjing Dai, Oscar C. Au, Feng Zou 0006 |
ICME | 2 |
| 2012 | From 2D Extrapolation to 1D Interpolation: Content Adaptive Image Bit-Depth ExpansionabstractIn this paper, we address the problem of image bit-depth expansion and present a novel method to generate high bit-depth (HBD) images from a single low bit-depth (LBD) image. We expand image bit-depth by reconstructing the least significant bits (LSBs) for the LBD image after it is rescaled to high bit-depth. For image regions whose intensities are neither locally maximum nor minimum, neighborhood flooding is applied to convert 2D interpolation problem into 1D interpolation, for local maxima/minima (LMM) regions where interpolation is not applicable, a virtual skeleton marking algorithm is proposed to convert problematic 2D extrapolation problem into 1D interpolation. At last, a content-adaptive reconstruction model is proposed to obtain the output HBD image. The experimental results show that proposed method significantly outperforms existing methods in PSNR and SSIM without contouring artifacts. Pengfei Wan 0001, Oscar C. Au, Ketan Tang, Yuanfang Guo, Lu Fang 0001 |
ICME | 2 |
| 2012 | Fast sub-pixel motion estimation with simplified modeling in HEVCabstractMotion estimation (ME) is one of the key elements in video coding standard which eliminates the temporal redundancies between successive frames. In recent international video coding standards, sub-pixel ME is proposed for its excellent coding performance. Compared with integer-pixel ME, sub-pixel ME needs interpolation to get the value in sub-pixel position. Also, Hadamard transform will be applied in order to achieve better performance. Therefore, it is becoming more and more critical to develop fast sub-pixel ME algorithms. In this paper, a novel fast sub-pixel ME algorithm is proposed which makes full use of 8 neighboring integer-pixel points. This algorithm models the error surface in sub-pixel position by a second order function with five parameters two times to predict the best sub-pixel position. Experimental results show that the proposed method can reduce the complexity significantly with negligible quality degradation. Wei Dai 0002, Oscar C. Au, Lin Sun 0004, Ruobing Zou |
ISCAS | 2 |
| 2012 | Adaptive depth map filter for blocking artifacts removal and edge preservingabstractIn depth map coding for 3D video coding systems, coding errors in edges can severely affect the synthesis quality. Edge errors mainly compose of two parts: one is blurring and ringing artifact around sharp edge and the other is fake edge caused by blocking artifact. In this paper, we propose an adaptive depth map filter to remove blocking artifacts while preserving depth edges. The proposed filter is designed based on bilateral filter, in which the range kernel parameter is changed adaptively considering the strength of edges and blocking artifacts. Experimental results demonstrate that the proposed depth map filter can achieve up to 0.41 dB gain on the synthesis quality compared to the deblocking filter in MVC at low bit rate. Wei Hu 0003, Oscar C. Au, Lin Sun 0004, Wenxiu Sun, Lingfeng Xu |
ISCAS | 2 |
| 2012 | Interpolation based symmetrical predictor structure for lossless image codingabstractPredictor based algorithms reported in literature uses only causal pixels and hence a non-symmetrical predictor structure for prediction. We observed that the performance of predictor is highly dependent on the predictor structure used. In view of this, we propose a novel interpolation based prediction scheme that enables us to use symmetrical predictor structure. In this sense, we have also used non causal pixels in our scheme. Also, from various interpolation algorithms available, we selected a simple one to ensure decoder simplicity, without any significant loss in performance. From performance evaluation, we found that our algorithm is significantly better in terms of compression performance as compared to some of the computationally complex methods. Vinit Jakhetiya, Sunil Prasad Jaiswal, Anil Kumar Tiwari, Oscar C. Au |
ISCAS | 4 |
| 2012 | Texture optimization for seamless view synthesis through energy minimizationabstractIn this paper, we present a view synthesis method named Visto which aims to generate seamless novel views from a monocular view input. We formulate the problem as joint optimization of inter-view texture similarity and geometry preservation, which significantly differs from traditional view synthesis framework. In this way, the image characteristics of virtual view are inherently inherited from the reference view without introducing any image prior or texture modeling technique. The energy function is minimized using Gauss-Seidel-like approach, and the quality of the virtual view is refined iteratively. The proposed approach also tolerates small depth map errors. Further more, the algorithm is parallel friendly. The simulation results outperform several existing state-of-the-art monocular view synthesis systems. Wenxiu Sun, Oscar C. Au, Lingfeng Xu, Wei Hu 0003, Zhiding Yu |
ACM Multimedia | 2 |
| 2012 | Similar images compression based on DCT pyramid multi-level low frequency templateabstractMedical imaging applications produce a huge amount of similar images. Instead of compressing each image individually, set redundancy compression (SRC) methods remove the inter image redundancy and reduce storage. However, in the previous SRC methods — -MMD, MMP and Centroid methods, the prediction templates for extracting set redundancy are not very efficient, especially when image sets are very large with several clusters. In this paper, inspired by face recognition techniques, a novel lossless SRC method is derived based onDCT pyramid multi-level low frequency template. The approximation subband is used as a prediction template for each image to calculate the residue. Intra prediction is also used to reduce the entropy of the residues. Experiments with 3 sets of MR brain images demonstrate the efficiency of our proposed algorithm in respect to bits/pixel (bpp). Oscar C. Au, Ruobing Zou, Lin Sun 0004, Wei Dai 0002 |
MMSP | 2 |
| 2012 | Modified distortion redistribution problem for High Efficiency Video Coding(HEVC)abstractAdaptive quantization matrix design for different block sizes is one of the possible methods to improve the RD performance in video coding and has recently attracted the focus of many researchers. In this paper, we first analyze the shortcomings of the evenly distributed distortion method which was proposed recently. In order to tackle these problems, we propose two modified methods, method I with relaxed distortion constraints and method II is iterative boundary distortion minimization problem considering variance adaptively. Both problems can be solved using convex optimization effectively and efficiently. Simulations have been conducted based on HM4.0, which is the reference software of the latest High Efficiency Video Coding (HEVC). Simulation results show the effect of our proposed methods. Both methods show their significance when evaluated by RD performance. Lin Sun 0004, Oscar C. Au, Wei Dai 0002, Ruobing Zou |
MMSP | 2 |
| 2012 | Adaptive search range algorithm based on Cauchy distributionabstractIn video coding standard, motion estimation (ME) always plays an important role in reducing temporal redundancies at the expense of higher computational complexity. Many fast ME algorithms have been proposed to reduce the coding complexity. Some papers focus on applying specific search patterns to reduce the search points within a fixed search range (SR). But there are only a few of them trying to reduce the size of SR. In this paper, an adaptive SR algorithm is presented. Cauchy distribution is used to model the SR for one frame and the information of motion vector differences in the neighboring blocks is used to adjust the SR for a particular block. Experimental results show that the proposed algorithm can reduce the size of SR significantly with negligible quality degradation. Wei Dai 0002, Oscar C. Au, Lin Sun 0004, Ruobing Zou |
VCIP | 2 |
| 2012 | Bit-depth expansion using Minimum Risk Based ClassificationabstractBit-depth expansion is an art of converting low bit-depth image into high bit-depth image. Bit-depth of an image represents the number of bits required to represent an intensity value of the image. Bit-depth expansion is an important field since it directly affects the display quality. In this paper, we propose a novel method for bit-depth expansion which uses Minimum Risk Based Classification to create high bit-depth image. Blurring and other annoying artifacts are lowered in this method. Our method gives better objective (PSNR) and superior visual quality as compared to recently developed bit-depth expansion algorithms. Gaurav Mittal, Vinit Jakhetiya, Sunil Prasad Jaiswal, Oscar C. Au, Anil Kumar Tiwari, Dai Wei |
VCIP | 4 |
| 2012 | Recent advances and future directions in multimedia and mobile computing
Jong Hyuk Park 0001, Oscar C. Au, Mikael Wiberg, Changhoon Lee |
Multim. Tools Appl. | 2 |
| 2012 | LMM-based frame-level rate control for H.264/AVC high-definition video coding
Oscar C. Au, Jingjing Dai, Feng Zou 0006 |
Signal Process. Image Commun. | 2 |
| 2012 | Determining the Capacity Parameters in PEE-Based Reversible Image WatermarkingabstractIn the existing prediction-error expansion (PEE)-based reversible image watermarking schemes, the capacity parameters are determined in a recursive manner until the payload is just accommodated. This class of methods requires many rounds of embedding iterations, especially when the payload is high, and therefore, is computationally inefficient. Moreover, when multiple capacity parameters need to be determined, the previous methods cannot guarantee optimality in the capacity-distortion sense. In this work, a capacity-distortion optimization (CDO) framework is built to estimate the optimal capacity parameters. We prove that the CDO problem is convex for any embedding rates, permitting efficient solution. The estimated capacity parameters then serve as starting point to facilitate a local search algorithm to find the optimal capacity parameters with much less rounds of embedding iterations. Experimental results are provided to validate our findings. Jiantao Zhou 0001, Oscar C. Au |
IEEE Signal Process. Lett. | 2 |
| 2012 | Novel 2-D MMSE Subpixel-Based Image Down-SamplingabstractSubpixel-based down-sampling is a method that can potentially improve apparent resolution of a down-scaled image on LCD by controlling individual subpixels rather than pixels. However, the increased luminance resolution comes at price of chrominance distortion. A major challenge is to suppress color fringing artifacts while maintaining sharpness. We propose a new subpixel-based down-sampling pattern called diagonal direct subpixel-based down-sampling (DDSD) for which we design a 2-D image reconstruction model. Then, we formulate subpixel-based down-sampling as a MMSE problem and derive the optimal solution called minimum mean square error for subpixel-based down-sampling (MMSE-SD). Unfortunately, straightforward implementation of MMSE-SD is computational intensive. We thus prove that the solution is equivalent to a 2-D linear filter followed by DDSD, which is much simpler. We further reduce computational complexity using a smallk×kfilter to approximate the much larger MMSE-SD filter. To compare the performances of pixel and subpixel-based down-sampling methods, we propose two novel objective measures: normalizedl1high frequency energy for apparent luminance sharpness and PSNRU(V)for chrominance distortion. Simulation results show that both MMSE-SD and MMSE-SD(k) can give sharper images compared with conventional down-sampling methods, with little color fringing artifacts. Lu Fang 0001, Oscar C. Au, Ketan Tang, Hanli Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Antialiasing Filter Design for Subpixel Downsampling via Frequency-Domain AnalysisabstractIn this paper, we are concerned with image downsampling using subpixel techniques to achieve superior sharpness for small liquid crystal displays (LCDs). Such a problem exists when a high-resolution image or video is to be displayed on low-resolution display terminals. Limited by the low-resolution display, we have to shrink the image. Signal-processing theory tells us that optimal decimation requires low-pass filtering with a suitable cutoff frequency, followed by downsampling. In doing so, we need to remove many useful image details causing blurring. Subpixel-based downsampling, taking advantage of the fact that each pixel on a color LCD is actually composed of individual red, green, and blue subpixel stripes, can provide apparent higher resolution. In this paper, we use frequency-domain analysis to explain what happens in subpixel-based downsampling and why it is possible to achieve a higher apparent resolution. According to our frequency-domain analysis and observation, the cutoff frequency of the low-pass filter for subpixel-based decimation can be effectively extended beyond the Nyquist frequency using a novel antialiasing filter. Applying the proposed filters to two existing subpixel downsampling schemes called direct subpixel-based downsampling (DSD) and diagonal DSD (DDSD), we obtain two improved schemes, i.e., DSD based on frequency-domain analysis (DSD-FA) and DDSD based on frequency-domain analysis (DDSD-FA). Experimental results verify that the proposed DSD-FA and DDSD-FA can provide superior results, compared with existing subpixel or pixel-based downsampling methods. Lu Fang 0001, Oscar C. Au, Ketan Tang, Aggelos K. Katsaggelos |
IEEE Trans. Image Process. | 2 |
| 2012 | Joint Demosaicing and Subpixel-Based Down-Sampling for Bayer Images: A Fast Frequency-Domain Analysis ApproachabstractA portable device such as a digital camera with a single sensor and Bayer color filter array (CFA) requires demosaicing to reconstruct a full color image. To display a high resolution image on a low resolution LCD screen of the portable device, it must be down-sampled. The two steps, demosaicing and down-sampling, influence each other. On one hand, the color artifacts introduced in demosaicing may be magnified when followed by down-sampling; on the other hand, the detail removed in the down-sampling cannot be recovered in the demosaicing. Therefore, it is very important to consider simultaneous demosaicing and down-sampling. Lu Fang 0001, Oscar C. Au, Yan Chen 0007, Aggelos K. Katsaggelos, Hanli Wang |
IEEE Trans. Multim. | 2 |
| 2011 | Nonparametric density estimation on a graph: Learning framework, fast approximation and application in image segmentationabstractWe present a novel framework for tree-structure embedded density estimation and its fast approximation for mode seeking. The proposed method could find diverse applications in computer vision and feature space analysis. Given any undirected, connected and weighted graph, the density function is defined as a joint representation of the feature space and the distance domain on the graph's spanning tree. Since the distance domain of a tree is a constrained one, mode seeking can not be directly achieved by traditional mean shift in both domain. we address this problem by introducing node shifting with force competition and its fast approximation. Our work is closely related to the previous literature of nonparametric methods. One shall see, however, that the new formulation of this problem can lead to many advantages and new characteristics in its application, as will be illustrated later in this paper. Zhiding Yu, Oscar C. Au, Ketan Tang, Chunjing Xu |
CVPR | 2 |
| 2011 | Anti-aliasing filter for subpixel down-sampling based on frequency analysisabstractNowadays, digital pictures are usually captured at very high resolution ranged up to 12 mega-pixels. Limited by low-resolution display, we have to shrink the image. Signal processing theory tells us that optimal decimation requires low-pass filtering with a suit able cut-off frequency followed by down-sampling. In doing so, we need to remove lots of details. Subpixel-based down-sampling, taking advantage of the fact that each pixel on a color LCD is actually composed of individual red, green, and blue subpixel stripes, can provide apparent higher resolution. In this paper, we use frequency domain analysis to explain what happens in subpixel-based down sampling and why it is possible to achieve a higher apparent resolution. According to our frequency domain analysis and observation, the cut-off frequency of the low-pass filter for subpixel-based decimation can be effectively extended beyond the Nyquist frequency using a novel anti-aliasing filter. Experimental results verify that the proposed subpixel down-sampling scheme based on frequency analysis (SDSFA) can give superior results compared with existing pixel-based down-sampling methods. Lu Fang 0001, Ketan Tang, Oscar C. Au, Aggelos K. Katsaggelos |
ICASSP | 3 |
| 2011 | Error compensation and reliability based view synthesisabstractView synthesis offers a great flexibility in generating free viewpoint television (FTV) and 3D video (3DV). However, the depth-image-based view synthesis approach is very sensitive to errors in the camera parameters or poorly estimated depth maps (also called depth images). Because of these errors, three kinds of artifacts (blurring, contour, hole) are possibly introduced during the general synthesis process. Comparing to conventional methods which implement the view synthesis only in ideal case, in this paper, we propose to design an error compensation and reliability based view synthesis system where the potential errors are considered. The main contributions are highlighted as follows: Firstly, the camera parameter errors are compensated by a global homography transformation matrix. Secondly, the depth maps are classified into both reliable and unreliable regions and the reliability based weighting masks are built to blend synthesized images from two different views together. Finally, a reliability depth map based hole-filling technique is used to fill the existing holes. The experimental results demonstrate that these artifacts are efficiently reduced in the synthesized images. Wenxiu Sun, Oscar C. Au, Lingfeng Xu, Sung Him Chui, Chun Wing Kwok |
ICASSP | 2 |
| 2011 | Adaptive Bilateral Filter Considering Local CharacteristicsabstractIn this paper, we propose a simple but effective adaptive bilateral filter considering local characteristics. The presented method exploits the local gaussian gradient information of the processing image and applies bilateral filter with changing the range filter parameter $\sigma_{r}$ adaptively. The proposed adaptive bilateral filter could preserve more details of the image compared with original bilateral filter in edge or texture regions. What is more, it is more robust to the noise. Also compared with the original bilateral filter and the state-of-art wavelet de-noising method, the proposed method also can effectively remove signal noise while preserving the original structure. Lin Sun 0004, Oscar C. Au, Ruobing Zou, Wei Dai 0002 |
ICIG | 2 |
| 2011 | Image Interpolation Using Autoregressive Model and Gauss-Seidel OptimizationabstractIn this paper we propose a simple yet effective image interpolation algorithm based on autoregressive model. Unlike existing algorithms which rely on low resolution pixels to estimate interpolation coefficients, we optimize the interpolation coefficients and high resolution pixel values jointly from one optimization problem. Although the two sets of variables are coupled in the cost function, the problem can be effectively solved using Gauss-Seidel method. We prove the iterations are guaranteed to converge. Experiments show that on average we have over 3dB gain compared to bicubic interpolation and over 0.1dB gain compared to SAI. Ketan Tang, Oscar C. Au, Lu Fang 0001, Zhiding Yu, Yuanfang Guo |
ICIG | 2 |
| 2011 | Video denoising based on transform domain minimum mean square errorabstractDenoising is one of the most common and important tasks in video processing systems and abundant efforts have been made on video denoising nowadays. In this paper, we propose a novel denoising scheme based on minimum mean square error (MMSE) filter in the 2D transform domain, which we call 2DTD-MMSE. The current input noisy frame is processed block-by-block, and for every block, the current noisy observation and multiple prediction blocks found by motion estimation (ME) in denoised previous frames as well as noisy future frames constitute a 2D observed representation array. Afterwards, 2D transform is applied to every block in the representation array, and every transform coefficient of current block is estimated by weighted average of the coefficients in the same frequency position of all the transformed blocks. The weighting coefficients are adaptively determined through MMSE for every block after the estimation of statistical parameters in the transform domain. Experimental results on commonly used test sequences demonstrate that the proposed 2DTD-MMSE achieves comparable or favorable performance when compared to several state-of-the-art algorithms. Jingjing Dai, Oscar C. Au, Feng Zou 0006 |
ICIP | 2 |
| 2011 | A convex-optimization approach to dense stereo matchingabstractWe present a novel convex-optimization approach to solving the dense stereo matching problem in computer vision. Instead of directly solving for disparities of pixels, by establishing the connection between a permutation matrix and a disparity vector, we directly formulate the stereo matching problem as a continuous convex quadratic program in a simple, elegant and straightforward manner without performing any complicated relaxation or approximation. By using CVX, the Matlab software for disciplined convex programming, our method is extremely simple to implement. Oscar C. Au, Lingfeng Xu, Wenxiu Sun, Sung Him Chui, Chun Wing Kwok |
ICIP | 2 |
| 2011 | Multi-scale analysis of color and texture for salient object detectionabstractIn this paper we propose a multi-scale segment-based framework for salient object detection. In this framework texture and color features are used together to provide diverse information of salient object. Segmentation is performed on three different scales so that the object boundary can be accurately captured with high probability. Besides, we propose a novel adaptive feature combination mechanism to combine the saliency maps produced with different features, in which the combining weight of each saliency map is learned using online learning. Experiment results demonstrate that the proposed method significantly outperforms the state-of-the-art methods. Ketan Tang, Oscar C. Au, Lu Fang 0001, Zhiding Yu, Yuanfang Guo |
ICIP | 2 |
| 2011 | Image rectification for single camera stereo systemabstractSingle camera stereo system utilizes mirrors and a single camera for computational stereo, where the mirrors provide extra views needed for stereo and 3D reconstruction. In this paper, we investigate the basic epiploar geometric properties of the single camera stereo image and propose a novel image rectification technique to map the epipolar lines in the original image into the horizontally aligned lines in the rectified image. Besides, the rotation angles of the single camera corresponding to the planar mirror are derived during rectification. Experimental results show the robustness and accuracy of our method. Lingfeng Xu, Oscar C. Au, Wenxiu Sun, Sung Him Chui, Chun Wing Kwok |
ICIP | 2 |
| 2011 | Improved combined inter-intra prediction using spatial-variant weighted coefficientabstractIn current video coding standard H.264/AVC, pixel prediction is applied to reduce spatial and temporal redundancy existed in video signal. Previously, it has been shown that better coding performance is achieved compared to H.264/AVC by combining inter and intra prediction to generate a more accurate prediction. In this paper, an improved combined prediction scheme is presented which allows the video codec to tune weighted coefficient for inter prediction and intra prediction adaptively to local signal characteristics. In order to avoid additional overhead signalling, statistics of already coded neighboring block is analyzed to predict the weighted coefficients of the combined prediction for current macroblock. Compared to H.264/AVC, coding performance is increased by up to 1.9%. And compared to the latest combined prediction method using spatial-invariant weighted coefficients, simulation results show that the proposed scheme achieves additional coding gain of up to 0.66%. Run Cha, Oscar C. Au, Xiaopeng Fan 0001 |
ICME | 2 |
| 2011 | An analysis on bitwise operations in the encrypted domainabstractWhen encrypted data need to be sent to an untrusted computer for processing, it is highly desirable to have a homomorphic cryptosystem that allows the untrusted computer to process the encrypted signals totally without decryption such that, when the encrypted domain result is decrypted, the decrypted value is the same as the equivalent plaintext domain operation. These operations range from basic arithmetic operations to complicated transformations. In this paper, we analyze the existence of equivalent operations in the encrypted domain for four of the common bitwise operations: OR, AND, NOR and NAND in plaintext domain. We will show that such equivalent operations should not exist. For otherwise, if such operations exist, the RSA cryptosystems can be broken by a low-complexity attack. Sung Him Chui, Oscar C. Au, Chun Wing Kwok, Lingfeng Xu, Wenxiu Sun |
ICME | 2 |
| 2011 | Adaptive joint demosaicing and Subpixel-based Down-sampling for Bayer imageabstractA digital camera provided with a Bayer pattern single sensor needs color interpolation to reconstruct a full color image. To show high resolution image on a lower resolution display, it must then be down-sampled. These two steps influence each other, i.e., the color artifacts introduced in demosaicing may be magnified in subsequent down-sampling process and vice versa. Thanks to the fact that LCD displays are actually composed of separable subpixels, which can be individually addressed to achieve a higher effective apparent resolution. This paper presents an Adaptive Joint Demosaicing and Subpixel-based Down-sampling scheme (AJDSD) for single-sensor camera image, where the subpixel-based down-sampling is adaptively and directly applied in Bayer domain, without the process of demosaicing. Simulation results demonstrate that when compared with conventional “demosaicing-first and down-sampling-later” methods, AJDSD achieves superior performance improvement in terms of computational complexity. As for visual quality, AJDSD is more effective in preserving high frequency details, leading to much sharper and clearer results. Lu Fang 0001, Oscar C. Au, Aggelos K. Katsaggelos |
ICME | 2 |
| 2011 | Data hiding in dot diffused halftone imagesabstractIn this paper, we propose two halftone image watermarking methods. Data Hiding by Conjugate Dot Diffusion (DHCDD) and Data Hiding by Dual Conjugate Dot Diffusion (DHD-CDD). DHDCDD is an improved method of DHCDD. Both of these two methods can embed a secret pattern into two halftone images. When the two halftone images are overlaid, the secret pattern will be revealed. Compared to the recent method Noise Balanced Dot Diffusion, the experimental results show that the proposed methods are better in both Correct Decoding Rate and the visual quality of the revealed hidden pattern. Yuanfang Guo, Oscar C. Au, Ketan Tang, Lu Fang 0001, Zhiding Yu |
ICME | 2 |
| 2011 | Optimal distortion redistribution in block-based image coding using successive convex optimizationabstractIn the block-based image coding systems, the pixels at the block boundaries usually have larger distortions than the pixels at other positions. This biased distortion distribution not only generates unpleasant blocking artifact, but potentially decreases the intra-prediction performance and eventually diminishes the coding efficiency. Motivated by this observation, a novel distortion redistribution approach is proposed in this paper which aims to achieve the best rate-distortion performance. Our contributions are summarized in the following aspects: First, the inter-block dependency model (IBDM) is introduced as a quantitative measure of the inter-block dependency. Then, based on the proposed IBDM, the distortion redistribution problem is formulated as a rate minimization problem under a certain distortion constraint. In the end, successive convex optimization method is employed to tackle the optimization problem, and the final optimal distortion can be achieved efficiently. Experimental results demonstrate that compared with the original distortion distribution, significant gain can be obtained with the proposed approach. Oscar C. Au, Feng Zou 0006, Jingjing Dai, Run Cha |
ICME | 2 |
| 2011 | Adaptive depth map assisted matting in 3D videoabstractDepth map is widely adopted and available in the 3D research area. Combining the depth map with the matting techniques is helpful to the original matte and depth image based rendering in 3D. Herein, in this paper, a novel adaptive depth map assisted matting approach with concise integration is presented and applied to achieve favorable matting results. In this approach, the Lagrange-multiplier-free closed form solution is firstly derived to reduce the computation complexity and to increase matting accuracy. Based on the work of Levin et al. on closed form matting, an improved alpha matte is then achieved by introducing an adaptive smoothness criterion which is the function of depth map variance. Finally, the matting system is capable of working in a full automatical way by generating the trimap from the depth information. Simulation results demonstrate that the proposed method is able to efficiently generate an alpha matte with an roughly user specified scribbles or an automatically generated trimap. Wenxiu Sun, Oscar C. Au, Lingfeng Xu, Zhiding Yu |
ICME | 2 |
| 2011 | How anti-aliasing filter affects image contrast: An analysis from majorization theory perspectiveabstractWhen we design an anti-aliasing low pass filter, it is usually an IIR filter. We need to truncate the filter to an FIR filter. One may think that the more taps there are, the better the image quality is. However, we find that there exists an optimal value of tap number that will give the best visual quality. Filters with larger or smaller number of taps will degrade the image quality, due to the fact that the image contrast is reduced. In this paper we analyze this phenomenon using majorization theory and find that the image contrast can be formulated as a Schur convex function on filter coefficients. We also propose an effective method to choose the best filter so that the image contrast is maximized, so as to give best visual quality. Ketan Tang, Oscar C. Au, Lu Fang 0001, Zhiding Yu, Yuanfang Guo |
ICME | 2 |
| 2011 | Compressing similar image sets using low frequency templateabstractIn advance of the imaging capturing technology, large amount of similar images are created. Instead of compressing each similar image individually, removing the inter-image redundancy would reduce the storage and transmission time. However, only a few set redundancy methods are proposed to deal with the problem. In this paper, a new method was derived from a theoretical model by extracting the low frequency in an image set. For the similar images, the values of their low frequency components are very close to that of their neighboring pixel in the spatial domain. In our model, a low frequency template is created and used as a prediction for each image to compute its residue. This model proves the reduction in the entropy and hence the bit rates. Experiments were conducted and proved there were up to 30% gains over the existing methods. Chi Ho Yeung, Oscar C. Au, Ketan Tang, Zhiding Yu, Enming Luo, Yannan Wu, Shing Fat Tu |
ICME | 2 |
| 2011 | Towards robust and efficient segmentation: An approach based on inter-region contour and intra-region content analysisabstractWe address the problem of boundary estimation by formulating it as inter-region contour and intra-region information analysis in the framework of graph-based segmentation. Given an image without any prior information about object model and class, we seek to approximate one's instant perception of visual similarity. The method can serve as a preprocessing step for many higher level operations that require regional support, such as scene understanding and object recognition. We show in this paper that the defined region comparison predicate makes a better boundary estimator than efficient graph-based image segmentation (EGS) - a well known and widely used segmentation method. We further illustrate, by making a small relaxation, further improvement of segmentation performance can be achieved. Experimental results have demonstrated the effectiveness of our proposed method. Zhiding Yu, Oscar C. Au, Ketan Tang, Lingfeng Xu, Wenxiu Sun, Yuanfang Guo |
ICME | 2 |
| 2011 | Security evaluation of a perceptual image hashing scheme based on virtual watermark detectionabstractThis paper evaluates the security of a recently proposed perceptual image hashing scheme based on virtual watermark detection. Under the known-hash attack where the attacker has access to several image/hash vector pairs, we show that the task of estimating the virtual watermark sequences serving as the secret key can be formulated as a simple convex optimization problem, and hence, can be solved efficiently. More specifically, we demonstrate that satisfactory level of estimation accuracy of a watermark sequence of length m could be achieved from approximately 2 · m image/hash vector pairs on average. Experimental results using artificial data and real image data are also provided to verify the effectiveness of our proposed attack approaches. Jiantao Zhou 0001, Oscar C. Au |
ICME | 2 |
| 2011 | Multiple sub-pixel interpolation filters with adaptive symmetry for high-resolution video codingabstractIn order to further improve video coding efficiency, a novel adaptive sub-pixel interpolation filter is presented in this paper. Considering the local image characteristics, the proposed method designs interpolation Alters for sub-pixels in low-frequent and high-frequent areas separately. And in order to reduce the header information, flexible symmetry is assumed for each filter. Experimental results show that this method achieves up to 0.39 dB coding gain which equals to a 11.39% average bit-rate reduction for high-resolution video materials compared to the standard non-adaptive interpolation method of H.264. Compared to the state-of-art 2D non-separable adaptive interpolation scheme, an average bit-rate saving of 0.83% is achieved for high-definition video coding. Run Cha, Oscar C. Au |
ISCAS | 2 |
| 2011 | Frame-level dependent bit allocation via geometric programmingabstractIn this paper, a novel frame-level dependent bit allocation (DBA) method is proposed. Our contribution is two fold: First, the dependency between adjacent frames is quantitatively measured by the introduced inter-frame dependency model(IFDM). The IFDM not only holds for slow video sequences, but remains valid for video sequences of median and high motion as well. Second, based on the IFDM, the conventional DBA problem is revisited and finally it is categorized to a geometric programming problem which can be optimally and effectively solved using interior-point methods. Experimental results show that a significant gain of up to 0.5dB in PSNR can be obtained. Oscar C. Au, Jingjing Dai, Feng Zou 0006 |
ISCAS | 2 |
| 2011 | Sub-pixel downsampling of video with matching highly data re-use hardware architectureabstractSubpixel-based down-sampling is a method that can potentially improve the apparent resolution of a down-scaled image by controlling individual subpixels rather than pixels. However, the increased luminance resolution often comes at the price of chrominance distortion. A major challenge is to suppress color fringing artifacts while maintaining sharpness. In [1], we proposed a novel human visual quality based (HVS) subpixel downsampling method. In this paper, we propose a hardware-friendly subpixel based downsampling scheme based on our previous work which can achieve similar performance as [1] but eliminate all floating point operations with limited bandwidth and much better performance than Direct Pixel based Downsampling (DPD) and Pixel-based downsampling with Anti-aliasing Filter (PDAF). We further propose a hardware architecture for our subpixel downsampling method which is highly data re-useable. We are the first few, if not the first, to implement subpixel downsampling method to video by hardware. The design is implemented with TSMC 0.18um CMOS technology and costs 244k gates. At a clock frequency of 63 MHz, the architecture achieves real-time 1920×1080 subpixel downsampling at 30fps. Oscar C. Au, Jiang Xu 0001, Lu Fang 0001, Run Cha |
ISCAS | 2 |
| 2011 | Alternative Anti-Forensics Method for Contrast Enhancement
Chun Wing Kwok, Oscar C. Au, Sung Him Chui |
IWDW | 2 |
| 2011 | Automatic object segmentation from large scale 3D urban point clouds through manifold embedded mode seekingabstractThis paper presents a system that can automatically segment objects in large scale 3D point clouds obtained from urban ranging images. The system consists of three steps: The first one involves a ground detection process that can detect relatively complex terrain and separate it from other objects. The second step superpixelizes the remaining objects to speed up the segmentation process. In the final step, a manifold embedded mode seeking method is adopted to segment the point clouds. Even though the segmentation of urban objects is a challenging problem in terms of accuracy and problem scale, our system can efficiently generate very good segmentation results. The proposed manifold learning effectively improves the segmentation performance due to the fact that continuous artificial objects often have manifold-like structures. Zhiding Yu, Chunjing Xu, Jianzhuang Liu, Oscar C. Au, Xiaoou Tang |
ACM Multimedia | 4 |
| 2011 | An analytical framework for frame-level dependent bit allocation in hybrid video codingabstractIn this paper, an analytical framework for frame-level dependent bit allocation (DBA) in hybrid video coding is proposed. First, the dependency of neighboring frames is quantitatively measured with the proposed inter-frame dependency model (IFDM). Based on the proposed IFDM, the problem of frame-level DBA among a number of frames of different frame types is studied, and the optimal solution is achieved through successive convex optimization. The prove the validity of the proposed framework, a case study of current state-of-the-art standard H.264/AVC is conducted. Experimental results show that significant gain of up to 0.9dB in PSNR can be obtained. Oscar C. Au, Feng Zou 0006, Jingjing Dai |
MMSP | 2 |
| 2011 | An Adaptive Motion Data Storage Reduction Method for Temporal Predictor
Ruobing Zou, Oscar C. Au, Lin Sun 0004, Wei Dai 0002 |
PSIVT (2) | 2 |
| 2011 | Progressive adaptive correlation estimation(PACE) for WZVCabstractWyner-Ziv video coding is a new paradigm for video compression, in which the prediction frames are possibly only available at the decoder. It exploits the redundancy between the source frame and the prediction frame at the decoder by utilizing their correlation information. However, such correlation information is difficult to estimate due to the absence of the prediction frame at the encoder and the lack of the source frame at the decoder. In this paper, we focus on this issue and propose a progressive adaptive correlation estimation (PACE) approach, in which the correlation information is progressively learned during the decoding process. Compared with our previous TRACE approach, PACE has similar performance in estimation accuracy as well as rate-distortion. Furthermore, it can be potentially integrated into more extensive WZVC applications, such as scalable applications and error-resilience applications. Xiaopeng Fan 0001, Jiayang Gao, Oscar C. Au |
VCIP | 4 |
| 2011 | Rate distortion optimized transform for intra block coding for HEVCabstractThe Discrete Cosine Transform is statistically optimal for first order Markov signals, which is widely used in image and video coding. However, in the intra frame coding of H.264/AVC, it is known that after directional intra prediction, there is still anisotropic features left in residue signals. And the features are related to the intra prediction modes. In order to represent the corresponding features, mode dependent directional transform (MDDT), which is based on Karhunen Loeve transform (KLT), was derived and adopted into JMKTA software, which is a preliminary software platform for High Efficiency Video Coding (HEVC)(). Within the MDDT scheme, each prediction mode has its correspondent transform. However, due to the data variation, this mode dependent classification may not lead to optimal residual data separation. It means that even though the residue blocks are using the same prediction mode, they may exhibit different statistical or structural properties. Sometimes, the mode dependent transform basis functions can not represent the residue signal very well. Therefore, in this paper, we propose a rate distortion optimized transform scheme, which provides a transform selection capability. The proposed scheme is implemented in HMO.9 for HEVC, achieving 3.2% BDBR in Intra Low Complexity (LoCo) condition and 2.0% BDBR in Intra High Efficiency (HE) condition. Feng Zou 0006, Oscar C. Au, Jingjing Dai |
VCIP | 2 |
| 2011 | Frame Complexity Guided Lagrange Multiplier Selection for H.264 Intra-Frame CodingabstractRate-distortion (R-D) optimized mode decision plays an essential role in H.264/AVC encoding. Among all the possible coding modes, it aims to select the one which has the best trade-off between bitrate and compression distortion. Specifically, this tradeoff is tuned through the choice of the Lagrange multiplier. However, in the H.264 reference software, the value of Lagrange multiplier is only determined by the quantization stepsize, with no consideration of the characteristics of the input signal. In this paper, we address the problem of optimal Lagrange determination for intra-frames in a signal dependent environment, and a novel frame complexity guided Lagrange multiplier selection method is introduced. Our contributions are two-fold. First, we propose a novel signal dependent R-D model which uses the average input frame gradient to measure the signal complexity. Second, based on the proposed signal dependent R-D model, we derive the closed-form solution of the optimal Lagrange multiplier. Experimental results suggest that up to 8.18% bitrate reduction can be obtained with negligible additional computation complexity. Oscar C. Au, Jingjing Dai, Feng Zou 0006 |
IEEE Signal Process. Lett. | 2 |
| 2011 | Novel RD-Optimized VBSME With Matching Highly Data Re-Usable Hardware ArchitectureabstractTo achieve superior performance, rate-distortion optimized motion estimation (ME) for variable block size (RDO VBSME) is often used in state-of-the-art video coding systems such as the H.264 JM software. However, the complexity of RDO-VBSME is very high both for software and hardware implementations. In this paper, we propose a hardware-friendly ME algorithm called RDOMFS with a novel hardware-friendly rate-distortion (RD)-like cost function, and a hardware-friendly modified motion vector predictor. Simulation results suggest that the proposed RDOMFS can achieve essentially the same RD performance as RDO-VBSME in JM. We also propose a matching hardware architecture with a novel Smart Snake Scanning order which can achieve very high data re-use ratio and data throughout. It is also reconfigurable because it can achieve variable data re-use ratio and can process variable frame size. The design is implemented with TSMC 0.18 μm CMOS technology and costs 103 k gates. At a clock frequency of 63 MHz, the architecture achieves real-time 1920 × 1080 RDO-VBSME at 30 frames/s. At a maximum clock frequency of 250 MHz, it can process 4096 × 2160 at 30 frames/s. Oscar C. Au, Jiang Xu 0001, Lu Fang 0001, Run Cha |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Introduction to the ICME2010 Special IssueabstractThe 15 papers in this special issue are extended versions of papers presented at the 2010 IEEE International Conference on Multimedia and Expo (ICME), held in Singapore on July 19-23, 2010. These papers cover a wide range of topics in multimedia including user interface, content understanding, mobility, 3-D processing, storage, and forensics. Zicheng Liu 0001, Ming-Ting Sun, Chia-Wen Lin, Zhengyou Zhang, Zhu Liu 0001, Homer H. Chen, Yap-Peng Tan, Oscar C. Au |
IEEE Trans. Multim. | 8 |
| 2010 | Film grain noise removal and synthesis in video codingabstractIn this paper, we propose new techniques for film grain noise removal and synthesis, which can be applied in video coding. Film grain noise is clearly noticeable in high-definition video, and should be preserved for the sake of natural look. However, film grain noise tends to reduce the coding efficiency because of its random nature. In our work, prior to video encoding, essential parameters of film grain noise are estimated and the noise is removed by temporal filtering; at the decoder size, film grain noise is modeled by an autoregressive (AR) model, synthesized using the estimated parameters, and added back to the decoded video. Simulation results show that the proposed algorithms can considerably reduce the bitrate and at the same time achieve good subjective quality. Jingjing Dai, Oscar C. Au, Feng Zou 0006 |
ICASSP | 2 |
| 2010 | Backward error concealment of redundantly coded videoabstractError concealment at the video decoder is to recover erroneous picture region based on correctly decoded region in the same frame or the neighboring frames. However, error concealment cannot give satisfactory result in some cases, e.g. when a whole frame is lost. In this paper, we propose a novel backward error concealment method. While the existing temporal error concealment methods recover an erroneous frame by using its reference frame, the proposed method uses a refreshed future frame to recover the previous corrupted reference frame. This is based on the observation that a future frame can be recovered before its reference frame when some error resilience tools such as intra refresh or redundant picture are used. The advantage of the proposed method is that it can recover most pixels in the corrupted reference frame by inverse motion compensation without any error, since typically the refreshed future frame has both MVs and the residues. In the experiments, the proposed method achieves up to 1.0dB gain over the state-of-the-art temporal error concealment method. Xiaopeng Fan 0001, Oscar C. Au, Jiantao Zhou 0001 |
ICASSP | 2 |
| 2010 | Novel 2-D MMSE subpixel-based image down-sampling for matrix displaysabstractSubpixel-based down-sampling is a method that can potentially improve the apparent resolution of a down-scaled image by controlling individual subpixels rather than pixels. However, the increased luminance resolution often comes at the price of chrominance distortion. In this paper, we propose a new subpixel-based down-sampling scheme for which we design a two-dimensional image reconstruction model. Then, we formulate subpixel-based down-sampling as a MMSE problem and derive the optimal solution called MMSESD. To compare the performance of subpixel-based down-sampling methods, we propose novel objective measures for the apparent luminance resolution and chrominance distortion. Simulation results show that MMSE-SD can give sharper images compared with the conventional down-sampling methods, with little color fringing artifacts. Lu Fang 0001, Oscar C. Au |
ICASSP | 2 |
| 2010 | Security and efficiency analysis of progressive audio scrambling in compressed domainabstractIn this paper, we address the security and efficiency issues of two recently proposed audio scrambling schemes. We show that these two audio scrambling schemes are actually vulnerable against various attacks such as ciphertext-only attack, known-plaintext attack and chosen-plaintext attack. We also demonstrate that one of these two schemes is lack of efficiency in terms of generating the key stream using the dynamic password generator (DPG). Furthermore, we briefly discuss the ways to improve the security and efficiency of these two audio scrambling schemes. Jiantao Zhou 0001, Oscar C. Au |
ICASSP | 2 |
| 2010 | Recent advances in high dynamic range imaging technologyabstractRecently, visual representations using high dynamic range (HDR) images become increasingly popular, with advancement of technologies for increasing the dynamic range of image. HDR image is expected to be used in wide-ranging applications such as digital cinema, digital photography and next generation broadcast, because of its high quality and its powerful expression ability. HDR imaging technologies will spread its sphere of influence in imaging industry. In this paper, we review the state-of-the-art studies and the trends of the HDR imaging, in terms of the following three points: (1) HDR imaging sensor and HDR image generation techniques as image acquisition technologies, (2) encode method of HDR images for efficient transmission and storage, (3) human visual system issues associated with reproduction of HDR image. Yukihiro Bandoh, Guoping Qiu, Masahiro Okuda, Scott Daly, Til Aach, Oscar C. Au |
ICIP | 6 |
| 2010 | Two-level optimized tone mapping for high dynamic range imagesabstractIn this paper, we propose a two-step tone-mapping algorithm to convert a high dynamic range (HDR) image to a low dynamic range (LDR) image. The first step S1 constructs a global tone mapping which optimizes between uniform quantization and histogram equalization. The second step S2 improves the visual quality by optimizing between global operation and local contrast maintenance. Both S1 and S2 can operate independent of each other. Simulation results suggest that proposed tone mapping algorithm can give good visual quality, with S1+S2 better than S1 alone. Simulation results also suggest that proposed S2 can be combined with other existing tone mapping methods to achieve improved local contrast. Chun-Hung Liu, Oscar C. Au, Cheuk Hong Cheng, Ka Yue Yip |
ICIP | 2 |
| 2010 | Inter-channel demosaicking traces for digital image forensicsabstractDigital image forensics seeks to detect statistical traces left by image acquisition or post-processing in order to establish an images source and authenticity. Digital cameras acquire an image with one sensor overlayed with a color filter array (CFA), capturing at each spatial location one sample from the three necessary color channels. The missing pixels must be interpolated in a process known as demosaicking. This process is highly nonlinear and can vary greatly between different camera brands and models. Most practical algorithms, however, introduce correlations between the color channels, which are often different between algorithms. In this paper, we show how these correlations can be used to construct a characteristic map that is useful in matching an image to its source. Results show that our method employing inter-channel traces can distinguish between sophisticated demosaicking algorithms. It can complement existing classifiers based on inter-pixel correlations by providing a new feature dimension. John S. Ho, Oscar C. Au, Jiantao Zhou 0001, Yuanfang Guo |
ICME | 2 |
| 2010 | Laplacian Mixture Model(LMM) based frame-layer rate control method for H.264/AVC high-definition video codingabstractAccurate statistical distribution for estimating the transformed residues is greatly important for us to analyze the rate-distortion behavior of video encoders. However, the previous work pays more attention to those sequences of low-resolution. In this paper, we address the statistical characteristics of DCT coefficients of high-definition videos coded by H.264/AVC. The contribution of this paper is threefold: First, Laplacian Mixture Model (LMM) is proposed to model the residues instead of using Laplacian or Cauchy distributions; the corresponding new rate-distortion model based on LMM is presented next; based on this new rate-distortion model, one frame-layer rate control algorithm is developed. Experimental results showed that the proposed rate control method achieves an improvement of PSNR up to 0.52dB with less visual quality variation compared to JM 11.0. Oscar C. Au, Jingjing Dai, Feng Zou 0006, Mingyuan Yang |
ICME | 2 |
| 2010 | A highly data reusable and standard-compliant motion estimation hardware architectureabstractMotion Estimation (ME) is the most computationally intensive part in the whole video compression process. The ME algorithms can be divided into full search ME (FS) and fast ME (FME). The FS is not suitable for high definition (HD) frame size videos because its relevant high computation load and hard to deal with complex motions in limited search range. A lot of FME algorithms have been proposed which can significantly reduce the computation load compared to FS. Though many kinds of hardware implementations of ME have been proposed, almost all of them fail to consider about the motion vector field (MVF) coherence and rate-distortion (RD) cost which have significant impact to the coding efficiency. In this paper, we propose a hardware friendly ME algorithm and corresponding highly data reusable hardware architecture. Simulation results show that the proposed ME algorithm performs better RD performance than conventional FME algorithm. The proposed reconfigurable ME hardware is implemented in VHDL and mapped to a low cost Xilinx XC3S1500 FPGA. It works at 100MHz and is capable to process 1920 × 1080 of 30fps video format in real time and have very high data reuse ratio. Oscar C. Au, Jiang Xu 0001, Lu Fang 0001, Run Cha |
ICME | 2 |
| 2010 | Graph segmentation revisited: Detailed analysis and density learning based implementationabstractIn this paper we give a step-by-step detailed analysis on the performance of shortest spanning tree (SST) and its revised version, recursive SST (RSST). We further propose a novel segmentation scheme based on recursive SST in the warped domain produced by density estimation. The proposed method is robust for variant natural image input and is easy to implement. Experimental results and comparisons with other methods have illustrated the effectiveness and robustness of the proposed method. Zhiding Yu, Oscar C. Au, Ketan Tang, Lingfeng Xu |
ICME | 2 |
| 2010 | Intra mode dependent quantization error estimation of each DCT coefficient in H.264/AVCabstractH.264/AVC employs intra prediction to reduce spatial redundancy between neighboring blocks. Different directional prediction modes are used to cater diversified video content. Although it achieves quite high coding efficiency, it is desirable to establish a proper theoretical quantization error model under different prediction modes, since this allows us to explain the behavior of existing codecs and to design better ones. Actually, residue after different intra prediction modes exhibits different characteristics in frequency domain. In this paper, an intra mode dependent quantization error estimation is presented. For a complete analysis, we investigate not only the coding distortion in current JM reference software, but also its effect on intra mode dependent residue. Based on the mode dependent residue characteristics, a mode dependent quantization error estimation for each frequency position is proposed. Simulation results show that the proposed model can estimate the mode dependent quantization error with high accuracy. Furthermore, the estimation accuracy remains high for various sequences and QP. Feng Zou 0006, Oscar C. Au, Jingjing Dai |
ICME | 2 |
| 2010 | Color video denoising based on adaptive color space conversionabstractDenoising is one of the most common and important task in video processing systems and abundant efforts have been made on video denoising nowadays. Multihypothesis motion compensated filter (MHMCF) is an effective video denoising method, which combines multiple hypotheses obtained from motion estimation through a number of reference frames by weighted average to suppress noise. However, MHMCF only considers denoising of grayscale video signal. In this paper, we apply MHMCF to color video denoising, where the RGB video is first transformed to the luminance-color difference space before denoising. Instead of using traditional YCbCr color conversion, we propose a novel color conversion matrix which is adaptive to the noise variance in R, G, B channels. Simulation results demonstrate that our proposed color space conversion method can successfully improve the denoising performance for color video. Jingjing Dai, Oscar C. Au, Feng Zou 0006 |
ISCAS | 2 |
| 2010 | Subpixel-based down-sampling via Min-Max Directional ErrorabstractSubpixel-based down-sampling is a method that can potentially improve the apparent resolution of a down-scaled image by controlling individual subpixels rather than pixels. However, the increased luminance resolution often comes at the expense of chrominance distortion. In this paper, we formulate the subpixel-based down-sampling as a Min-Max problem (Min-Max Directional Error) which we call MMDE. Unfortunately, the solution of MMDE is computational intensive, especially for large images. We thus relax the MMDE by determining the maximum error based on HVS, which largely reduces the number of constraints. We call such relaxation as MMDE-VR (Visual Relaxation). Simulation results illustrate that MMDE-VR can effectively reduce visible color fringing artifacts while still maintaining sharpness. Lu Fang 0001, Oscar C. Au |
ISCAS | 2 |
| 2010 | An efficient motion vector coding algorithm based on adaptive predictor selectionabstractMotion Estimation is a core part of modern video coding standards, which significantly improves the compression efficiency. On the other hand, motion information takes considerable portion of compressed bit stream, especially in low bit rate situation. In this paper, an efficient motion vector prediction algorithm is proposed to minimize the bits used for coding the motion information. Several spatial and temporal neighboring motion vectors are selected as the motion vector predictor (MVP) candidates. By applying template matching to each block, a near-optimal MVP can be obtained both at the encoder and decoder side, thus no predictor index is needed to signal to the decoder. We also embed the MVP into current motion estimation process. Furthermore, a correction technique is executed as a remedy when template matching picks out a non-efficient predictor. Simulation results indicate that a bit rate reduction of up to 7.29% over H.264/AVC is achieved by the proposed scheme. Oscar C. Au, Jingjing Dai, Feng Zou 0006 |
ISCAS | 2 |
| 2010 | Cryptanalysis of chaotic convolutional coderabstractIn this paper, we evaluate the security of a recently proposed joint error correction and encryption approach called chaotic convolutional coder, which integrates the chaotic encryption into the convolutional coding. We show that the probability of recovering the key vector controlling the chaotic switch is at least 0.289 under known-plaintext attack, if the number of available plaintext/ciphertext pairs p is equal to the constraint length k of the chaotic convolutional coder. In the case that p = k + e, where e > 0, we prove that the probability to recover the key vector is lower bounded by 1-2-e. We also consider the security of the chaotic con-volutional coder under chosen-plaintext attack. We propose two approaches to efficiently derive the key vector without leaving tractable pattern to the register. In particular, one of these two methods based on an efficient erasure code is capable of recovering the key vector with complexity of order O(k log k). Jiantao Zhou 0001, Oscar C. Au |
ISCAS | 2 |
| 2010 | Motion vector coding algorithm based on adaptive template matchingabstractMotion estimation as well as the corresponding motion compensation is a core part of modern video coding standards, which highly improves the compression efficiency. On the other hand, motion information takes considerable portion of compressed bit stream, especially in low bit rate situation. In this paper, an efficient motion vector prediction algorithm is proposed to minimize the bits used for coding the motion information. First, a possible motion vector predictor (MVP) candidate set (CS) including several scaled spatial and temporal predictors is defined. To increase the diversity of predictors, the spatial predictor is adaptively changed based on current distribution of neighboring motion vectors. After that, adaptive template matching technique is applied to remove non-effective predictors from the CS so that the bits used for the MVP index can be significantly reduced. As the final MVP is chosen based on minimum motion vector difference criterion, a guessing strategy is further introduced so that in some situations the bits consumed by signaling the MVP index to the decoder can be totally omitted. The experimental results indicate that the proposed method can achieve an average bit rate reduction of 5.9% compared with the H.264 standard. Oscar C. Au, Jingjing Dai, Feng Zou 0006 |
MMSP | 2 |
| 2010 | Edge-based Adaptive Directional Intra PredictionabstractH.264/AVC employs intra prediction to reduce spatial redundancy between neighboring blocks. Different directional prediction modes are used to cater diversified video content. Although it achieves quite high coding efficiency, it is desirable to analyze its drawbacks in the existing video coding standard, since it allows us to design better ones. Basically, even after intra prediction, the residue still contains a lot of edge or texture information. Unfortunately, these high frequency components consume a large quantity of bits and the distortion is usually quite high. Based on this drawback, an Edge-based Adaptive Directional Intra Prediction is proposed (EADIP) to reduce the residue energy especially for the edge region. In particular, we establish an edge model in EADIP, which is quite flexible for natural images. Within the model, the edge splits the macroblock into two regions, each being predicted separately. In implementation, we consider the current trend of mode selection and complexity issues. A mode extension is made on INTRA 16 × 16 in H.264/AVC. Experimental results show that the proposed algorithm outperforms H.264/AVC. And the proposed mode is more likely to be chosen in low bitrate situations. Feng Zou 0006, Oscar C. Au, Jingjing Dai |
PCS | 2 |
| 2010 | An adaptive unsupervised approach toward pixel clustering and color image segmentation
Zhiding Yu, Oscar C. Au, Ruobing Zou, Weiyu Yu, Jing Tian 0002 |
Pattern Recognit. | 2 |
| 2010 | Successive refinement based Wyner-Ziv video compression
Xiaopeng Fan 0001, Oscar C. Au, Ngai-Man Cheung, Yan Chen 0007, Jiantao Zhou 0001 |
Signal Process. Image Commun. | 2 |
| 2010 | Error recovery of variable length code over BSC with arbitrary crossover probabilityabstractThe error recovery capability of variable length code (VLC) has been considered as an important performance and design criterion in addition to its coding efficiency. However, almost all of the existing methods for evaluating the error recovery capability of VLC assume that the transmission fault is a random single bit inversion. In this paper, we consider a more generalized problem of precisely evaluating the error recovery capability of VLC in the case that the encoded bit stream is transmitted over a BSC with arbitrary crossover probability. By making use of the Perron-Frobenius Theorem, we derive a very simple expression for the exact mean error propagation rate (MEPR), and show that the variance of error propagation rate (VEPR) is zero. We also prove that in the regime of very low crossover probability, the mean error propagation length (MEPL) derived for single inversion error case approaches a scaled value of the MEPR. Furthermore, we briefly discuss the problem of evaluating the error detection capability of non-exhaustive code over BSC. Jiantao Zhou 0001, Oscar C. Au |
IEEE Trans. Commun. | 2 |
| 2010 | Transform-Domain Adaptive Correlation Estimation (TRACE) for Wyner-Ziv Video CodingabstractWyner-Ziv video coding (WZVC) is a newly emerged video coding scheme which compresses the input video frames with the side information (SI) frames only available at the decoder. WZVC exploits the statistics between the source frame and the SI frame at the decoder by utilizing their correlation information. This correlation information is important but also difficult to estimate due to the absence of the SI frame at the encoder, and the lack of the source frame at the decoder. In this paper, we focus on this problem and propose a novel transform-domain adaptive correlation estimation method called TRACE for WZVC. In TRACE, the correlation information is progressively learned during the decoding process of each frame. Within TRACE, we also propose a convex optimization based band-level correlation estimation method which is optimal in the sense of minimizing the theoretical bit rate. Experiments suggest that, when applied in motion compensated interpolation-based low complexity WZVC, TRACE yields competitive results against the state-of-the-art correlation estimation algorithms. More importantly, different from the existing coefficient-level correlation estimation algorithms, the proposed TRACE can be applied in many other WZVC schemes and can provide considerable gain over the popular band-level correlation estimation methods. Xiaopeng Fan 0001, Oscar C. Au, Ngai-Man Cheung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Integration of Recursive Temporal LMMSE Denoising Filter Into Video CodecabstractThe presence of noise can dramatically affect the efficiency of video compression systems. For performance improvement, most practical video compression systems adopt a denoising filter as a pre-processing module for the video encoder, or as a post-processing module for the video decoder, but the complexity introduced by denoising can be very high. This paper first presents a recursive temporal linear minimum mean squared error (LMMSE) filter for video denoising. Based on the analysis of the hybrid video compression process, two novel schemes are presented, one for video encoding and the other for video decoding, in which the proposed recursive temporal LMMSE filter is seamlessly integrated into the encoding and the decoding processes, respectively. For both of these two schemes, the denoising is implemented with nearly no extra computation introduced. Experimental results validate the effectiveness of the proposed schemes on encoding and decoding noisy video sequences. Oscar C. Au, Mengyao Ma, Peter Hon-Wah Wong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Edge-Directed Error ConcealmentabstractIn this paper we propose an edge-directed error concealment (EDEC) algorithm, to recover lost slices in video sequences encoded by flexible macroblock ordering. First, the strong edges in a corrupted frame are estimated based on the edges in the neighboring frames and the received area of the current frame. Next, the lost regions along these estimated edges are recovered using both spatial and temporal neighboring pixels. Finally, the remaining parts of the lost regions are estimated. Simulation results show that compared to the existing boundary matching algorithm [1] and the exemplar-based inpainting approach [2] , the proposed EDEC algorithm can reconstruct the corrupted frame with both a better visual quality and a higher decoder peak signal-to-noise ratio. Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan, Ming-Ting Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | A Fast Converging Algorithm for Acoustic Echo Cancellation in Time-Varying ChannelsabstractIn this paper, we propose an acoustic echo cancellation method called MPNLMS++ which can adapt to time varying channels with fast convergence. The MPNLMS++ can converge much faster than conventional algorithms when the channel changes from dispersive to sparse, and at least as fast as other fast algorithms in other situations. The effectiveness of MPNLMS++ is verified in the simulation. Tsz-Kin Hon, Oscar C. Au |
AVSS | 2 |
| 2009 | Transcoding based robust streaming of compressed videoabstractA variety of techniques have been proposed to enhance the error robustness of the video streaming system. However, most of them improves the error resilience during compression rather than after compression. In this paper, we propose a novel transcoding based scheme called lossless inter frame transcoding (LIFT) scheme to improve the error resilience of existing compressed video stream. In the LIFT scheme, inter coded blocks are selectively transcoded into new kind of blocks called ‘L-block’. At the decoder, the L-block can be transcoded back to the original P-block when the prediction is available and can also be robustly decoded as I-block when the prediction is unavailable. By offline transcoding and online adjusting the ratio of P-blocks and L-blocks, the proposed streaming server achieves error robustness scalability. Experimental results demonstrate the correctness and effectiveness of the proposed method. Xiaopeng Fan 0001, Oscar C. Au, Mengyao Ma, Ling Hou, Jiantao Zhou 0001, Ngai-Man Cheung |
ICASSP | 2 |
| 2009 | Secure Exp-Golomb coding using stream cipherabstractIn this paper, we propose a secure Exp-Golomb coding scheme by incorporating with a stream cipher. Different from the traditional case of using stream cipher where the key stream is directly XORed with the plaintext, we here use the key stream to control the switching between two coding conventions (leading zeros and leading ones). Security analysis results show that the proposed system can provide high level of security with the same coding efficiency and negligible additional cost, compared with a regular Exp-Golomb coding. This scheme could potentially be applied to the state-of-the-art multimedia compression systems, e.g., H. 264, to offer security features. Jiantao Zhou 0001, Oscar C. Au, Amanda Yannan Wu |
ICASSP | 2 |
| 2009 | Parallel rate-distortion optimized intra mode decision on multi-core graphics processors using greedy-based encoding ordersabstractRate-distortion (RD) optimized intra-prediction mode selection can lead to significant improvement in coding efficiency in intra-frame encoding. However, it would incur considerable increase in encoding complexity. In this paper, we investigate how multi-core Graphics Processing Units (GPUs) can be efficiently utilized to undertake the task of RD optimized intra mode selection in AVS and H.264 video encoding. Achieving efficient GPU-based intra mode decision, however, could be non-trivial. It is because the mode decision of the current block would depend on the reconstructed data of the neighboring blocks. Therefore, the coding modes of neighboring blocks would need to be computed first before that of the current block can be determined. This dependency poses challenge to computation on multi-core GPUs, which rely heavily on parallel data processing to achieve superior speedups. To address this issue, we analyze the data dependency in intra mode decision, and propose novel greedy-based encoding orders to achieve highly parallel processing. We also prove that the proposed greedy-based orders are optimal in terms of execution time. Experimental results suggest that the proposed GPU-based intra mode decision compares favorably to the counterpart implemented on a single-core CPU. Ngai-Man Cheung, Oscar C. Au, Man Cheung Kung, Xiaopeng Fan 0001 |
ICIP | 2 |
| 2009 | Adaptive correlation estimation for general Wyner-Ziv video codingabstractWyner-Ziv video coding (WZVC) is a new paradigm for video compression with the prediction frames possibly only available at the decoder. It exploits the statistics between the source frame and the prediction frame at the decoder by utilizing their correlation information. This correlation information is important but also difficult to estimate due to the absolute absence of the prediction frame at the encoder, and the lack of the source frame at the decoder. In this paper, we focus on this issue and derive a coefficient-level adaptive correlation model for general Wyner-Ziv video coding. Based on this model, we propose an online transform-domain adaptive correlation estimation (TRACE) approach, in which the correlation information is progressively learned during the decoding process. In our experiments, the proposed approach outperforms the existing approaches up to 4 dB. More importantly, different from existing coefficient-level variance estimation approaches, the proposed on-line TRACE is applicable for not only low complexity WZVC but also other WZVCs such as flexible WZVC as demonstrated in the experiments. Xiaopeng Fan 0001, Oscar C. Au, Ngai-Man Cheung |
ICIP | 2 |
| 2009 | Subpixel-based image downsampling-some analysis and observationabstractOften we need to shrink a high resolution image (e.g. 10-mega pixel) in order to display it on a low resolution display (e.g mobile phone). Signal processing theory tells us that optimal decimation requires low-pass filtering with a suitable cutoff frequency followed by downsampling. In doing so, we need to remove lots of details in the original high resolution image. In this paper, we review some little known results on an interesting topic called subpixel rendering, which can provide apparent higher resolution at the expense of color fringing artifacts. We attempt to explain what happens and why this is even possible. Lu Fang 0001, Oscar C. Au, Yi Yang 0041, Weiran Tang |
ICME | 2 |
| 2009 | LMMSE frequency merging for demosaickingabstractFor raw images captured by most digital cameras, every pixel has only on color in R, G and B. Kinds demosaicking algorithms are proposed for interpolating the missing two colors. In this article, the relationships inter and intra color channels are analyzed, and basing on the features, we propose a method to divide raw images into sub images and merge them in frequency domain with linear combination. Optimal weights are calculated with estimation values and raw values based on minimum mean square error criteria. Experiments results with different estimations are presented and discussed. Weiran Tang, Oscar C. Au, Yi Yang 0041, Lu Fang 0001 |
ICME | 2 |
| 2009 | A robust spatial-temporal line-warping based deinterlacing methodabstractIn this paper, a line-warping based deinterlacing method will be introduced. The missing pixels in interlaced videos can be derived from the warping of pixels in horizontal line pairs. In order to increase the accuracy of temporal prediction, multiple temporal-line pairs, selected according to constant velocity model, are used for warping. The stationary pixels can be well-preserved by accuracy stationary detection. A soft switching between spatial-temporal interpolated values and temporal average is introduced in order to prevent unstable switching. Owing to above novelties, the proposed method can yield higher visual quality deinterlaced videos than conventional methods. Moreover, this method can suppress most deinterlaced visual artifacts, such as line-crawling, flickering and ghost-shadow. Shing Fat Tu, Oscar C. Au, Yannan Wu, Enming Luo, Chi Ho Yeung |
ICME | 2 |
| 2009 | A novel deringing method based on MAP image restorationabstractLong has it become a hot topic that reconstructing images of better visual quality from one or a serial of degraded ones. Although there are thousands of different restoration methods, in this paper, we focus on removing the ringing artifact caused by lossy video compression. Being a sort of restoration method, we choose the max-a-posterior (MAP) method to model this optimization problem. Quantification of the ringing artifacts serves as a prior information of the images. So it is also analyzed in this paper. The MAP optimization is further solved using a gradient decent solver. Although this is a quite classical method, there are still lots of problems with it. For settling them, we transform the solver to a filter format operation named iterative optimization filter. Experiments show that, such method could give an averaging 0.3 dB gain and in special regions more than 0.7 dB gain. What is more, as analyzed in this paper, the method is very preferable in terms of hardware implementation in several aspects. Yannan Wu, Oscar C. Au, Enming Luo, Dennis Tu, Leo Yeung |
ICME | 2 |
| 2009 | Perceptual compressive sensing for image signalsabstractHuman eyes have different sensitivity to different frequency components of image signals, typically, low frequency components are relatively more crucial to the perceptual quality of images than high frequency components. Based on this observation, we propose a novel sampling scheme for compressive sensing framework by designing a weighting scheme for the sampling matrix. By adjusting the weighting coefficients, we can tune the structure of the sampling matrix to favor the frequency components that are important to human perception, so that those components could be more precisely recovered in the reconstruction procedure. Experimental results reveal that our proposed scheme can greatly enhance the performance of compressive sensing framework in both PSNR and visual quality without increasing the complexity of the framework structure or computational procedure. Yi Yang 0041, Oscar C. Au, Lu Fang 0001, Weiran Tang |
ICME | 2 |
| 2009 | Bit-depth Expansion by Contour Region ReconstructionabstractColor bit-depth is an important attribute to image quality. However, the precision in various image capture devices limits the color bit-depth and introduces loss in visual quality. Expanding the color bit-depth is an important image enhancement issue. A good bit-depth expansion system manipulates low color bit-depth image for best visual quality as displayed on high color bit-depth monitors. However, in most color bit-depth expansion algorithms, severe contouring effect is observed in smooth gradient area which degrades the visual quality. In this paper, a novel approach is proposed. By considering the distance from contour edges, fine gradient value are applied to fill the contour gaps to achieve gradual transaction. Cheuk Hong Cheng, Oscar C. Au, Chun-Hung Liu, Ka Yue Yip |
ISCAS | 2 |
| 2009 | On Improving the Robustness of Compressed Video by Slepian-Wolf based Lossless TranscodingabstractA variety of techniques have been proposed to enhance the error robustness of the video streaming system. However, most of them improve the error resilience during compression rather than after compression. In this paper, we propose a novel transcoding based scheme called Slepian-Wolf based inter frame transcoding (SWIFT) to improve the error resilience of existing compressed video stream. In the SWIFT scheme, inter coded blocks are selectively transcoded into new kind of blocks called ‘X-block’. At the decoder, the X-block can be transcoded back to the original P-block when there is no error in the prediction, and can also be robustly decoded as I-block when there are errors in the prediction. In the experiments, the proposed SWIFT scheme does not introduce transcoding distortion as expected, and always improves the robustness of the compressed video at all packet loss rate. Compared with the H.264 based transcoder, SWIFT achieves better RD performance and error resilience performance. Xiaopeng Fan 0001, Oscar C. Au, Mengyao Ma, Ling Hou, Jiantao Zhou 0001, Ngai-Man Cheung |
ISCAS | 2 |
| 2009 | A New Adaptive Subpixel-based Downsampling Scheme using Edge DetectionabstractIn this paper, a new adaptive subpixel-based downsampling scheme is proposed. Inside this scheme, we take full advantage of subpixels by adaptively choosing the sample directions based on edge information which has not been addressed before. Then, an adaptive filter is designed to suppress color fringing artifacts. Moreover, a good cut-off frequency is derived and deployed in our filter to obtain extra information. Simulation results illustrate that the proposed adaptive subpixel-based downsampling scheme successfully improves the resolution while efficiently removes visible color fringing artifacts. Lu Fang 0001, Oscar C. Au, Yi Yang 0041, Weiran Tang |
ISCAS | 2 |
| 2009 | A Novel Ray-space based Color Correction Algorithm for Multi-view VideoabstractIn multi-view video, color inconsistency among different views always exists because of imperfect camera calibration, CCD noise, etc. Since color inconsistency greatly reduces the coding efficiency and rendering quality of multi-view video, a novel ray-space based color correction algorithm is proposed in this paper. Firstly, for each epipolar plane image (EPI) in ray-space domain, feature points are extracted to form a corresponding feature EPI (FEPI). Secondly, radon transform is applied to each FEPI to detect corresponding points from different views and the average color is calculated from the detected corresponding points. Finally, for each viewpoint image, the optimal color correction matrix is calculated by minimizing the error energy between the color of the current view and the average color based on the least square error criteria. Experimental results show that the proposed algorithm greatly improves the color consistency among different views. Moreover, the coding efficiency of the corrected multi-view images is greatly improved compared to that of the original ones and the ones corrected by histogram matching method [1]. Ling Hou, Oscar C. Au, Xiaopeng Fan 0001, Mengyao Ma |
ISCAS | 2 |
| 2009 | Frequency Selection and Merging with Universal Matrices for Color Filter Array DemosaickingabstractIn most digital cameras, color filter array (CFA) is used for sampling only one color value for each pixel on CCD. So the other two color values need to be interpolated. The interpolation process is commonly known as demosaicking. In this paper, we discuss the relationship among the down sampled CFA images in frequency domain. According to the correlation, we propose an interpolation method that select specified parts from every down sampled frequency CFA image and merge them together. For more accurate selection, we calculate the minimum error of points in frequency images, and produce some universal matrices to record the source of every point in merging. And then, interpolation is based on the universal matrices. This algorithm is a non-adaptive method and regular in hardware implementation. Weiran Tang, Oscar C. Au, Yi Yang 0041, Lu Fang 0001 |
ISCAS | 2 |
| 2009 | A Novel Multiple Description Video Coding based on H.264/AVC Video Coding StandardabstractMultiple description coding (MDC) is a source coding technique that exploits path diversity to solve packet losses over error-prone channels. In this paper, we propose an improved drift-free multi-state MDC method based on H.264/AVC coding scheme. At the encoder side, we compress original video into multiple independent H.264 streams with different coding parameters, which can help us to control correlations between the descriptions. At the decoder side, each description is considered as a noisy observation of the original video, and a linear minimum mean square error (LMMSE) based merge algorithm is proposed to combine the descriptions. Experimental results show that the proposed algorithm can achieve better coding efficiency and visual quality than temporary MDC present in [1]. The error resilience ability is also improved by the fact that each frame is coded twice with different parameters. Oscar C. Au, Jiang Xu 0001, Zhiqin Liang, Yi Yang 0041, Weiran Tang |
ISCAS | 2 |
| 2009 | Image Registration Method based on Local High Order ApproachabstractImage registration based on gradient and least square optimization technique is one of the most edge-cutting registration algorithms. Such method, especially useful for sub-pixel motion, searches for the best motion in an iterative way. This paper solves the same motion registration problem following this direction. And the well-known Gauss-Newton method (GNM) is employed here as the optimization tool. To achieve a speed-up and reduction of the arithmetic calculation, a simplified high order approach(SHoA) used to calculate several parameters for GNM is introduced. Detailed complexity analysis and performance comparison are presented showing that such an approach is a better trade-off between registration error and the number of math operations. Yannan Wu, Oscar C. Au, Enming Luo, Chi Ho Yeung, Shing Fat Tu |
ISCAS | 2 |
| 2009 | Motion vector coding based on predictor selection and boundary-matching estimationabstractIn most of recent video coding standards based on block-based hybrid coding scheme, especially in the state-of-the-art H.264/AVC standard, motion vector information occupies a considerable portion of the whole compressed bitstream. There-fore, the efficient coding of motion vectors has become an essential objective to further reduce the bitrate. In this paper, we propose a novel motion vector coding method based on predictor selection and estimation. First, the optimal motion vector predictor (MVP) is chosen from a predefined predictor candidate set consisting of both spatial predictors and temporal predictors to minimize the number of bits used for encoding motion vector difference (MVD). Then, in order to avoid sending extra bits for informing the decoder which candidate is selected by the encoder, boundary-matching (BM) estimation is applied at the decoder side to find out the optimal predictor. The basic principle of BM estimation is to preserve the spatial continuity of boundaries between currently reconstructed block and its neighbors. Simulation results show that compared to the original H.264/AVC codec, the proposed scheme improves coding efficiency for various video sequences. Jingjing Dai, Oscar C. Au, Feng Zou 0006 |
MMSP | 2 |
| 2009 | Maximum-likelihood versus maximum a posteriori based local illumination and color correction algorithm for multi-view videoabstractIn multi-view video, illumination and color inconsistency among different views always exist because of imperfect camera calibration, CCD noise, camera positions and orientations, etc. Since illumination and color inconsistency greatly reduce the coding efficiency and rendering quality of multiview video, effective illumination and color correction modules are necessary for practical multi-view video processing system. In this paper, we proposed two local illumination and color correction algorithms. In these two algorithms, the correction matrix is estimated by applying maximum likelihood (ML) and maximum a posteriori (MAP) methods respectively. According to the Bayes rule, the MAP estimate is determined by two terms: error conditional density model (likelihood model)and priori conditional density model. Experimental results show that both the ML and MAP based correction matrices greatly improve the illumination and color consistency among different views. Moreover, images corrected by MAP based correction matrix look much nicer than those corrected by ML based correction matrix. Ling Hou, Oscar C. Au, Xiaopeng Fan 0001, Jiantao Zhou 0001 |
MMSP | 2 |
| 2009 | A fast NL-Means method in image denoising based on the similarity of spatially sampled pixelsabstractAs one of the best image denoising methods, the Non-Local Means(NL-Means)algorithm[5] proposed by Buades et al. generates state-of-the-art performance. However, due to the high computational complexity, it is difficult to be directly used in practical applications. In this paper, a novel fast algorithm based on the similarity of spatially sampled pixels is introduced. Compared with other fast approaches, the result has shown that our method always uses the shortest time. Meanwhile, it keeps a similar or even better visual results. A maximum of 0.9dB improvement can be attained in comparison with the original method when the noise variance is small. Oscar C. Au, Jingjing Dai, Feng Zou 0006 |
MMSP | 2 |
| 2009 | Motion vector predictor selection for the enhancement layer in the H.264/AVC extension-spatial SVCabstractScalable Video Coding (SVC) has been approved as the extension of the H.264/AVC video coding standard recently [1]. In current spatial scalability scheme, the technique to examine both inter modes with residual prediction and without residual prediction for enhancement layers can achieve the highest possible coding efficiency, but it's typically one of the most time-consuming parts of a video encoder. In this paper, we propose a method to skip half of motion estimation processes under “inter modes without residual prediction” based on the motion vector predictor selection under “inter modes with residual prediction”. Experimental results show that our proposed scheme can achieve up to 24% time saving for motion estimation in the enhancement layer and meanwhile the coding efficiency can be preserved very well. Enming Luo, Oscar C. Au, Yannan Wu, Shing Fat Tu, Chi Ho Yeung |
PCS | 2 |
| 2009 | Reweighted Compressive Sampling for image compressionabstractCompressive Sampling (CS), is an emerging theory which points us a promising direction of designing novel efficient data compression techniques. However, the conventional CS adopts a non-discriminated sampling scheme which usually gives poor performance on realistic complex signals. In this paper we propose a reweighted Compressive Sampling for image compression. It introduces a weighting scheme into the conventional CS framework whose coefficients are determined in encoding side according to the statistics of image signals. Experimental results demonstrate that our proposed method notably outperforms the conventional Compressive Sampling framework in coding performance in the sense that the reconstruction quality is greatly enhanced with same number of measurements and computational complexity. Yi Yang 0041, Oscar C. Au, Lu Fang 0001, Weiran Tang |
PCS | 2 |
| 2009 | Wyner-Ziv-based bidirectionally decodable video coding
Xiaopeng Fan 0001, Oscar C. Au, Yan Chen 0007, Jiantao Zhou 0001, Mengyao Ma, Peter Hon-Wah Wong |
J. Vis. Commun. Image Represent. | 2 |
| 2009 | Simultaneous MAP-Based Video Denoising and Rate-Distortion Optimized Video EncodingabstractIn this paper, a simultaneous MAP-based video denoising and rate-distortion optimized video encoding algorithm is proposed. We begin with formulating the denoising problem as a maximumaposteriori(MAP) estimate problem. Then, according to the Bayes rule, we show that the MAP estimate is determined by two terms: noise conditional density model andprioriconditional density model. Based on the assumptions that the noise satisfies Gaussian distribution and the priori model is measured by the bit-rate, the MAP estimate can be expressed as a rate distortion optimization problem. With this, we are able to simultaneously perform MAP-based video denoising and rate-distortion optimized video encoding under some assumptions. Moreover, we describe in details how to select suitable coding parameters, i.e., quantization parameter, mode, motion vector, reference index, and regularization parameter. Finally, we conduct several experiments to verify our proposed algorithm. Yan Chen 0007, Oscar C. Au, Xiaopeng Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Highly Parallel Rate-Distortion Optimized Intra-Mode Decision on Multicore Graphics ProcessorsabstractRate-distortion (RD)-based mode selections are important techniques in video coding. In these methods, an encoder may compute the RD costs for all the possible coding modes, and select the one which achieves the best trade-off between encoding rate and compression distortion. Previous papers have demonstrated that RD-based mode selections can lead to significant improvements in coding efficiency. RD-based mode selections, however, would incur considerable increases in encoding complexity, since these methods require computing the RD costs for numerous candidate coding modes. In this paper, we consider the scenario where software-based video encoding is performed on personal computers or game consoles, and investigate how multicore graphics processing units (GPUs) may be efficiently utilized to undertake the task of RD optimized intra-prediction mode selections in audio and video coding standards and H.264 video encoding. Achieving efficient GPU-based intra-mode decisions, however, could be nontrivial for two reasons. First, intra-mode decision tends to be sequential. Specifically, the mode decision of the current block would depend on thereconstructed dataof the neighboring blocks. Therefore, the coding modes of neighboring blocks would need to be computed first before that of the current block can be determined. This dependency poses challenges to GPU-based computation, which relies heavily on parallel data processing to achieve superior speedups. Second, RD-based intra-mode decision may require conditional branchings to determine the encoding bit-rate, and these branching operations may incur substantial performance penalties when being executed on GPUs due to pipeline architectural designs. To address these issues, we analyze the data dependency in intra-mode decision, and propose novel greedy-based encoding orders to achieve highly parallel processing of data blocks. We also prove that the proposed greedy-based orders are optimal in our problem, i.e., they require the minimum number of iterations to process a video frame given the dependency constraints. In addition, we propose a method to estimate the coding rate suitable for GPU implementation. Experimental results suggest our proposed solution can be more than 50 times faster than the previously proposed parallel intra-prediction, since our work can efficiently exploit the massive parallel opportunity in GPUs. Ngai-Man Cheung, Oscar C. Au, Man Cheung Kung, Peter Hon-Wah Wong, Chun-Hung Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | A Novel Analytic Quantization-Distortion Model for Hybrid Video CodingabstractA proper theoretical quantization-distortion model for hybrid video coding is always desirable, since this allows us to explain the behavior of existing codecs and to design better ones. However, due to the existence of motion-compensated prediction, hybrid video coding introduces interframe dependency into the encoded video, which makes its quantization-distortion characteristics difficult to analyze. In this paper, a joint analysis of quantization and motion-compensated prediction is presented. For a complete analysis, we investigate not only the distortion that quantization introduces into video signal, but also its effect on motion-compensated prediction. Based on the joint analysis, a quantization-distortion model of hybrid video coding is proposed. Our extensive experimental results show that the proposed model can estimate the quantization-distortion curve of hybrid video coding with high accuracy. Furthermore, the estimation accuracy remains high for various video sequences and encoder configurations. Oscar C. Au, Mengyao Ma, Zhiqin Liang, Peter Hon-Wah Wong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Error Resilient Video Coding Using B Pictures in H.264abstractSince the quality of compressed video is vulnerable to errors, video transmission over unreliable Internet is very challenging today. Multi-hypothesis motion-compensated prediction (MHMCP) has been shown to have error resilience capability for video transmission, where each macroblock is predicted by a linear combination of multiple signals (hypotheses). B picture prediction is a special case of MHMCP. In H.264/AVC, the prediction of B pictures is generalized such that both of the two predictions can be selected from the past pictures or from the subsequent pictures. The multiple reference picture framework in H.264/AVC also allows previously decoded B pictures to be used as references for B picture coding. In this paper, we will discuss the error resilience characteristics of the generalized B pictures in H.264/AVC. Three prediction patterns of B pictures are analyzed in terms of their error-suppressing abilities. Both theoretical models (picture level error propagation) and simulation results are given for the comparison. Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Simultaneous RD-optimized rate control and video de-noisingabstractIn this paper, we propose a simultaneous rate control and video de-noising algorithm based on rate distortion optimization. According to our previous works, video de-noising can be performed by using rate distortion optimization with a lower bound quantization parameter (QP) constraint, where the lower bound QP is determined by the noise variance. Then, we find that the macroblock level rate control method in H.264 can be seen as an approximate solution of a rate distortion optimization problem with a specified rate distortion function. Based on these two studies, we integrate the video de-noising problem and rate control problem to a rate distortion optimization problem. We show the convexity of the problem and derive the optimal solution. To reduce the complexity, we propose to use a suboptimal solution based on simply thresholding. Some experiments are conducted to demonstrate the efficiency and effectiveness of the proposed method. Yan Chen 0007, Oscar C. Au |
ICASSP | 2 |
| 2008 | Temporal search range prediction based on a linear model of motion-compensated residueabstractAn efficient temporal search range prediction method is proposed to reduce the complexity of multiple reference frames motion estimation (MRFME) in video coding. Based on a linear model of motion-compensated residue, the behavior of residues under MRFME is investigated, and the gain of multiple reference frames is analyzed. The proposed method utilizes the current residue to estimate the gain of searching more reference frames, and predicts the temporal search range that maintains the coding performance with minimum complexity. Experimental results show that the proposed scheme can significantly reduce the complexity in motion estimation while the degradation of the coding performance is negligible. Oscar C. Au, Mengyao Ma, Zhiqin Liang, Peter Hon-Wah Wong |
ICASSP | 2 |
| 2008 | Complexity adaptive H.264 encoding using multiple reference framesabstractThe state-of-the-art H.264/AVC video coding standard achieves significant improvements in coding efficiency by introducing many new coding techniques. However the computation complexity is inevitably increased during both the encoding and decoding process. Many previous works, such as fast motion estimation and fast mode decision algorithms, have been proposed aiming at reducing the encoder complexity while maintaining the coding efficiency. In this paper, we propose a new encoding approach which accounts for the decoding complexity. Simulation results show that the decoding complexity can be reduced by up to 15% in terms of motion compensation operations, which is the most complex part of the decoder, while maintaining the R-D performance with only about 0.1 dB degradation. Sui-Yuk Lam, Oscar C. Au, Peter Hon-Wah Wong |
ICASSP | 2 |
| 2008 | Cryptanalysis of secure arithmetic codingabstractThis work investigates the security issues of the recently proposed secure arithmetic coding (AC), which is an encryption scheme incorporating the interval splitting AC with a series of symbol and codeword permutations. We propose a chosen-ciphertext attack which is capable of recovering the key vectors for codeword permutations with complexity O(N), where N is the symbol sequence length. After getting the key vectors for codeword permutations, we can remove the code-word permutation module, and the resulting system has already been shown to be insecure in the original paper [5]. Jiantao Zhou 0001, Oscar C. Au, Peter Hon-Wah Wong, Xiaopeng Fan 0001 |
ICASSP | 2 |
| 2008 | Image deblocking using convex optimizationabstractImages encoded at low-bit rate may suffer from blocking artifacts, which can dramatically degrade the visual quality. In this paper, a novel approach to image deblocking is presented. Based on the analysis of image coding process and the property of natural images, an objective function and a set of constraint functions are proposed, and image deblocking is formulated as a convex optimization problem which can be easily solved using numerical methods. The feasibility of the convex optimization problem is utilized to detect the true object edges and avoid blurring. Experimental results demonstrates the effectiveness of the proposed approach. Oscar C. Au, Mengyao Ma, Xiaopeng Fan 0001, Peter Hon-Wah Wong, Daniel Pérez Palomar |
ICIP | 2 |
| 2008 | Joint security and performance enhancement for secure arithmetic codingabstractThis paper studies the joint security and performance enhancement of secure arithmetic coding (AC) for digital rights management applications. The proposed cryptosystem incorporates the interval splitting AC with a simple bit-wise XOR operation step. Security analysis results show that the proposed scheme provides satisfactory level of security against the cipher-only attack, the chosen-plaintext attack and the chosen-ciphertext attack. Due to the elimination of the input symbol-wise permutation step, our proposed scheme can be extended conveniently to any context-based coding scenarios. In addition, the implementation complexity of our proposed scheme is lower than the original secure AC. Finally, we suggest a selective encryption version of our proposed scheme, which further reduces the implementation complexity. Jiantao Zhou 0001, Oscar C. Au, Xiaopeng Fan 0001, Peter Hon-Wah Wong |
ICIP | 2 |
| 2008 | Improved bidirectionally decodable Wyner-Ziv video codingabstractReverse playback is one of the most common video cassette recording (VCR) functions for video streaming systems. However, the predictive processing techniques employed in traditional hybrid video coding schemes severely complicate the reverse-play operation. In this paper, we enhance our previously proposed bidirectionally decodable Wyner-Ziv video coding scheme which supports both forward decoding and backward decoding. We derive that in our scheme the optimal Lagrangian multiplier for the backward motion estimation should be averagely two times larger than for the forward motion estimation. The new multiplier contributes 0.2dB gain in average. We propose an optimal P-frame/M-frame selection scheme to improve rate-distortion performance when the video is transmitted over error prone channels. The new scheme outperforms both H.264 and our previous scheme at all tested loss rate, and gain up to 0.55dB over our previous scheme in low loss rate case. Xiaopeng Fan 0001, Oscar C. Au, Jiantao Zhou 0001, Mengyao Ma |
ICME | 2 |
| 2008 | Image characteristic oriented tone mapping for high dynamic range imagesabstractThis paper presents a novel and efficient tone mapping algorithm for converting high dynamic range (HDR) images back to low dynamic range (LDR) images for displaying purpose because of the limited contrast ratio of common displays and printers. As the ratio between the maximum and minimum values of common HDR images is always very large and also the population usually deflects to one side, for convenient processing, most researchers first take the logarithm on the luminance layer or use another adaptive mapping to shorter the range of the distribution in their tone mapping methods. However, these mappings have already distorted the original imagespsila characteristics. In this paper, there is no such adverse mapping applied on the luminance layer in the proposed tone mapping algorithm. The paper does produce a tone reproduction curve to convert HDR images to LDR images. Adaptive techniques are also manipulated to provide better visual quality. The whole process is automatic and no parameter is required for manual input. The result will be a superior visual quality tone mapped LDR image with original HDR imagepsilas characteristics. Chun-Hung Liu, Oscar C. Au, Peter Hon-Wah Wong, Man Cheung Kung |
ICME | 2 |
| 2008 | Secure Lempel-Ziv-Welch (LZW) algorithm with random dictionary insertion and permutationabstractIn this paper, we propose an efficient encryption scheme by introducing randomness into the Lempel-Ziv-Welch (LZW) algorithm. This scheme utilizes random dictionary insertion and permutation, and incorporates with a bit-wise XOR module. Security analysis results show that the proposed scheme provides high level of security without any coding efficiency loss, compared with a standard LZW algorithm. Jiantao Zhou 0001, Oscar C. Au, Xiaopeng Fan 0001, Peter Hon-Wah Wong |
ICME | 2 |
| 2008 | Bidirectionally decodable Wyner-Ziv video codingabstractInter frame prediction technique significantly improves the compression efficiency in the hybrid video coding schemes. However, this technique causes the decoding dependency of each inter frame on all of its reference frames. This dependency complicates the reverse play operation which is the most common video cassette recording (VCR) functions. This dependency also causes error propagation when the video is transmitted over error prone channel. In this paper, we propose a novel bidirectionally decodable Wyner-Ziv video coding scheme which relaxes this inter frame dependency. The proposed bidirectionally decodable Wyner-Ziv frame can be decoded by using whether forward prediction or backward prediction as side information at the decoder, i.e. the proposed stream supports forward decoding and backward decoding simultaneously. Compared with the other schemes which support reverse playback, our scheme requires much lower bandwidth and smaller storage space. In error resilient test, our scheme outperforms H.264 up to 4dB at same bitrate. Our proposed frames also support video splicing and stream switching at arbitrary time point like I-frames. Xiaopeng Fan 0001, Oscar C. Au, Yan Chen 0007, Jiantao Zhou 0001, Mengyao Ma |
ISCAS | 2 |
| 2008 | Video decoder embedded with temporal LMMSE denoising filterabstractUnder noisy circumstance, the encoded video can be corrupted by noise and the decoded video may be very noisy. In this paper, a novel scheme is proposed for decoding noisy video bitstreams. First a temporal linear minimum mean squared error (LMMSE) estimator for video denoising is presented. Based on the analysis of hybrid video decoder, this temporal LMMSE denoising filtering is seamlessly embedded into hybrid video decoding process. The operations performed for denoising are very few, and the total complexity of the proposed scheme is very similar to that of a regular decoder. Experimental results show that compared to the standard decoder, with the proposed scheme, the subjective quality of the decoded video can be significantly improved. Oscar C. Au, Mengyao Ma, Peter Hon-Wah Wong |
ISCAS | 2 |
| 2008 | Bit-depth expansion by adaptive filterabstractBit-depth expansion is important for displaying a low bit-depth image in a high bit-depth monitor. Existing methods tend to give disturbing contouring or blurring artifacts. In this paper, we propose a novel, simple and efficient adaptive method to increase bit-depth taking advantage of the existing techniques to give superior image quality. Chun-Hung Liu, Oscar C. Au, Peter Hon-Wah Wong, Man Cheung Kung, Shen Chang Chao |
ISCAS | 2 |
| 2008 | A multi-hypothesis decoder for multiple description video codingabstractMultiple Description Coding (MDC) can be used as an Error Resilience (ER) technique for video coding. In case of transmission errors, Error Concealment (EC) can be combined with MDC to reconstruct the lost frame, such that the propagated error to the following frames is reduced. In this paper we propose a novel algorithm based on a Multi-hypothesis Decoder (MHD), to improve the reconstructed video quality of MDC over packet loss networks. Both subjective and objective results show that MHD can help to achieve a better video quality than a traditional EC algorithm. Mengyao Ma, Oscar C. Au, Xiaopeng Fan 0001, Ling Hou, Shueng-Han Gary Chan |
ISCAS | 2 |
| 2008 | Error recovery of variable length codes over BSC with arbitrary crossover probabilityabstractThe error recovery capability of variable length code (VLC) has been considered as an important performance and design criterion in addition to its coding efficiency. However, almost all of the existing methods for evaluating the error recovery capability of VLC assume that the transmission fault is a random single bit inversion. In this paper, we consider a more generalized problem of precisely evaluating the error recovery capability of VLC in the case that the encoded bit stream is transmitted over a BSC with arbitrary crossover probability. By making use of the Perron-Frobenius Theorem, we derive a very simple expression for the exact mean symbol error rate (MSER) in Levenshtein distance sense. We also prove that in the very low crossover probability region, the mean error propagation length (MEPL) derived for single inversion error case approaches a scaled value of MSER. In addition, we briefly discuss the error recovery of VLC over Gilbert-Elliott channel, which is one of the simplest and practical models for a channel with memory. Jiantao Zhou 0001, Oscar C. Au, Xiaopeng Fan 0001, Peter Hon-Wah Wong |
ISIT | 2 |
| 2008 | A novel radon based Ray-Space interpolation algorithmabstractRay-Space interpolation is one of the key technologies to make Ray-Space based FTV (Free Viewpoint Television) system feasible. Since Ray-Space data is composed of straight lines with different slopes, the problem in Ray-Space interpolation is to find the slope of the straight lines which tells us the interpolation direction. In this paper, a radon based Ray-Space interpolation algorithm is proposed. First, feature points of epipolar plane image (EPI) are extracted to form a feature EPI (FEPI). Then, radon transform is applied to FEPI to find the possible interpolation direction. Finally, a new cost function is proposed to improve the smoothness of disparity map and find the optimal interpolation direction. Experimental results show that both the performance and the robustness of the proposed algorithm are much higher than that of the traditional pixel matching based interpolation (PMI) and block matching based interpolation (BMI). Ling Hou, Oscar C. Au, Xiaopeng Fan 0001, Mengyao Ma |
MMSP | 2 |
| 2008 | Alternate motion-compensated prediction for error resilient video coding
Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan, Xiaopeng Fan 0001, Ling Hou |
J. Vis. Commun. Image Represent. | 2 |
| 2008 | Video Error Concealment Using Spatio-Temporal Boundary Matching and Partial Differential EquationabstractError concealment techniques are very important for video communication since compressed video sequences may be corrupted or lost when transmitted over error-prone networks. In this paper, we propose a novel two-stage error concealment scheme for erroneously received video sequences. In the first stage, we propose a novel spatio-temporal boundary matching algorithm (STBMA) to reconstruct the lost motion vectors (MV). A well defined cost function is introduced which exploits both spatial and temporal smoothness properties of video signals. By minimizing the cost function, the MV of each lost macroblock (MB) is recovered and the corresponding reference MB in the reference frame is obtained using this MV. In the second stage, instead of directly copying the reference MB as the final recovered pixel values, we use a novel partial differential equation (PDE) based algorithm to refine the reconstruction. We minimize, in a weighted manner, the difference between the gradient field of the reconstructed MB in current frame and that of the reference MB in the reference frame under given boundary condition. A weighting factor is used to control the regulation level according to the local blockiness degree. With this algorithm, the annoying blocking artifacts are effectively reduced while the structures of the reference MB are well preserved. Compared with the error concealment feature implemented in the H.264 reference software, our algorithm is able to achieve significantly higher PSNR as well as better visual quality. Yan Chen 0007, Yang Hu 0006, Oscar C. Au, Houqiang Li, Chang Wen Chen |
IEEE Trans. Multim. | 3 |
| 2008 | Error Concealment for Frame Losses in MDCabstractMultiple description coding(MDC) is an effectiveerror resilience(ER) technique for video coding. In case of frame loss,error concealment(EC) techniques can be used in MDC to reconstruct the lost frame, with error, from which subsequent frames can be decoded directly. With such direct decoding, the subsequent decoded frames will gradually recover from the frame loss, though slowly. In this paper we propose a novel algorithm usingmultihypothesis error concealment(MHC) to improve the error recovery rate of any EC in the temporal subsampling MDC. In MHC, the simultaneous temporal-interpolated frame is used as an additional hypothesis to improve the reconstructed video quality after the lost frame. Both subjective and objective results show that MHC can achieve significantly better video quality than direct decoding. Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan, Peter Hon-Wah Wong |
IEEE Trans. Multim. | 2 |
| 2007 | Halftone Image Data Hiding with Block-Overlapping Parity CheckabstractIn this paper, we propose an innovative algorithm, namely block-overlapping parity check (BOPC), that can be applied to the existing halftone image data hiding algorithms data hiding smart pair toggling (DHSPT) to achieve improvement in visual quality. The proposed algorithm utilizes the properties of block-overlapping parity check to reduce the number of pair toggling required in DHSPT. Our experiments suggest that the proposed algorithm reduces the number of pixel pair toggling significantly and has a better performance in terms of modified peak-signal-to-noise ratio (MPSNR), especially when the watermark payload is high. Richard Y. M. Li, Oscar C. Au, Carman K. M. Yuk, Shu-Kei Yip, Sui-Yuk Lam |
ICASSP (2) | 2 |
| 2007 | Advanced Macro-block Entropy Coding in H.264abstractExisting video encoders, such as MPEG4, H.263 and H.264, adopt variable length coding (VLC) as entropy coding method, focusing on residue data coding. In low bit-rate video coding, larger quantization parameters are used to give a smaller number of bits spent on residue data. As a result, the header overhead is a dominant factor of yielding the overall bit rates. In this paper, macro-block (MB) header behavior is investigated. It is found that the relative percentage of the number of bits of MB header increases with quantization parameter (QP) and header correlation between neighboring MBs is high. Based on these MB header properties, we propose an advanced coding algorithm for entropy coding of MB header. The experimental results suggest that our proposed algorithm achieves lower total number of encoded bits over JM10.2, up to 10% bit reduction, with the same PSNR quality. At the same bit rates, our algorithm has about 0.3-0.4 dB PSNR gain. Chi Wah Wong, Oscar C. Au, Raymond Chi-Wing Wong |
ICASSP (1) | 2 |
| 2007 | Exact Symbol Error Rate for Variable Length Codes Over Binary Symmetric ChannelabstractIn this paper, we analyze the error recovery performance of variable length codes (VLCs) transmitted over binary symmetric channel (BSC). Simple expressions for the exact mean symbol error rate (MSER) and the exact variance of symbol error rate (VSER) for any crossover probability pe are presented. We also prove that the mean error propagation length (MEPL) derived for single bit inversion error case is a scaled value of MSER when pe tends to zero. Comparisons with simulations demonstrate the accuracy of the MSER and VSER expressions. Jiantao Zhou 0001, Xiaopeng Fan 0001, Zhiqin Liang, Oscar C. Au |
ICASSP (3) | 4 |
| 2007 | Advanced Real-time Rate Control in H.264abstractMost existing rate control schemes in the literature use one rate model and calculate quantization parameters of the macro-blocks (MB), regardless of MB types. In advanced video coding standards such as H.264, MBs belong to more advanced MB types, such as skipped and non-skipped MBs. In non-skipped MBs, the encoder determines whether each of 8times8 luminance sub-blocks and 4times4 chrominance sub-block of a MB is to be encoded, giving the different number of sub-blocks at each MB encoding times. As a result, a traditional single rate model is insufficient to represent each MB accurately. In this work, it is found that different MB types have different rate behavior. Under different conditions of MB types, we establish novel different rate models and distortion models. Our rate control scheme is proposed based on these models. The experimental results suggest that our scheme can achieve PSNR gain over JM10.2 and TMN8. Chi Wah Wong, Oscar C. Au, Raymond Chi-Wing Wong |
ICIP (1) | 2 |
| 2007 | Color Demosaicking using Direction CategorizationabstractWe propose a novel color demosaicking algorithm using direction categorization. Each pixel is classified as vertical, horizontal or smooth before interpolation. The categorization is based on gradient change within same channel, color differences and neighbors categories, so it explores relationship between intra-and inter-color channels. Color artifacts in reconstructed images are significantly reduced because directions of interpolation across edges are greatly avoided. Experimental results show that our proposed algorithm has high PSNR and the visual quality of reconstructed images is also obviously improved. Carman K. M. Yuk, Oscar C. Au, Richard Y. M. Li, Sui-Yuk Lam |
ICIP (4) | 2 |
| 2007 | Maximum a Posteriori Based (MAP-Based) Video Denoising VIA Rate Distortion OptimizationabstractIn this paper, a maximum a posteriori based (MAP-based) video denoising algorithm is proposed. According to the Bayes rule, the MAP estimate is determined by two terms: noise conditional density model and priori conditional density model. Based on the assumptions that the noise satisfies Gaussian distribution and the priori model is measured by the bit rate, the MAP estimate can be expressed as a rate distortion optimization problem. In order to find a suitable Lagrangian parameter, we re-write the problem as a constraint minimization problem by setting the rate as an objective function and the distortion as a constraint. In this way, we find that the Lagrangian parameter is determined by the distortion constraint. Fixing the distortion constraint, we can get the optimal Lagrangian parameter, which leads to the optimal denoising result. Some experiments are conducted to demonstrate the efficiency and effectiveness of the proposed method. Yan Chen 0007, Oscar C. Au, Xiaopeng Fan 0001, Peter Hon-Wah Wong |
ICME | 2 |
| 2007 | Wyner-Ziv Successive Refinement of Video and Rate Distortion AnalysisabstractIn Wyner-Ziv video coding system, motion estimation efficiency is much lower than that in conventional video coding system because current frame is not available when doing motion estimation. In this paper, we propose a successive resolution refinement algorithm to improve motion estimation efficiency. Based on our rate distortion analysis, we derive optimal down-sample ratio for two-stage successive resolution refinement system. We also analyze the performance of multistage case, and find that it approaches the performance of ideal motion compensated Wyner-Ziv video coding system, with at most 2.17 dB loss in PSNR. Experimental results demonstrate the correctness of the analysis and show that the proposed method out-performs original bit-plane refinement scheme up to 2.5 dB, with much lower complexity. Xiaopeng Fan 0001, Oscar C. Au, Yan Chen 0007, Jiantao Zhou 0001, Peter Hon-Wah Wong |
ICME | 2 |
| 2007 | Data Hiding with Tree Based Parity CheckabstractIn this paper, we propose a novel algorithm namely tree based parity check (TBPC) that can be applied to most of the existing data hiding algorithms to achieve improvement in visual quality. In data hiding process, distortion is created when the original image is modified. Most existing data hiding algorithms try to minimize the visual artifacts introduced by the modifications. The proposed algorithm tries to reduce the probability of modifying the original host image. Theoretical analysis and experimental results are given in this paper. Both measures suggest that an improvement in visual quality is achieved in the watermarked image. Richard Y. M. Li, Oscar C. Au, Kelvin K. Lai, Carman K. M. Yuk, Sui-Yuk Lam |
ICME | 2 |
| 2007 | Joint Decoding of Multiple Video StreamsabstractDue to the storage limit, the original raw video may not exist after compressing it into multiple video bitstreams using different compression parameters. In this paper, we suggest a least square error (LSE) algorithm to jointly decode the multiple video bitstreams, aiming to achieve better quality of the reconstructed video. The experimental results by joint decoding of multiple H.263 streams show that the proposed algorithm can significantly enhance the decoded video quality. Zhiqin Liang, Jiantao Zhou 0001, Mengyao Ma, Oscar C. Au |
ICME | 5 |
| 2007 | Enhanced Image Trans-coding Using Reversible Data HidingabstractThe primary application of watermarking and data hiding is for authentication, or to prove the ownership of digital media. In this paper, a new trans-coding system with the help of the technique of watermarking and data hiding is proposed. Side information is extracted before the image transcoding, such as resizing. In this paper, we focus on the problem of resizing in "thin edge" region. "Thin edge" structure normally cannot preserve after resizing process and become discrete. As the "thin edge" region is difficult to analyze real time and with a high degree of accuracy, data hiding can be used. The information of the "thin edge" region is generated and embedding into the multimedia content in encoder. Experimental results shown that there is a great improvement in the visual quality by using side information. Richard Y. M. Li, Oscar C. Au, Carman K. M. Yuk, Shu-Kei Yip, Tai-Wai Chan |
ISCAS | 2 |
| 2007 | Color Demosaicking Using Direction Similarity in Color Difference SpacesabstractThis paper proposes a color demosaicking algorithm which adaptively selects direction for interpolation based on similarity of the directions in the color difference spaces. As color differences are usually smooth along the edge, the paper proposed to estimate missing channels by selecting a direction which has higher score in similarity measurement. The complexity of our proposed method is low that only addition, subtraction and shifting are required. Experimental results show that our proposed algorithm has relative high PSNR and improvement in the visual quality of reconstructed images comparing to that of state-of-art demosaicking algorithms Carman K. M. Yuk, Oscar C. Au, Richard Y. M. Li, Sui-Yuk Lam |
ISCAS | 2 |
| 2007 | Soft-Decision Color Demosaicking with Direction Vector SelectionabstractWe propose a soft-decision color demosaicking algorithm with direction vector selection which effectively minimizing color artifacts. Since our interpolation uses soft decision and decision making bases on direction vectors which consists of three primary colors together with same direction, it not only maintains the direction consistency, but also significantly reduces color artifacts by largely avoiding interpolation across the edge. Experimental results show that our proposed algorithm outperforms state-of-art methods and the visual quality of reconstructed images is also obviously improved. Carman K. M. Yuk, Oscar C. Au, Richard Y. M. Li, Sui-Yuk Lam |
MMSP | 2 |
| 2007 | Security Analysis of Multimedia Encryption Schemes Based on Multiple Huffman TableabstractThis letter addresses the security issues of the multimedia encryption schemes using multiple Huffman table (MHT). A known-plaintext attack is presented to show that the MHTs used for encryption should be carefully selected to avoid the weak keys problem. We then propose chosen-plaintext attacks on the basic MHT algorithm as well as the enhanced scheme with random bit insertion. In addition, we suggest two empirical criteria for Huffman table selection, based on which we can simplify the stream cipher integrated scheme, while ensuring a high level of security. Jiantao Zhou 0001, Zhiqin Liang, Yan Chen 0007, Oscar C. Au |
IEEE Signal Process. Lett. | 4 |
| 2007 | Temporal Video Denoising Based on Multihypothesis Motion CompensationabstractDenoising module is required by any practical video processing systems. Most existing denoising schemes are spatio-temporal filters which operate on data over three dimensions. However, to limit the number of inputs, these filters only utilize one reference frame and cannot fully exploit temporal correlation. In this paper, a recursive temporal denoising filter named multihypothesis motion compensated filter (MHMCF) is proposed. To fully exploit temporal correlation, MHMCF performs motion estimation in a number of reference frames to construct multiple hypotheses (temporal predictions) of the current pixel. These hypotheses are combined by weighted averaging to suppress noise and estimate the actual current pixel value. Based on the multihypothesis motion compensated residue model presented in this paper, we investigate the efficiency of MHMCF, and some numerical evaluations are revealed. Experimental results show that MHMCF demonstrates quite good denoising performance while the inputs are much fewer than spatio-temporal filters. Moreover, as a purely temporal filter, it can well preserve spatial details and achieve satisfactory visual quality. Oscar C. Au, Mengyao Ma, Zhiqin Liang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Sketch-Guided Texture-Based Image InpaintingabstractIn this paper, we propose a novel framework for image inpainting, named sketch-guided texture-based image inpainting. Inspired by the well-known primal sketch model, we present a penetrating perspective into the process of image formation, where each image is seen as a variety of texture organized by some underlying structure. Based on this conceptual foundation, our approach of image inpainting integrates two unified stages: it first reconstructs the image structure with the sketch model, and then guided by the structure, it restores the missing region by patch-based texture synthesis. The major superiority of the framework over other ones consists in its capability of simultaneously recovering the structure and texture in the missing regions. Comprehensive experiments are performed to compare our method with other state-of- the-art ones; the encouraging results obtained convincingly demonstrate the effectiveness of our method. Yan Chen 0007, Qing Luan, Houqiang Li, Oscar C. Au |
ICIP | 4 |
| 2006 | A Multihypothesis Motion-Compensated Temporal Filter for Video DenoisingabstractMost existing filters for video denoising are spatio-temporal filters which operate on data over 3 dimensions. This paper presents a purely linear temporal filter with multihypothesis motion compensation (MC). Compared to the spatio-temporal filters, the proposed one needs much fewer inputs and much simpler operation. Experimental results show that just by involving 3 pixels, the proposed filter can achieve very good noise suppression performance. Moreover, as a purely temporal filter, it can well preserve spatial details and achieve satisfactory visual quality. Oscar C. Au, Mengyao Ma, Zhiqin Liang, Carman K. M. Yuk |
ICIP | 2 |
| 2006 | Low-Complexity Rate Control for Efficient H.263 to H.264/AVC Video TranscodingabstractRate control is a complicated problem in the H.264/AVC coding standard, extra computation is usually needed for the existing rate control schemes to estimate the complexity of frames or macroblocks (MBs). However, during transcoding, information from preceded video could be used to simplify the rate control. In this paper, we propose a low-complexity rate control scheme for transcoding from H.263 to H.264/AVC. The relationship between the rate of the pre-coded video and both the rate and distortion of the transcoded video are studied. By using only the rate information from the preceded video, we introduce a row-layer bit allocation and perform average rate shaping across a row of MBs. Estimation error diffusion is also introduced. The proposed scheme has sufficiently lower computational complexity than other methods as there is no explicit complexity measurement of MBs and complicated parameters updating. Experimental results show the effectiveness of the proposed scheme. Chi-Wang Ho, Oscar C. Au, Shueng-Han Gary Chan, Shu-Kei Yip, Hoi-Ming Wong |
ICIP | 2 |
| 2006 | New Digital Watermarking for Few-Color ImagesabstractIn this paper, we propose a new digital watermarking algorithm named spatial unified key insertion (SUKI) for few-color images. By performing the adaptive threshold halftoning and contour shaping and modification, we can embed a binary logo into a digital image with superior JPEG compression resistance ability. The bit error rate (BER) of JPEG attack is equal to zero for all testing images with quality factor≥80. The BER is much lower than the traditional watermarking algorithms, such as spread spectrum watermarking and quantization index modulation watermarking. The payload is considerably high and there are no "salt-and-peppers" artifacts in the watermarked images. No additional color is introduced and the palette keeps unchanged after watermark embedding. The proposed algorithm can be applied to natural images with distinct boundaries by segmenting the images into different parts. Shu-Kei Yip, Oscar C. Au, Chi-Wang Ho, Hoi-Ming Wong, Richard Y. M. Li |
ICIP | 2 |
| 2006 | Adaptively Switching Between Directional Interpolation and Region Matching for Spatial Error Concealment Based on DCT CoefficientsabstractIn this paper, a novel spatial error concealment algorithm, which adaptively switches between directional interpolation and region matching, is proposed. Different from the previous spatial error concealment methods, which just utilize smooth property, the algorithm exploits both smooth property and texture information to recover the lost blocks. Based on the DCT coefficients in the available neighboring MBs, the algorithm automatically analyzes whether the MB is "smooth-like" or "texture-like" and adaptively select directional interpolation or region matching to recover the lost MB. The proposed algorithm has been evaluated on H.264 reference software JM 9.0. The experimental results demonstrate that the proposed method can achieve better PSNR performance and visual quality, compared with weighted pixel average (WPA) which is adopted in H.264, directional interpolation-only and region matching-only. Yan Chen 0007, Oscar C. Au, Jiantao Zhou 0001, Chi-Wang Ho |
ICME | 2 |
| 2006 | An Encoder-Embedded Video Denoising Filter Based on the Temporal LMMSE EstimatorabstractNoise not only degrades the visual quality of video contents, but also significantly affects the coding efficiency. Based on the temporal linear minimum mean square error (LMMSE) estimator, an innovative denoising filter is proposed in this paper. The proposed filter only requires simple operations manipulating on the individual residue coefficients and can be seamlessly integrated into video encoders. Compared to traditional filter-encoder cascaded scheme, embedding the proposed filter into the video encoder can save a large amount of computation. The experimental results show that with the proposed filter embedded, both the noise suppression capability and the coding efficiency of the video encoder can be dramatically improved. Furthermore, as a purely temporal filter, it can well preserve the fine details of video contents and satisfactory visual quality can be achieved Oscar C. Au, Mengyao Ma, Zhiqin Liang |
ICME | 2 |
| 2006 | Motion Estimation for H.264/AVC using Programmable Graphics HardwareabstractWe present an efficient implementation of motion estimation (ME) for H.264/AVC using programmable graphics hardware. The cost function for ME in H.264/AVC depends on the motion vector (MV) predictor which is the median MV of three neighboring coded blocks. Previous implementations assume no dependency among adjacent blocks, which is not true for H.264/AVC, they also perform unsatisfactorily because of their low arithmetic intensity, which is defined as operation per word transferred. To overcome the dependency problem, we introduce a new implementation which performs ME on block-by-block basis. Moreover, we can adjust the arithmetic intensity easily to optimize the performance on different graphics cards. Experimental results show that our implementation is substantially faster (by 10 times) than our SIMD optimized CPU implementation Chi-Wang Ho, Oscar C. Au, Shueng-Han Gary Chan, Shu-Kei Yip, Hoi-Ming Wong |
ICME | 2 |
| 2006 | COSMOS: Peer-to-Peer Collaborative Streaming Among MobilesabstractIn traditional mobile streaming networks such as 3G cellular networks, all users pull streams from a server. Such pull model leads to high streaming cost and problem in system scalability. In this paper, we propose and investigate a scalable and cost-effective protocol to distribute multimedia content to mobiles in a peer-to-peer manner. Our protocol, termed collaborative streaming among mobiles (COSMOS), makes use of multiple description coding (MDC) and data sharing to achieve high performance. In COSMOS, only a few peers pull video descriptions through a telecommunication channel. Using a free broadcast channel (such as Wi-Fi and Bluetooth), they share the descriptions to nearby neighbors in an ad-hoc manner. This way reduces greatly the telecommunication cost and cellular bandwidth requirement. As video descriptions are supplied by multiple peers, COSMOS is robust to peer failure. Since broadcasting is used to distribute video data, the protocol is highly scalable to large number of users. By taking turns to pull descriptions, we show through simulation that peers can effectively share, and hence substantially reduce, streaming cost. As peers can often obtain a number of descriptions from nearby neighbors, they enjoy lower delay as compared to a recent scheme CHUM Man-Fung Leung, Shueng-Han Gary Chan, Oscar C. Au |
ICME | 3 |
| 2006 | Lossless Visible WatermarkingabstractThe embedding distortion of visible watermarking is usually larger than that of invisible watermarking. In order to maintain the signal fidelity after the watermark extraction, "lossless" property is highlighted in the visible watermarking. In this paper, we propose two lossless visible watermarking algorithms, pixel value matching algorithm (PVMA) and pixel position shift algorithm (PPSA). PVMA uses the bijective intensity mapping function to watermark a visible logo whereas PPSA uses circular pixel shift to improve the visibility of the watermark in the high variance region. For the application of medical and military, as they are sensitive to distortion, PVMA and PPSA can be used to insert a visible logo to prevent unauthorized use Shu-Kei Yip, Oscar C. Au, Chi-Wang Ho, Hoi-Ming Wong |
ICME | 2 |
| 2006 | Block-Based Lossless Data Hiding in Delta DomainabstractDigital watermarking is one of the ways to prove the ownership and the authenticity of the media. However, some applications, such as medical and military, are sensitive to distortion, this highlights the needs of lossless watermarking. In this paper, we propose a new lossless data hiding algorithm in delta domain. A MSE discount is obtained by using checkerboard-pattern watermark sequences. The PSNR between the watermarked image and the original image is high and there is no "salt-and-peppers" artifact. The proposed algorithm can be extended to withstand the JPEG attack Shu-Kei Yip, Oscar C. Au, Hoi-Ming Wong, Chi-Wang Ho |
ICME | 2 |
| 2006 | On the Security of Multimedia Encryption Schemes Based on Multiple Huffman Table (MHT)abstractThis paper addresses the security issues of the multimedia encryption schemes based on multiple Huffman table (MHT). A detailed analysis of known-plaintext attack is presented to show that the Huffman tables used for encryption should be carefully selected to avoid the weak keys problem. Further, we propose an efficient chosen-plaintext attack on the basic MHT method as well as the enhanced scheme inserting random bits. We also show that random rotation in partitioned bit stream cannot essentially improve the security. Jiantao Zhou 0001, Zhiqin Liang, Yan Chen 0007, Oscar C. Au |
ICME | 4 |
| 2006 | Spatio-temporal boundary matching algorithm for temporal error concealmentabstractIn this paper, a novel temporal error concealment algorithm, called spatio-temporal boundary matching algorithm (STBMA), is proposed to recover the information lost in the video transmission. Different from the classical boundary matching algorithm (BMA), which just considers the spatial smoothness property, the proposed algorithm introduces a new distortion function to exploit both the spatial and temporal smoothness properties to recover the lost motion vector (MV) from candidates. The new distortion function involves two terms: spatial distortion term and temporal distortion term. Since both the spatial and temporal smoothness properties are involved, the proposed method can better minimize the distortion of the recovered block and recover more accurate MV. The proposed algorithm has been tested on H.264 reference software JM 9.0. The experimental results demonstrate the proposed algorithm can obtain better PSNR performance and visual quality, compared with BMA which is adopted in H.264. Yan Chen 0007, Oscar C. Au, Chi-Wang Ho, Jiantao Zhou 0001 |
ISCAS | 2 |
| 2006 | Improved refinement search for H.263 to H.264/AVC transcoding based on the minimum cost tendency searchabstractAn improved refinement search method for transcoding from H.263 to H.264/AVC is proposed in this paper. Many existing motion re-estimation methods refine the input motion vector (MV) with a small search range, which is usually input MV biased. Motion estimation (ME) in H.263 usually does not consider the rate required for coding the MV, and hence, the input MV may incur a large cost in H.264/AVC. To overcome this problem, we introduce a refinement search method, called minimum cost tendency search (MCTS), which takes the difference between the cost functions for ME in H.263 and H.264/AVC into consideration. The input MV and the predictor MV are used as two anchor points. The proposed MCTS starts searching from the anchor point with a higher cost to another. Finally, the best point is chosen as the center for further refinement. The performance of MCTS is evaluated by comparing with full search, FME in JM software and refinement scheme using small diamond pattern around the input MV (RSD). Experimental results show the proposed MCTS performs more stable than FME and RSD over a wide range of output video quality. Chi-Wang Ho, Oscar C. Au, Shueng-Han Gary Chan, Hoi-Ming Wong, Shu-Kei Yip |
ISCAS | 2 |
| 2006 | Three-loop temporal interpolation for error concealment of MDCabstractMultiple description coding (MDC) can be used as an error resilience (ER) technique for video coding. In case of transmission errors, error concealment can be combined with MDC to reconstruct the lost frame, such that the propagated error to the following frames is reduced. In this paper, we propose a new temporal error concealment method named three-loop temporal interpolation (TLTI). TLTI can be well combined with temporal sub-sampling ER methods, such as MDC and alternative motion-compensated prediction. In the simulation, we compare the performance of TLTI with unidirectional motion compensated temporal interpolation (UMCTI). Both visual and quantitive results show that TLTI can achieve a better video quality than UMCTI. Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan, Zhiqin Liang |
ISCAS | 2 |
| 2006 | Fast mode decision and motion estimation for H.264 (FMDME)abstractThe H.264 video coding standard achieves highest coding gain. In addition to common tools in previous standards such as H.263 and MPEG4, it supports multiple frames, multiple block sizes, and 1/4 pixel motion estimation, deblocking filter, intra-prediction, and so on. However, full search of for all block modes and reference frames requires heavy computational complexity. The algorithm predictive motion vector field adaptive search technique (PMVFAST) (Tourapis, 2000) was previously accepted into MPEG standard to achieve hundreds of time of speed up. FMBME (Chang, 2004), and some other algorithms (Yanfei Shen, 2004) have been developed to do fast multiple block size motion estimation. For multiple reference frame motion estimation, FMFME (Chang, 2003), and some other algorithms (Xiang Li, 2004) have also been developed. In this paper a new algorithm is presented named as fast mode decision and motion estimation (FMDME). This algorithm achieves high speed up factor by involving fast skip mode checking, fast multiple block size motion estimation, fast multiple reference frame selection, fast motion estimation, and fast intra-block skipping for P frame in H.264. Hoi-Ming Wong, Oscar C. Au, Andy Chang, Shu-Kei Yip, Chi-Wang Ho |
ISCAS | 2 |
| 2006 | Generalized lossless data hiding by multiple predictorsabstractDigital watermarking is to prove the ownership and the authenticity of the media. However, as some applications, such as medical and military, are sensitive to distortion, this highlights the needs of lossless watermarking. In this paper, we propose a new lossless data hiding algorithm by using multiple predictors, which extends and generalizes our previous watermarking idea (Yip, 2005). By using different predictors with different characteristics, we can choose the embedding location to be low variance region or high variance region. The PSNR and the payload capacity are high and there are no "salt-and-peppers" artifacts. Shu-Kei Yip, Oscar C. Au, Hoi-Ming Wong, Chi-Wang Ho |
ISCAS | 2 |
| 2006 | Content-adaptive Temporal Search Range Control Based on Frame Buffer UtilizationabstractMultiple reference frame selection adopted by the state-of-art H.264 video compression standard offers substantial performance gain. The temporal search range control, as a consequence, is crucial for maintaining the coding performance with minimum complexity. In this paper, we investigate the relationships between the reference frame buffer utilization and the optimal search range. A content-adaptive algorithm is proposed to control the search range dynamically during the encoding process. Experimental results show that our algorithm can rapidly adapt to the video characteristics and effectively reduce the complexity with negligible coding performance penalty. Zhiqin Liang, Jiantao Zhou 0001, Oscar C. Au |
MMSP | 3 |
| 2005 | A novel content-adaptive video denoising filterabstractWe propose a simple non-linear content-adaptive filter that is efficient in removing noise from a video. The proposed filter is called spatiotemporal varying filter (STVF) and is able to produce optimal results in the sense that it minimizes the weighted least square error. STVF combines the advantages of conventional denoising filters that enable it to decrease the noise variance in smooth areas but at the same time retains the sharpness of edges in object boundaries. Simulation results show that STVF outperforms the conventional denoising methods like low-pass filtering, median filtering and Wiener filtering. Tai-Wai Chan, Oscar C. Au, Tak-Song Chong, Wing-San Chau |
ICASSP (2) | 2 |
| 2005 | Sub-optimal quarter-pixel inter-prediction algorithm (SQIA)abstractMotion estimation (ME) is an important part of modern video coding systems to exploit temporary redundancy in a video. Motion estimation is typically per-formed firstly with integer-pixel accuracy and then at sub-pixel accuracy, which includes half-pixel and quarter-pixel accuracy. When sophisticated fast integer-pixel accuracy motion-search algorithms are used to decrease the number of search points for integer-pixel motion search, quarter-pixel motion search becomes another important processing bottleneck in the encoding process. The conventional method is to search 8 half-pixel positions around the motion vector (MV) obtained from integer-pixel motion search, then do motion search in the same way on 8 quarter-pixel positions around the MV obtained from the half-pixel motion search, therefore, in total, 16 search points are needed. The proposed algorithm, named sub-optimal quarter-pixel inter-prediction algorithm (SQIA), successfully optimizes the quarter-pixel motion search part and improves the processing speed with low PSNR penalty. Hoi-Ming Wong, Oscar C. Au, Jinxin Huang, Shiju Zhang, Winnie N. Yan |
ICASSP (2) | 2 |
| 2005 | Piecewise Linear Model for Real-Time Rate ControlabstractMost existing rate control schemes in the literature evaluate quantization parameters of macro-blocks (MB) based on models which estimate well at low bit rates (e.g. quadratic model). These models are not flexible because this model does not estimate well for different situations with different rates. This work investigates the relationships of the rate and the distortion of a residue MB with two parameters - MB quantization step size and standard deviation of MB prediction errors. A piecewise linear rate model and linear distortion model are proposed. Then, we present the proposed rate control scheme based on these models. The experimental results suggest that our scheme achieves PSNR gain over TMN8 and Li's scheme. Chi Wah Wong, Oscar C. Au, Raymond Chi-Wing Wong, Hong-Kwai Lam |
ICASSP (2) | 2 |
| 2005 | A new motion compensation approach for error resilient video codingabstractMultihypothesis motion-compensated prediction (MHMCP) can be used as an error resilience technique for video coding. Motivated by MHMCP, we propose a new error resilience approach named alternative motion-compensated prediction (AMCP), where two-hypothesis and one-hypothesis predictions are alternatively used with some mechanism. Both theory and simulation results show that in case of one frame loss, the expected converged error using AMCP is smaller than that using two-hypothesis MCP. Mengyao Ma, Oscar C. Au, Shueng-Han Gary Chan |
ICIP (1) | 2 |
| 2005 | Content adaptive watermarking using a 2-stage predictorabstractDigital watermarking is one of the solutions to protect intellectual properties and copyright by hiding information, such as a random sequence or a logo, into digital media. In this paper, a new watermarking scheme is proposed. The embedding process takes place in the spatial domain. A bi-level logo is embedded into digital media by comparing the absolute difference between the original pixel value and the predicted pixel value, which is computed and predicted by using the concept of activity measurement, followed by spatial varying filter (SVF). The logo will be extracted from a possibly corrupted image, without the help of original uncorrupted image. The new proposed algorithm can withstand the geometric attacks as well as common signal processing, and has a high payload capacity. Shu-Kei Yip, Oscar C. Au, Chi-Wang Ho, Hoi-Ming Wong |
ICIP (1) | 2 |
| 2005 | A spatial-temporal de-interlacing algorithmabstractIn this paper, we proposed a spatial-temporal de-interlacing algorithm for conversion of interlaced video to progressive video. Our proposed algorithm estimates the motion trajectory of three consecutive fields interpolates the missing field along the motion trajectory. In the motion estimator, the unidirectional motion estimation and the bidirectional motion estimation processes are combined by multiple objective minimization technique. The unidirectional motion estimation estimates the motion trajectory by comparing the blocks from opposite parity fields while the bi-directional motion estimation compares blocks from the same parity fields. By combining the two motion estimations, the motion trajectory can be accurately predicted. In addition, a quality analyzer is proposed to evaluate the visual quality of the reconstructed frame, which chooses the appropriate interpolation scheme in order to provide maximum de-interlacing performance. Simulation results show the proposed algorithm has better performance over existing de-interlacing algorithm. Tak-Song Chong, Oscar C. Au, Tai-Wai Chan, Wing-San Chau |
ICME | 2 |
| 2005 | Multiple Objective Frame Rate Up ConversionabstractIn this paper, we propose a multiple objective frame rate up conversion algorithm (MOFRUC), which utilizes two different models. The first model is a constant velocity model that assumes the objects position is a linear function of time. The second model exploits the spatial correlation between neighboring blocks, and assumes the pixel intensity is highly correlated in a small local area. In this model, the perceptual quality of interpolated frame is also taken into account and the blocking artifact is minimized. Our proposed MOFRUC estimates the motion trajectory by the first model and interpolate the frame along the motion trajectory. At the same time, the algorithm refines the motion trajectory by maximizing a spatial correlation measurement defined in the second model and interpolates the frame with minimum blocking artifact. Simulation results show that our proposed MOFRUC outperforms other existing algorithms and produces high quality interpolated frame. Tak-Song Chong, Oscar C. Au, Wing-San Chau, Tai-Wai Chan |
ICME | 2 |
| 2005 | Enhanced predictive motion vector field adaptive search technique (E-PMVFAST)-based on future MV predictionabstractMotion estimation (ME) is a core part of most modern video coding standard, and it directly affects the compression efficiency and visual quality of a video. If full search (FS) algorithm is used, ME could takes over 70% of computational power. Many algorithms such as TSS and PMVFAST, have been developed to achieve great speed up for ME. In this paper, a new algorithm enhanced-PMVFAST (E-PMVFAST) is proposed, which performs better than most if not all other existing algorithms in terms of speed up factor while keeping the similar PSNR with FS. Our experiments have also verified the robustness of the proposed E-PMVFAST algorithm. Hoi-Ming Wong, Oscar C. Au, Chi-Wang Ho, Shu-Kei Yip |
ICME | 2 |
| 2005 | Rate Control Based on Zero-Residue Pre-Selection for Video TranscodingabstractA common issue in video transcoding for heterogeneous network environment is to efficiently and accurately reduce the bit-rate such that the distortion is minimized under a given rate constraint. To convert the bit-rate of an encoded video to match the channel capacity, in general, re-quantization is done on the DCT coefficients with larger quantization step size. Most existing rate control algorithms for video transcoding in the literature calculate quantization parameters (QPs) of macroblocks (MBs) based on a relationship between certain properties of coded video and bit-rate. They reduce the computational complexity by simplifying the R-D model and reusing the statistics information of input video. In this paper, we propose a zero-residue pre-selection (ZRPS) mechanism to select only a portion of MBs to apply the rate control in video transcoding. TMN-8 is used to evaluate the impact of ZRPS. Experimental results show that, as compared to the original TMN-8 rate control scheme, TMN-8 with ZRPS achieves up to 1.60 dB gain, in term of PSNR, and requires less than 50% of the computational complexity compared to TMN-8, depending on the characteristics of the video content Chi-Wang Ho, Oscar C. Au, Shueng-Han Gary Chan, Hoi-Ming Wong, Shu-Kei Yip |
MMSP | 2 |
| 2005 | Efficient rate control for JPEG2000 image codingabstractJPEG2000 is the new image coding standard which can provide superior rate-distortion performance over the old JPEG standard. However, the conventional post-compression rate-distortion (PCRD) optimization scheme in JPEG2000 is not efficient. It requires entropy encoding all available data even though a large portion of them will not be included in the final output. In this paper, three rate control methods are proposed to efficiently reduce both the computational complexity and memory usage over the conventional PCRD method. The first method, called successive bit-plane rate allocation (SBRA), allocates the bit rate by using the currently available rate-distortion information only. The second method, called priority scanning rate allocation (PSRA), allocates bits according to certain prioritized ordering. The third method uses PSRA to achieve optimal truncation as PCRD without encoding of all the image details and is called priority scanning with optimal truncation (PSOT). Simulation results suggest that the three proposed methods provide different tradeoff among visual quality, computational complexity, coding delay and working memory size. SBRA is memoryless and causal and requires the least computational complexity, lowest coding delay and achieves good visual quality. PSRA achieves higher PSNR than SBRA at the expense of larger working memory size and longer coding delay. PSOT gives the best PSNR but requires even more computation, delay and memory. Yick Ming Yeung, Oscar C. Au |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2004 | Watermarking technique for color halftone imagesabstractIn this paper, we propose a method to hide invisible patterns in color error diffused halftone images. The hidden pattern is embedded in different color components. The hidden patterns would be revealed when the watermarked color halftone images are under Boolean operation or overlaid. Simulation results show that the watermarked color halftone images have good visual quality and the hidden pattern is visible clearly. Ming Sun Fu, Oscar C. Au |
ICASSP (3) | 2 |
| 2004 | Automatic white balancing using luminance component and standard deviation of RGB components [image preprocessing]abstractAutomatic white balancing is an essential image preprocessing component in consumer digital still cameras, and it can greatly improve the final image quality of the captured image. In this paper, a novel automatic white balancing algorithm based on both the luminance component and standard deviation of RGB components of the pre-captured image is proposed. A light source model for evaluation of an automatic white balancing is also described. The simulation results indicate that the proposed algorithm can improve the final image quality of the captured image. Hong-Kwai Lam, Oscar C. Au, Chi Wah Wong |
ICASSP (3) | 2 |
| 2004 | A sequential multiple watermarks embedding techniqueabstractIn this paper, we propose a novel multiple watermarks embedding scheme. We assume M watermarks have already embedded in the image using M sets of secret key. With the availability of these M sets secret key, another N watermarks can be embedded using the proposed technique while the energies of the watermarks are minimized. And the watermarks embedded later will not interfere with the first M watermarks. Experimental results show watermarked images have good visual quality and the watermarks are robust to JPEG compression and noise attacks. Peter Hon-Wah Wong, Andy Chang, Oscar C. Au |
ICASSP (3) | 3 |
| 2004 | Fast integer motion estimation for H.264 video coding standardabstractMultiple block size motion estimation is adopted in the latest JVT/H.264 video coding standard to achieve higher coding efficiency. However, full exhaustive search of all block sizes is computational intensive with motion estimation complexity increasing linearly with the number of block sizes allowed. We propose a fast integer motion estimation method for multiple block size motion estimation (FMBME) which reduces computation significantly. Experimental result shows that, compared with full search, the proposed method can have a speed up factor of four with bit-rate increases within 1%. Andy Chang, Peter Hon-Wah Wong, Yick Ming Yeung, Oscar C. Au |
ICME | 4 |
| 2004 | Key frame selection by macroblock type and motion vector analysisabstractKey frame selection is the process to select the most representative frames of a long video sequence. The paper describes a new algorithm for determining key frames in a video shot. We use macroblock (MB) type and motion vector information of MPEG compressed video to determine the local maxima of visual content change in a shot, i.e. the frames that are most noticeable to the viewer. This algorithm exploits the important cue of the intensity of motion activity and visual content change within a shot from the corresponding MB type and motion vector characteristics. The advantages of this algorithm are the efficient extraction of the required information in the MPEG compressed domain with low complexity VLC decoding and high accuracy. Results show that the proposed algorithm can successfully select the key frames that are sufficient to represent a shot. Wing-San Chau, Oscar C. Au, Tak-Song Chong |
ICME | 2 |
| 2004 | Temporal error concealment for video transmissionabstractWe propose a temporal error concealment algorithm for video transmission in an error-prone environment. The error concealment algorithm employed an edge detection algorithm and progressive median motion vector concealment. First, the edges are detected and concealed portion by portion. Then, the corrupted MB is partitioned by the reconstructed edges and each partition is concealed by progressive median motion vector individually. The proposed algorithm shows better performance on both objective and subjective quality over the existing temporal error concealment algorithm. Tak-Song Chong, Oscar C. Au, Wing-San Chau, Tai-Wai Chan |
ICME | 2 |
| 2004 | Joint visual cryptography and watermarkingabstractIn this paper, we discuss how to use the watermarking technique for visual cryptography. Both halftone watermarking and visual cryptography involve a hidden secret image. However, their concepts are different. For visual cryptography, a set of shared binary images is used to protect the content of the hidden image. The hidden image can only be revealed when enough shared images are obtained. For watermarking, the hidden image is usually embedded in a single halftone image while preserving the quality of the watermarked halftone image. In this paper, we propose a joint visual-cryptography and watermarking (JVW) algorithm that has the merits of both visual cryptography and watermarking Ming Sun Fu, Oscar C. Au |
ICME | 2 |
| 2004 | Fast motion vector re-estimation for arbitrary video downsizing using spatial-variant filterabstractTo downsize a compressed video to a lower resolution, the straightforward method is to decompress the compressed video, downsample the video in spatial domain and then re-compress the downsampled video. This process can be computationally expensive unless fast motion vector re-estimation methods are used. The use of a spatial-variant filter (SVF) as a fast motion vector re-estimation method for arbitrary downsizing factor is investigated. Simulation results suggest that the SVF can improve the quality of the existing fast algorithms in terms of PSNR. With a little bit of local search, its PSNR can be close to the full search. Hong-Kwai Lam, Oscar C. Au, Chi Wah Wong |
ICME | 2 |
| 2004 | Automatic white balancing using adjacent channels adjustment in RGB domainabstractAutomatic white balancing is one of the key essential image pre-processing components in consumer digital still cameras, and it can highly improve the final image quality of the captured image. A novel automatic white balancing algorithm based on the adjacent channels adjustment in the RGB domain using both luminance information and standard deviation of RGB components of the precaptured image is proposed. A light source model for evaluating different automatic white balancing methods is also described. The simulation results show that the proposed method can improve the final image quality of the captured image. Hong-Kwai Lam, Oscar C. Au, Chi Wah Wong |
ICME | 2 |
| 2004 | PID-based real-time rate controlabstractMany existing rate control schemes consider spatial quality only. As a result, this may introduce a large distortion variation over frames to which people are sensitive. In fact, people are also sensitive to temporal quality. We design a real-time rate control based on a PID controller to have better tradeoff between spatial and temporal quality. From different estimated number of bits per frame and buffer status, different target bits for each frame are used in order to reduce the flickering effect (smaller distortion variation) which is one factor affecting temporal quality. Experimental results suggest that our scheme can obtain more consistent quality while keeping high spatial quality. Chi Wah Wong, Oscar C. Au, Hong-Kwai Lam |
ICME | 2 |
| 2004 | Rate control using probability of non-zero quantized coefficientsabstractMost existing rate control schemes in the literature calculate quantization parameters of the macro-blocks (MB) based on the standard deviation of the residue before quantization. The probability of non-zero coefficients after quantization is a new factor that has an interesting property. The rate tends to be linear with this probability. In this work, we design the rate control based on this special factor with its linear property. This work investigates the relationships of bit rate and distortion with MB quantization step size and its interesting probability. We establish linear rate and linear distortion models and propose a rate control scheme based on these models. The proposed algorithm has a low quantization overhead, low computation complexity and is independent of the distortion factors, compared with TMN8 rate control. The experimental results suggest that our scheme can achieve PSNR gain over TMN8. Chi Wah Wong, Oscar C. Au, Hong-Kwai Lam |
ICME | 2 |
| 2004 | Arbitrarily-shaped video coding: smart padding versus MPEG-4 LPE/zero paddingabstractAn effective padding scheme, called the smart padding (SmartPad), has been developed recently for the DCT coding of arbitrarily-shaped image/video objects; whereas its superior performance over the MPEG-4 LPE padding has been confirmed solidly. In the present paper, we propose to extend the use of SmartPad to all INTER frames (of arbitrary shapes), i.e., to use SmartPad to replace the zero padding scheme (as recommended in MPEG-4). Our simulation results show that a very substantial performance gain (3-7 dB) has been achieved, as compared to the MPEG-4 LPE/zero padding scheme. A. C. Yu, Guobin Shen, Bing Zeng 0001, Oscar C. Au |
ICME | 4 |
| 2003 | A symmetric key watermark for halftone imagesabstractIn many printer and publishing applications, it is desirable to embed data in halftone images for copyright control and authentication purposes. While intentional attacks on printed matters may not be likely, unintentional attacks such as cropping and distortion due to dirt or human writing/marking are likely. In this paper, we proposed a novel halftone image watermarking method called watermarking error diffusion (WED) to embed a watermark in the parity domain of halftone images with symmetric key during halftoning while introducing minimal distortion. Oscar C. Au, Ming Sun Fu |
ICASSP (3) | 1 |
| 2003 | A novel approach to fast multi-frame selection for H.264 video codingabstractThe latest under development video coding standard, H.264, uses multiple reference frames to improve the rate-distortion performance. However, the motion estimation process involved is computational intensive and increases linearly with the number of allowed reference frames. In this paper, a novel fast multi-frame selection method is proposed for H.264 video coding. The proposed scheme can efficiently reduce the computational cost up to 70% with similiar quality and bit-rate. So the method is highly suitable for real-time (or low delay) applications while the benefits from multi-frame motion compensation can be preserved. Andy Chang, Oscar C. Au, Yick Ming Yeung |
ICASSP (3) | 2 |
| 2003 | A novel method to embed watermark in different halftone images: data hiding by conjugate error diffusion (DHCED)abstractIn this paper, we propose a novel way called DHCED to hide invisible patterns in two or more visually different halftone images (e.g. Lena and Harbor) such that the hidden patterns would appear on the halftone images when they are overlaid. Conjugate error diffusion is used to embed the binary visual pattern in the two distinct halftone images. Simulation results show that the two halftone images have good visual quality, and the hidden pattern is visible when the two distinct halftone images are overlaid. Ming Sun Fu, Oscar C. Au |
ICASSP (3) | 2 |
| 2003 | Fast intra-prediction mode selection for 4A blocks in H.264abstractIn the upcoming H.264 advanced video coding standard, intra-prediction for each 4/spl times/4 block is used to compress I-frames. However, the full search algorithm to choose one of the 9 prediction modes is computationally expensive. We propose a fast intra-prediction mode selection (FIPMS) method based on partial computation of the cost function, early termination and selective computation of highly probable modes. The proposed FIPMS can reduce the complexity considerably while maintaining similar PSNR and bit rate. Bojun Meng, Oscar C. Au |
ICASSP (3) | 2 |
| 2003 | Successive bit-plane rate allocation technique for JPEG2000 image codingabstractA novel rate control scheme using successive bit-plane rate allocation (SBRA) is proposed for JPEG2000 image coding. By using the current rate-distortion information only, the proposed method can achieve a quality close to the post-compression rate-distortion (PCRD) optimization scheme adopted in JPEG2000. The proposed scheme can efficiently reduce both the computational cost and working memory size of the entropy coding process up to about 90%, in the case of 0.25bpp (1/32) compression. Without using the future rate-distortion information, the sequential property of the proposed method is highly suitable for real-time (or low delay) applications and implementation. Yick Ming Yeung, Oscar C. Au, Andy Chang |
ICASSP (3) | 2 |
| 2003 | A novel color interpolation framework in modified YCbCr domain for digital camerasabstractMost digital cameras use single-CCD with color filter array to capture sub-sampled digital color images and need color interpolation to generate full resolution color details. While most color interpolation algorithms are performed in the RGB domain, we propose a new framework to perform the color interpolation in the YCbCr domain. Simulation results suggest that the proposed framework can produce significantly better quality color images compared with existing algorithms. Wing Cheong Chan, Oscar C. Au, Ming Fai Fu |
ICIP (2) | 2 |
| 2003 | A set of mutually watermarked halftone imagesabstractIn this paper, we propose a method called, modified stochastic error diffusion (MSED), to hide an invisible binary visual pattern in a set of error diffused halftone images. MSED can embed a dithered watermark in multiple halftone images while preserving the visual quality of each halftone image. When any of these images overlapping with any others, the watermark will appear. In case that many of these image overlapped together, the contrast of watermark becomes clearer. Ming Sun Fu, Oscar C. Au |
ICIP (2) | 2 |
| 2003 | Fast global motion estimation based on local motion segmentationabstractIn MPEG-4, global motion estimation (GME) and global motion compensation can improve the visual quality of local motion estimation/compensation. However, the computational complexity of GME is very high. Even the existing fast algorithm, FFRGMET, in the MPEG-4 OM3 software is not fast enough. In this paper, we propose a fast GME algorithm called APSGME. It uses motion vectors from the PMVFAST local motion estimation to perform early termination, foreground/background block-based segmentation, feature point selection. Experimental results shown that APSGME can achieve considerable speed-up over FFRGMET while improving the visual quality simultaneously. Ming Fai Fu, Oscar C. Au, Wing Cheong Chan |
ICIP (2) | 2 |
| 2003 | Efficient intra-prediction algorithm in H.264abstractIn the upcoming H.264, intra-prediction with block sizes of 4/spl times/4 and 16x16 is used to compress I-frame. However, there are 9 or 4 candidate modes for the 4x4 or 16x16 intra-prediction. The full search algorithm to choose the modes is computationally expensive. In this paper, we suggest an efficient intra-prediction (EIP) algorithm based on early termination, selective computation of highly probable modes, and partial computation of the cost function. The proposed EIP reduces the complexity considerably while maintaining similar PSNR and bit rate. Bojun Meng, Oscar C. Au, Chi Wah Wong, Hong-Kwai Lam |
ICIP (3) | 2 |
| 2003 | An efficient optimal rate control scheme for JPEG2000 image codingabstractMost of the computation and memory usage of the post-compression rate-distortion (PCRD) optimization scheme in JPEG2000 are redundant. In this paper, an efficient rate PCRD scheme based on priority scanning (PS) is proposed to alleviate the problem. By encoding the truncation points in a different order based on the priority information, the proposed method can efficiently reduce the redundancy while keeping the same quality as the conventional PCRD scheme. The proposed scheme can efficiently reduce both the computational cost and working memory size of the entropy coding process by up to 52% and 71%, in the case of 0.25 bpp (1/32) compression, respectively. Yick Ming Yeung, Oscar C. Au, Andy Chang |
ICIP (3) | 2 |
| 2003 | A novel approach to fast multi-block motion estimation for H.264 video codingabstractThe upcoming video coding standard, H.264, uses motion estimation with multiple block sizes to improve the rate-distortion performance. However, full exhaustive search of all block sizes is computational intensive with complexity increasing linearly with the number of allowed block size. In this paper, a novel fast multi-block motion estimation (FMBME) is proposed for H.264 video coding. Experimental results show that the proposed FMBME can efficiently reduce the computational cost by 40.73% with similar visual quality and bit rate. Andy Chang, Oscar C. Au, Yick Ming Yeung |
ICME | 2 |
| 2003 | A multi-bit robust watermark for halftone imagesabstractIn many printer and publishing applications, it is desirable to embed data in halftone images for copyright control and authentication purposes. While intentional attacks on printed matters may not be likely, unintentional attacks such as cropping and distortion due to dirt or human writing/marking are likely. In this paper, we proposed a novel halftone image watermarking method to embed a robust, invisible, multi-bit watermark in the halftone images during halftoning while introducing minimal distortion. Ming Sun Fu, Oscar C. Au |
ICME | 2 |
| 2003 | A novel method to embed watermark in different halftone images: data hiding by conjugate error diffusion (DHCED)abstractIn this paper, we propose a novel way called DHCED to hide invisible patterns in two or more visually different halftone images (e.g. Lena and Harbor) such that the hidden patterns would appear on the halftone images when they are overlaid. Conjugate error diffusion is used to embed the binary visual pattern in the two distinct halftone images. Simulation results show that the two halftone images have good visual quality, and the hidden pattern is visible when the two distinct halftone images are overlaid. Ming Sun Fu, Oscar C. Au |
ICME | 2 |
| 2003 | Efficient intra-prediction mode selection for 4×4 blocks in H.264abstractIn the upcoming H.264, intra-prediction for 4/spl times/4 and 16/spl times/16 blocks is used to compress I-frame. However, we need to choose one of the 9 prediction modes for each 4/spl times/4 block. The full search (FS) algorithm to select the mode is computationally expensive. In this paper, an efficient intra-prediction mode selection (EIPMS) method is proposed based on partial computation of the cost function, selective computation of highly probable modes, early termination, and variable threshold setting. The proposed EIPMS can reduce the complexity considerably while maintaining similar PSNR and bit rate. Bojun Meng, Oscar C. Au, Chi Wah Wong, Hong-Kwai Lam |
ICME | 2 |
| 2003 | Perceptual rate control for low-delay video communicationsabstractThe existing rate control schemes in the literature evaluate their video quality performance in terms of traditional PSNR. Actually, this evaluation may not be good because this measurement is not the same as that judged by human. This work defines a perceptual quality measure based on the human vision system. The modified rate control includes the perceptual consideration. The results show that the perceptual rate control achieves average perceptual quality gain over TMN8. Chi Wah Wong, Oscar C. Au, Bojun Meng, Hong-Kwai Lam |
ICME | 2 |
| 2003 | Capacity for JPEG2000-to-JPEG2000 images watermarkingabstractA data capacity estimation method is proposed for image watermarking. The data capacity is maximum number of bits that can be embedded in an image such that the image has no perceptual loss. It is assumed that the input is a JPEG2000 image file and after watermark embedding, the image is JPEG2000-compressed using the same quantization factors. This is called JPEG2000-to-JPEG2000 (J2K2J2K) watermarking. A human visual systems (HVS) model is used to estimate the just noticeable difference (JND) of each discrete wavelet transform (DWT) coefficients. Peter Hon-Wah Wong, Yick Ming Yeung, Oscar C. Au |
ICME | 3 |
| 2003 | Efficient rate control technique for JPEG2000 image coding using priority scanningabstractMost of the computation and memory usage of the post-compression rate-distortion (PCRD) optimization scheme in JPEG2000 are redundant. In this paper, we propose a novel rate control scheme using priority scanning (PS) to address the problem. By encoding the truncation points in a different order based on the priority information, the proposed method can totally remove the redundancy while keeping a similar quality to the PCRD scheme. The proposed scheme can efficiently reduce both the computational cost and working memory size of the entropy coding process by up to about 90%, in the case of 0.25 bpp (1/32) compression. Yick Ming Yeung, Oscar C. Au, Andy Chang |
ICME | 2 |
| 2003 | Steganography in halftone images: conjugate error diffusion
Ming Sun Fu, Oscar C. Au |
Signal Process. | 2 |
| 2003 | A capacity estimation technique for JPEG-to-JPEG image watermarkingabstractIn JPEG-to-JPEG image watermarking (J2J), the input is a JPEG image file. After watermark embedding, the image is JPEG-compressed such that the output file is also a JPEG file. We use the human visual system (HVS) model to estimate the J2J data hiding capacity of JPEG images, or the maximum number of bits that can be embedded in JPEG-compressed images. A.B. Watson's HVS model (Proc. SPIE Human Vision, Visual Process., and Digital Display IV, p.202-16, 1993) is modified to estimate the just noticeable difference (JND) for DCT coefficients. The number of modifications to DCT coefficients is limited by JND in order to guarantee the invisibility of the watermark. Our capacity estimation method does not assume any specific watermarking method and thus would apply to any watermarking method in the J2J framework. Peter Hon-Wah Wong, Oscar C. Au |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2003 | Novel blind multiple watermarking technique for imagesabstractThree novel blind watermarking techniques are proposed to embed watermarks into digital images for different purposes. The watermarks are designed to be decoded or detected without the original images. The first one, called single watermark embedding (SWE), is used to embed a watermark bit sequence into digital images using two secret keys. The second technique, called multiple watermark embedding (MWE), extends SWE to embed multiple watermarks simultaneously in the same watermark space while minimizing the watermark (distortion) energy. The third technique, called iterative watermark embedding (IWE), embeds watermarks into JPEG-compressed images. The iterative approach of IWE can prevent the potential removal of a watermark in the JPEG recompression process. Experimental results show that embedded watermarks using the proposed techniques can give good image quality and are robust in varying degree to JPEG compression, low-pass filtering, noise contamination, and print-and-scan. Peter Hon-Wah Wong, Oscar C. Au, Yick Ming Yeung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | Novel motion compensation for wavelet video coding using overlappingabstractThe discrete wavelet transform (DWT) has several advantages in multi-resolution analysis and subband decomposition, which has been successfully used in image compression. In this paper, we will extend the usage of DWT to video compression and focus on motion estimation (ME) and motion compensation (MC) in wavelet domain. The ME used in this paper is Low-Band-Shift Motion Estimation (LBSME) [1] and a novel MC in wavelet domain is proposed. With this novel method, ME/MC in wavelet domain now is comparable with ME/MC in spatial domain and the performance is superior to other ME/MC methods in wavelet domain. Ming Fai Fu, Oscar C. Au, Wing Cheong Chan |
ICASSP | 2 |
| 2002 | Fast SOLA-based time scale modification using Modified Envelope MatchingabstractTime scale modification (TSM) of speech and audio is useful in many applications. Synchronized Overlap-and-Add (SOLA) is a time-domain TSM algorithm known to achieve good speech and audio quality. One problem of SOLA is that it requires a large amount of computation. In this paper, we propose a technique called Modified Envelope-Matching TSM (MEM-TSM) to simplify the computation. The. MEM-TSM improves the previously proposed EM-TSM in quality and speed up factor. The proposed algorithm can reduce computation by 300 times with good perceptual quality of time-scaled speech and audio. Our experimental results show that the quality of MEM-TSM is almost the same as SOLA. Peter Hon-Wah Wong, Oscar C. Au |
ICASSP | 2 |
| 2002 | A blind watermarking technique in JPEG compressed domainabstractWe propose a blind watermarking technique to embed a watermark in the JPEG compressed domain. Low frequency DCT coefficients are extracted to form an M-dimensional vector. Watermarking is achieved by modifying this vector in order to point to the centroid of a particular cell. This cell is determined according to the extracted vector, private keys and the watermark. A dual-key system is used to reduce the chance of the removal of watermark. An iterative approach is used to prevent the removal of the watermark by JPEG re-quantization. Experimental results show that the watermark can be detected when the watermarked image is further compressed using a larger scaling factor. Oscar C. Au, Peter Hon-Wah Wong |
ICIP (3) | 1 |
| 2002 | Temporal interpolation using wavelet domain motion estimation and motion compensationabstractTemporal frame subsampling is a simple but efficiency technique to achieve very high compression ratio in slow motion video. The discrete wavelet transform (DWT) is successful in still image compression and many researchers try to extend the usage of DWT to a video coder. In this paper, we incorporate these two ideas, together with the newly developed wavelet domain motion estimation method - low-band-shift motion estimation (LBSME) to develop a novel temporal interpolation scheme in the wavelet domain. Ming Fai Fu, Oscar C. Au, Wing Cheong Chan |
ICIP (3) | 2 |
| 2002 | A novel motion estimation algorithm for arbitrarily shaped video codingabstractIn this paper, we present a fast motion estimation algorithm for arbitrarily shaped video coding in MPEG-4. This novel algorithm takes advantage of our new discovery about the close relation between the best matching block and its shape information-alpha plane. Without any additional padding computation, this new motion estimation algorithm computes the sum of absolute difference (SAD) between two alpha planes rather than the pixels' intensities. Compared with the full search block-matching algorithm (recommended in the MPEG-4 standard), the proposed algorithm achieves an impressive speed-up ratio with very minor quality degradation and little bit-count increase. Extensive simulations are provided in this paper to demonstrate this fact. Andy C.-W. Yu, Bing Zeng 0001, Oscar C. Au |
ICME (1) | 3 |
| 2002 | Highly efficient predictive zonal algorithms for fast block-matching motion estimationabstractMotion estimation (ME) is an important part of any video encoding system since it could significantly affect the output quality of an encoded sequence. Unfortunately, this feature requires a significant part of the encoding time especially when using the straightforward full search (FS) algorithm. We propose two techniques, the generalized motion vector (MV) predictor and the adaptive threshold calculation, that can be used to significantly improve the performance of many existing fast ME algorithms. In particular, we apply them to create two new algorithms, named advanced predictive diamond zonal search and predictive MV field adaptive search technique, respectively, which can considerably reduce, if not essentially remove, the computational cost of ME at the encoder, while at the same time give similar, and in many cases better, visual quality with the brute force full search algorithm. The proposed algorithms mainly rely upon very robust and reliable predictive techniques and early termination criteria with parameters adapted to the local characteristics combined with the zonal based patterns. Our experiments verify the considerable superiority of the proposed algorithms versus the performance of possibly all other known fast algorithms, and FS. Alexis M. Tourapis, Oscar C. Au, Ming Lei Liou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | Data hiding watermarking for halftone imagesabstractIn many printer and publishing applications, it is desirable to embed data in halftone images. We proposed some novel data hiding methods for halftone images. For the situation in which only the halftone image is available, we propose data hiding smart pair toggling (DHSPT) to hide data by forced complementary toggling at pseudo-random locations within a halftone image. The complementary pixels are chosen to minimize the chance of forming visually undesirable clusters. Our experimental results suggest that DHSPT can hide a large amount of hidden data while maintaining good visual quality. For the situation in which the original multitone image is available and the halftoning method is error diffusion, we propose the modified data hiding error diffusion (MDHED) that integrates the data hiding operation into the error diffusion process. In MDHED, the error due to the data hiding is diffused effectively to both past and future pixels. Our experimental results suggest that MDHED can give better visual quality than DHSPT. Both DHSPT and MDHED are computationally inexpensive. Ming Sun Fu, Oscar C. Au |
IEEE Trans. Image Process. | 2 |
| 2001 | Data hiding in halftone images by stochastic error diffusionabstractWe propose a novel method called DHSED (data hiding stochastic error diffusion) to hide binary visual patterns in two error diffused halftone images. While one halftone image is only a regular error diffused image, stochastic error diffusion is applied to the other image to generate special stochastic characteristics with respect to the first image such that the visual pattern would appear when the two halftone images are overlaid Simulation results show that the two halftone images have good visual quality, and the hidden pattern appears with "normal" and "lower-than-normal" intensity when the two halftone images are overlaid. Ming Sun Fu, Oscar C. Au |
ICASSP | 2 |
| 2001 | N-dimensional zonal algorithms. The future of block based motion estimation?abstractThe popularity of zonal based algorithms for block based motion estimation has been increasing due to their superior performance in both terms of reduced complexity and superior quality versus other preexisting algorithms. In our previous work we mainly focused on generalizing the different parameters used in these algorithms and finding the possibly most efficient implementation. In this paper a further generalization of these algorithms is presented, where instead we mainly consider the way zones can be designed and what should be the ultimate goal of such an algorithm. As a result we present a framework of algorithms which can have applications not only in video coding, but also in other video signal processing areas, such as computer vision, video analysis, salient stills etc. We do so by initially considering the dimensionality of video data and how it can be most efficiently analyzed and exploited in the context of zonal algorithms. A formulization of these algorithms is then presented, according to which different implementations for different applications can be selected. Simulation results, for the simple 3-D case using the predictive diamond search (PDS) algorithm, a low complexity zonal algorithm, demonstrate the efficacy of the proposed techniques while still having low complexity. Higher order implementations using more dimensions and more efficient zonal algorithms can also be considered. Alexis M. Tourapis, Hye-Yeon Cheong, Ming Lei Liou, Oscar C. Au |
ICIP (3) | 4 |
| 2001 | Temporal interpolation of video sequences using zonal based algorithmsabstractTemporal interpolation has been proposed as a solution for increasing temporal resolution or even for predicting missing or corrupted frames within a video sequence. In this paper new techniques on temporal interpolation are presented, by mainly exploiting properties of the very popular and highly efficient zonal based motion estimation algorithms, and by introducing several other techniques such as multihypothesis motion compensation, motion classification and temporal/spatial filtering. In addition we further give an analysis on when temporal interpolation should be employed, thus possibly avoiding unwanted artifacts created from this process, while at the same time significantly improving the overall performance of the interpolation. Alexis M. Tourapis, Hye-Yeon Cheong, Ming Lei Liou, Oscar C. Au |
ICIP (3) | 4 |
| 2001 | Fast adaptive-spatial-varying filtering for inverse halftoning
Ming Sun Fu, Oscar C. Au |
VCIP | 2 |
| 2001 | Predictive motion vector field adaptive search technique (PMVFAST): enhancing block-based motion estimation
Alexis M. Tourapis, Oscar C. Au, Ming Lei Liou |
VCIP | 2 |
| 2001 | Advanced deinterlacing techniques with the use of zonal-based algorithms
Alexis M. Tourapis, Oscar C. Au, Ming Lei Liou |
VCIP | 2 |
| 2001 | Halftone image data hiding with intensity selection and connection selection
Ming Sun Fu, Oscar C. Au |
Signal Process. Image Commun. | 2 |
| 2000 | Data hiding by smart pair toggling for halftone imagesabstractThere is growing interest in hiding data for authentication and copyright control in halftone images printed in books, newspapers and by computer printers. A previous data hiding method, the data hiding by pair-toggling (DHPT), is reasonably good but introduce considerable visual artifacts. In this paper, we analyze the sources of the artifacts in DHPT and propose an improvement by using smart pair toggling. Simulation results suggest that the proposed data hiding by smart pair-toggling (DHSPT) algorithm can hide the same amount of data while generating halftone images with considerably better visual quality than DHPT. Ming Sun Fu, Oscar C. Au |
ICASSP | 2 |
| 2000 | Improved acoustics modeling for speech recognition using transformation techniquesabstractIn statistical speech recognition, misclassification often occurs when there is a mismatch between the incoming signal and the acoustics model inside the recognizer. In order to combat this problem, techniques such as Cepstral Mean Subtraction, Vocal Tract Normalization, adaptation and pronunciation model can be used. In this paper, we proposed a new approach based on transformation technique where the output distribution function in the HMM model, a Gaussian probability density function, could be transformed to match the estimated distribution of the incoming signal by using a memoryless invertible nonlinearity function. Since the new density still has a Gaussian form, the function could be completely characterized by using the Expectation Maximization (EM) algorithm. 1. Carrson C. Fung, Oscar C. Au, Chi H. Yim, Cyan L. Keung |
INTERSPEECH | 2 |
| 2000 | Probabilistic compensation of unreliable feature components for robust speech recognition
Cyan L. Keung, Oscar C. Au, Chi H. Yim, Carrson C. Fung |
INTERSPEECH | 2 |
| 2000 | Auditory spectrum based features (ASBF) for robust speech recognition
Chi H. Yim, Oscar C. Au, Cyan L. Keung, Carrson C. Fung |
INTERSPEECH | 2 |
| 2000 | Parameter estimation for image/video transcodingabstractIn most of the current image/video compression standard, such as JPEG, MPEG, H.263, etc, the DCT is used. The DCT coefficients for every AC frequency component usually have Laplacian distribution. Such a distribution is critical in the image or video transcoding system. Therefore, the parameter in the distribution should be well estimated. Different from the normal parameter estimation, in the transcoding system, the original DCT value is unavailable, and only the dequantized value is available. In this paper, we propose 3 estimation methods based on the dequantized value for the transcoding system. The simulation demonstrates that they can yield very good results. In addition, under a fixed bit rate, they can work even better than estimation with the original DCT value in the sense of image/frame's PSNR. Zihua Guo, Oscar C. Au, Khaled Ben Letaief |
ISCAS | 2 |
| 2000 | Optimizing the MPEG-4 encoder-advanced diamond zonal searchabstractMotion estimation (ME) is an important part of the MPEG-4 encoder, due to its significant impact on the bitrate and the output quality of the encoded sequence. Unfortunately this feature occupies a significant part of the encoding time especially when using the straightforward full search (FS) algorithm. The diamond search (DS) was recently accepted as a fast motion estimation algorithm for the MPEG-4 VM. In this paper we propose a new algorithm named advanced diamond zonal search (ADZS), which is significantly faster than DS (in terms of number of checking points and total encoding time) and gives similar, if not better, quality (in terms of PSNR) of the output sequence. This is more obvious in the high bit rate cases. Our experiments verify the superiority of the proposed algorithm. Alexis M. Tourapis, Oscar C. Au, Ming Lei Liou, Guobin Shen, Ishfaq Ahmad 0001 |
ISCAS | 2 |
| 2000 | Image watermarking using spread spectrum technique in log-2-spatio domainabstractIn this paper, we propose to embed the watermark information in the log-2-spatio domain by means of spread spectrum technique. In log-2-spatio domain, the variance of the information is reduced significantly. This improves the efficiency and robustness of spread spectrum technique. Low intensity and mid-band regions are selected to embed the information in order to guarantee an invisible watermark as well as the robustness to JPEG compression. Simulation results show that the embedded information still survives up to the JPEG compression ratio of 14.7. Peter Hon-Wah Wong, Oscar C. Au, Justy W. C. Wong |
ISCAS | 2 |
| 2000 | Novel fast motion estimation for frame rate/structure conversionabstractDifferent multimedia applications and transmission channels require different resolution, frame rates/structures and bitrate; there is often a need to transcode the stored compressed video to suit the needs of these various applications. This paper is concerned about fast motion estimation for frame rate/structure conversion. In this paper, we proposed several novel algorithms that exploit the correlation of the motion vectors in the original video and those in the transcoded video. We achieve a much higher quality than existing fast search algorithms with much lower complexity. Justy W. C. Wong, Oscar C. Au, Peter Hon-Wah Wong |
ISCAS | 2 |
| 2000 | Hiding data in halftone image using modified data hiding error diffusion
Ming Sun Fu, Oscar C. Au |
VCIP | 2 |
| 2000 | New predictive diamond search algorithm for block-based motion estimation
Alexis M. Tourapis, Guobin Shen, Ming Lei Liou, Oscar C. Au, Ishfaq Ahmad 0001 |
VCIP | 4 |
| 1999 | Fast ad-hoc inverse halftoning using adaptive filteringabstractWe propose a novel fast inverse halftoning technique using an adaptive spatial varying filtering. The proposed algorithm is significantly simpler than most existing algorithm while achieving a PSNR close to that of the set theoretic POCS. Oscar C. Au |
ICASSP | 1 |
| 1999 | An Advanced Zonal Block Based Algorithm for Motion EstimationabstractEfficient motion estimation is very important for compressing video in standards like MPEG1/2/4 and ITU-T H.261/263. In this paper a new algorithm is presented which can outperform most of the traditional fast motion estimation algorithms in both speed and quality. In addition, in some cases this algorithm can achieve even better visual quality, than the “optimal” but computational intensive “full search” algorithm. Alexis M. Tourapis, Oscar C. Au, Ming Lei Liou, Guobin Shen |
ICIP (2) | 2 |
| 1999 | Sinusoidal representation and auditory model-based parametric matching and smoothing and its application in speech analysis/synthesis
Oscar C. Au, Cyan L. Keung, Chi H. Yim |
EUROSPEECH | 1 |
| 1999 | A novel approach of low bit-rate speech coding based on sinusoidal representation and auditory model
Oscar C. Au, Cyan L. Keung, Chi H. Yim |
EUROSPEECH | 2 |
| 1999 | Modified one-bit transform for motion estimationabstractMotion estimation using the one-bit transform (1BT) was proposed by Natarajan, Bhaskaran and Konstantinides (see ibid., vol.7, p.702-06, 1997) to achieve large computation reduction. However, it degrades the predicted image by almost 1 dB as compared with full search. We propose a modification to the 1BT by adding conditional local searches. Simulation results show that the proposed modification improves the peak signal-to-noise ratio (PSNR) significantly at the expense of slightly increased computational complexity. A variant of the proposed modification called the multiple-candidate two-step search (M2SSFS) is found to be particularly good for high quality, high bit rate video coding. In the MPEG-1 simulation, its PSNR is within 0.1 dB from that of full search at bit rates higher than 1 Mbit/s with a computation reduction factor of ten. Peter Hon-Wah Wong, Oscar C. Au |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1998 | Predictive Motion Estimation for Reduced-Resolution Video from High-Resolution Compressed VideoabstractTo convert a compressed video sequence to a lower-resolution compressed video, one typically needs to decompress the original sequence, down-sample each frame, and recompress it. It involves motion estimation in the reduced sequence which is computational intensive. We propose a novel fast algorithm to predict the motion vector of the reduced video by using the original motion information in the compressed bitstream. We achieve a much higher quality than existing algorithms with low additional complexity. Justy W. C. Wong, Oscar C. Au, Peter Hon-Wah Wong, Alexis M. Tourapis |
ICIP (2) | 2 |
| 1997 | Fast Fractal Encoding in Frequency DomainabstractFractal image compression applies the self-similarity property of an image. Much research has been done to study the properties of fractal coding in the image domain. In this paper, however, we try to explore the features of fractal coding in the frequency domain. We firstly overview the properties of fractal coding in the image domain, then we derive the corresponding formula of scaling factor and offset of affine transform in the DCT domain. Applying the energy compaction property of the DCT, we propose a fast fractal encoding algorithm by using only a small number of low frequency DCT coefficients in measuring the similarity between range block and domain block. We further propose a possible fast hybrid fractal encoding algorithm which combines existing fast search methods, statistical normalization and frequency domain comparison. Oscar C. Au, Ming Lei Liou, L. K. Ma |
ICIP (2) | 1 |
| 1996 | Objective speech quality measure for cellular phoneabstractCellular phone network speech quality monitoring is a regular task performed by the cellular service providers. Objective speech quality measures are needed in such tasks to provide a reasonably accurate estimate of subjective quality of the network. We performed an experiment to collect real distorted data, conducted a survey to obtain subjective quality measure of the collected speech samples and studied the statistical correlation of 32 objective speech quality measures with the subjective measures. Four of the objective measures were found to be good. Synchronization was found to be important. Kai Hung Lam, Oscar C. Au, Ching Chuen Chan, K. F. Hui, S. F. Lau |
ICASSP | 2 |
| 1996 | Modified motion compensated temporal frame interpolation for very low bit rate videoabstractCommon techniques such as frame repetition or linear interpolation for reconstructing skipped frames in temporally subsampled video sequences tend to introduce undesirable artifacts. A previously proposed technique, fast motion compensated temporal interpolation (FMCTI), can interpolate video frames in the time domain with good interpolated image quality. However, the reconstructed frames tend to be blocky. We propose to incorporate a pixel-based motion classifier into FMCTI to solve the problem. The resulting technique, modified fast motion compensated temporal interpolation (MFMCTI), was found by simulation to give significantly reduced blocking artifacts. Chi-Kong Wong, Oscar C. Au |
ICASSP | 2 |
| 1996 | Bit rate and blocking artifact reduction by iterative pre-distortionabstractBlock transform coding is widely used in image and video compression methods. Due to quantization error, there are blocking artifacts in the decoded image. This paper presents a pre-processing method, called pre-distortion, to reduce both the bit rate and blocking artifacts. The proposed algorithm is an iterative algorithm which reduces the blocking artifacts of the decoded image by introducing distortion intentionally in the encoding process. Simulation shows that, with pre-distortion, additional compression can be achieved with slight degradation in visual quality. Moreover, the blocking artifacts of the decoded image are reduced. Yiu-Hung Fok, Oscar C. Au, Corina Chang |
ICIP (2) | 2 |
| 1996 | On sensitivity of detector structures to contaminants in non-white mixture noise model
Oscar C. Au |
Signal Process. | 1 |
| 1995 | Objective speech measure for Chinese in wireless environmentabstractThe cellular phone is becoming an important means of mobile wireless communication, especially in metropolitan areas. One of the important operating considerations of the cellular phone service providers is the maintainence of the speech quality of the cellular phone network. Subjective evaluation by repeated listening tests at various sites within the coverage area is impractical due to its intrinsic laborious and expensive nature. As a result, it would be much desirable to have an automatic objective evaluation system which applies a good objective speech measure to estimate the statistical average of subjective opinions of the typical conversational speech sentences sent through the cellular network. While extensive work was done for objective speech measures for languages such as English, Japanese, French, and other western languages, little has been done for Chinese. In addition, little has been done to quantify speech quality in the wireless environment. Kai Hung Lam, Oscar C. Au, Ching Chuen Chan, K. F. Hui, S. F. Lau |
ICASSP | 2 |
| 1995 | Novel fast block motion estimation in feature subspaceabstractMotion estimation and compensation are widely used in video coding. This paper presents two fast block matching algorithms for motion estimation. These algorithms use the subspace features of a block to determine the block distance. With a search block size of N/spl times/N, the proposed algorithms can achieve a computation reduction factor of N/2 while retaining close-to-optimal performance in the mean absolute difference (MAD) sense. Yiu-Hung Fok, Oscar C. Au, Ross Murch |
ICIP | 2 |
| 1994 | An Improved Fast Feature-Based Block Motion EstimationabstractA fast block matching algorithm in the feature domain was proposed by Fok and Au (1993) which achieves a computation reduction factor of N/2 with search block size of N/spl times/N. This paper presents an improved fast block matching algorithm in the integral projections feature domain. This algorithm is based on the fast block matching algorithm in the feature domain and a new motion field sub-sampling scheme. By utilizing the properties of the features and the motion field sub-sampling scheme, with a search block size of N/spl times/N, the improved algorithm can achieve a computation reduction factor of N while retaining close-to-optimal performance in the mean absolute difference (MAD) sense.> Yiu-Hung Fok, Oscar C. Au |
ICIP (3) | 2 |
| 1991 | On transformation noise as a model for correlated noiseabstractThe authors consider transformation noise processes, which are those that can be generated by passing an underlying noise process through a memoryless invertible nonlinearity. They point out the close relationship between the dependency structure of the transformation noise and that of the underlying noise. The use of the class of transformation noise generated from multivariate Gaussian background noise to approximate other noise processes is proposed. This class has tractable, closed-form joint densities. Some ad hoc ways to find the parameters in the model are suggested. It is shown that the nonlinearities in the optimal detector, in the case of transformation noise with a mixture marginal density are approximately the sum of the nominal term and a term proportional to the parameter in .> Oscar C. Au, John B. Thomas |
ICASSP | 1 |