VLDB 2026 Research / reviewers in the wild / expert
Yaowu Xu
dblp:81/5014
· DBLP profile ↗
41ranked-venue papers
7as first author
9since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Artificial intelligence and machine learning · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Saliency Map Approach to Optimize VMAF for Video and Image CompressionabstractThe Video Multi-method Assessment Fusion (VMAF) has demonstrated a better correlation with Human Visual System than the conventional objective metrics, and has gradually gained adoption in the industry that needs to monitor the visual quality of compressed videos. However, due to its machine learning nature, it can not be expressed through a simple and explicit mathematical formula, which makes it difficult to incorporate VMAF into the rate-distortion optimization framework in video compression. In this work, we propose a new perspective that decomposes the VMAF as a superposition of spatial and temporal factors. The spatial factor, which also directly applies to image quality evaluation, is approximated by a saliency map. It in conjunction with the temporal factor approximated by the motion quantities allows a simple analytical formula that translates the mean squared distortion at each pixel to its impact to the overall VMAF metric. The proposed hypothesis is embedded into the rate-distortion optimization framework, and is experimentally shown to provide considerable coding gains in VMAF for both image and video compression. Jingning Han, Yaowu Xu |
DCC | 3 |
| 2023 | Learned Image Compression Guided Adaptive Quantization for Perceptual QualityabstractNeural network based image compression has made significant progress in recent years. The learned image codecs are commonly reported to outperform their conventional counterparts in perceptual quality. Despite the superior performance, the learned image codecs are much more complex to decode, which hinders their usage in practice. Without a significant advance in hardware capability, the conventional image codec will likely remain a primary component for large scale image services. It is therefore desirable to improve the quality of conventional image codecs. In this paper, we present an adaptive quantization approach to the conventional image codec with the help of learned image codecs to improve its perceptual quality. It exploits the bit allocation of the neural network based image codec to adapt the quantizers on a block basis. It is experimentally shown that the proposed method provides considerable perceptual quality improvements over other leading contenders. Ruiqi Geng, Bohan Li 0006, Maryla Ustarroz-Calonge, Frank Galligan, Jingning Han, Yaowu Xu |
ICIP | 7 |
| 2022 | An Efficient Scheme of Multi-Hypothesis Motion Compensated Prediction for Video Coding ApplicationsabstractPrior research has demonstrated that the multi-hypothesis motion compensated prediction (MCP) can theoretically provide a better prediction quality than single-reference MCP, thereby improving the compression efficiency in video coding. However, the existing multi-hypothesis MCP methods typically require either additional rate cost to transmit the motion vectors, or significant decoding complexity to conduct the motion search at the decoder end, which is usually expensive. In this work, we propose a novel scheme to materialize the multi-hypothesis MCP that requires no additional rate cost, nor extra motion search on either the encoder or decoder side. Various approaches to synthesize these available multiple references to form the inter prediction are presented. We experimentally demonstrate that the proposed scheme provides considerable and consistent coding gains across a wide range of operating points. Bohan Li 0006, Jingning Han, Yaowu Xu |
ICIP | 3 |
| 2022 | Differential Contrast Based Adaptive Quantization for Perceptual Quality Optimization in Image CodingabstractWe consider the perceptual quality optimization in image coding through adaptive quantization. A differential contrast model is proposed to measure the visual sensitivity to the quantization distortions, and thereby deriving the spatially adaptive quantization strategy. A complementary quantitative approach is provided as a means to efficiently calculate the proposed differential contrast model. The resulting visual quality improvement is experimentally demonstrated. Jingning Han, Frank Galligan, Pascal Massimino, Paul Wilkins, Wan-Teh Chang, Yannis Guyon, Yaowu Xu, Jim Bankoski |
ICIP | 8 |
| 2022 | Probability Model Estimation for M-Ary Random VariablesabstractThe entropy coding system in AV1 processes syntax elements as M-ary random variables. In comparison to the binarization approach used in its predecessor VP9 that converts an M-ary random variables into a series of binary symbols for entropy coding, the M-ary random variable approach provides higher throughput for hardware decoders. The non-binary probability table associated with the M-ary random variable, however, poses new challenges in the probability model estimation process beyond the binary case. This paper provides a retrospect of the probability model estimation for M-ary random variables used in AV1, and proposes new algorithms for the probability estimation process to improve the compression efficiency. Its efficacy is experimentally demonstrated under various testing conditions. Jingning Han, Yaowu Xu |
ICIP | 2 |
| 2021 | Adaptive GOP Size Decision for Multi-Pass Video Coding Based on Hidden Markov ModelabstractMulti-pass coding is a widely utilized technique to improve the compression efficiency in video coding, where frame statistics are collected from the previous passes and then analyzed to provide better encoder decisions, such as rate control parameters, prediction mode selection, motion estimation, etc. In this paper, a novel method to determine the size of each group of picture (GOP) using the multi-pass information is presented. In particular, we propose to categorize frames into regions with different natures, including stationary, high-variance, blending, and scene cut, through analyzing the frame statistics generated from the previous passes using a hidden Markov model. The GOP size is then determined based on the region types and the inter frame correlations. It is experimentally shown that the proposed adaptive GOP size decision provides considerable coding performance improvements over conventional fixed GOP length. Bohan Li 0006, Jingning Han, Yaowu Xu |
ICASSP | 3 |
| 2021 | Study On Coding Tools Beyond AV1abstractThe Alliance for Open Media has recently initiated coding tool exploration activities towards the next-generation video coding beyond AV1. With this regard, this paper presents a package of coding tools that have been investigated, implemented and tested on top of the codebase, known as libaom, which is used for the exploration of next-generation video compression tools. The proposed tools cover several technical areas based on a traditional hybrid video coding structure, including block partitioning, prediction, transform and loop filtering. The proposed coding tools are integrated as a package, and a combined coding gain over AV1 is demonstrated in this paper. Furthermore, to better understand the behavior of each tool, besides the combined coding gain, the tool-on and tool-off tests are also simulated and reported for each individual coding tool. Experimental results show that, compared to libaom, the proposed methods achieve an average 8.0% (up to 22.0%) overall BD-rate reduction for All Intra coding configuration a wide range of image and video content. Xin Zhao 0003, Madhu Peringassery Krishnan, Yixin Du, Shan Liu 0001, Debargha Mukherjee, Yaowu Xu, Adrian Grange |
ICME | 7 |
| 2021 | A Temporal Filtering Approach Based on Optical Flow Estimation for Video CodingabstractVideo coding uses motion compensated prediction to exploit temporal correlations for compression efficiency. Prior works have demonstrated that substantial coding gains can be achieved by decomposing a long-term reference frame into a synthetic reference-only (non-displayable) frame and an overlay displayable frame that resembles the original frame. The source of the reference-only frame is typically generated by temporal filtering along the motion trajectories across nearby frames, where the motion trajectories are built using block matching algorithms (BMAs), thereby reducing the noise level within this synthetic frame. Noting that the efficacy of the conventional BMAs are limited to capturing translational motion activities, this paper proposes a novel approach that uses a per-pixel motion field generated by an optical flow estimation to form the motion trajectory for more efficient temporal filtering. It is experimentally shown that the proposed method better captures non-translational motion activities, which translates into considerable coding gains for video signals with such complicate motion patterns. Bohan Li 0006, Lauren Partin, Jingning Han, Yaowu Xu |
MMSP | 4 |
| 2021 | A Technical Overview of AV1abstractThe AV1 video compression format is developed by the Alliance for Open Media consortium. It achieves more than a 30% reduction in bit rate compared to its predecessor VP9 for the same decoded video quality. This article provides a technical overview of the AV1 codec design that enables the compression performance gains with considerations for hardware feasibility. Jingning Han, Bohan Li 0006, Debargha Mukherjee, Ching-Han Chiang, Adrian Grange, Hui Su, Sarah Parker, Sai Deng, Urvang Joshi, Yue Chen 0040, Yunqing Wang, Paul Wilkins, Yaowu Xu, Jim Bankoski |
Proc. IEEE | 14 |
| 2020 | Video Denoising for the Hierarchical Coding Structure in Video CodingabstractModern video codecs explore the temporal and spatial correlations of video signal to achieve the goal of compression. The noise in video signal corrupts such temporal and spatial correlations and thus is difficult to compress. Denoising of video signal is a potential solution to this problem. Despite the significant progress in video denoising in recent years, there is few research exploring the feasibility of denoising for video compression. In this work, we demonstrate that video denoising is able to significantly reduce bit rates while maintaining the subjective and objective quality when appropriately incorporated into the hierarchical coding structure of video coding. We present a temporal filtering algorithm for denoising and apply it to AV1 for lossy video compression. We obtain a significant compression efficiency improvement over videos of different resolutions, types, and noise. Jingning Han, Yaowu Xu |
DCC | 3 |
| 2020 | Online Probability Model Estimation for Video CompressionabstractModern video codec uses arithmetic coding for entropy coding. The arithmetic coding asymptotically achieves the entropy bound provided the true probability distribution. Hence the compression efficiency heavily relies on the ability to capture the time-variant probability model in video signals. Variants of first-order linear probability model update schemes have been used in recent generation video codecs. Built on top of those, a multimodal estimation scheme that forms a higher order probability model update has been proposed in this work. We experimentally demonstrate its coding efficiency. Jingning Han, Yaowu Xu |
DCC | 3 |
| 2020 | A Non-local Mean Temporal Filter for Video CompressionabstractModern video codecs exploit the temporal and spatial correlations of video signal to achieve compression. The noise in video signal corrupts such correlations and impairs the coding efficiency. Prior works in VP8, VP9, and HEVC exploit the use of temporal filtering to remove certain noise from the source signal. They typically compare a pair of pixels along a motion trajectory and decide the filter coefficients based on the pixel value difference. It is observed that such noise removal allows better rate-distortion performance trade off and hence improves the objective compression efficiency. Note that the compression distortion is evaluated against the original video signal in all cases. This work proposes a non-local mean temporal filter for noise removal. Instead of comparing a pair of pixels along the motion trajectory, it compares two pixel blocks surrounding the pixels of interest. Their distance in L2 norm is then normalized by the frame noise level, which is used to determine the temporal filter coefficients in a non-parametric model. It is experimentally shown that the proposed non-local mean filter approach achieves improved compression efficiency over other contenders. Jingning Han, Yaowu Xu |
ICIP | 3 |
| 2020 | Machine Learning Based Symbol Probability Distribution Prediction For Entropy Coding In Av1abstractEntropy coding is a lossless data compression technique that is widely applied in video codecs to encode syntax elements into bitstreams. Efficient entropy coding requires accurate prediction of the probability distribution of the encoded symbols. In AV1, multi-symbol arithmetic coding is adopted. The symbol probability is derived with handcrafted context models and lookup tables that store the predicted probabilities corresponding to different entropy contexts. The lookup table based scheme has some fundamental deficiencies. The entropy context features have to be discrete so that they can be used to index the lookup tables. To reduce the size of the lookup table, the number of contexts cannot be very large. Moreover, the probability distributions stored in the lookup tables are maintained separately without taking their correlations into consideration. In this paper, we propose a machine learning based scheme that achieves more accurate symbol probability prediction for entropy coding. The proposed approach is implemented in AV1 for the entropy coding of intra prediction modes. Experimental results demonstrate that it can improve the efficiency of entropy coding significantly. Mingliang Chen 0001, Hui Su, Sai Deng, Yaowu Xu |
ICIP | 4 |
| 2020 | VMAF Based Rate-Distortion Optimization for Video CodingabstractVideo Multi-method Assessment Fusion (VMAF) is a machine-learning based video quality metric. It is experimentally shown to provide higher correlation with human visual system as compared to conventional metrics like peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) in many scenarios and has drawn considerable interest as an alternative metric to evaluate the perceptual quality. This work proposes a systematic approach to improve the video compression performance in VMAF. It is composed of multiple components including a pre-processing stage with a complement automatic filter parameter selection, and a modified rate-distortion optimization framework tailored for VMAF metric. The proposed scheme achieves on average 37% BD-rate reduction in VMAF, as compared to conventional video codec optimized for PSNR. Sai Deng, Jingning Han, Yaowu Xu |
MMSP | 3 |
| 2020 | Optical Flow Based Co-Located Reference Frame for Video CompressionabstractThis paper proposes a novel bi-directional motion compensation framework that extracts existing motion information associated with the reference frames and interpolates an additional reference frame candidate that is co-located with the current frame. The approach generates a dense motion field by performing optical flow estimation, so as to capture complex motion between the reference frames without recourse to additional side information. The estimated optical flow is then complemented by transmission of offset motion vectors to correct for possible deviation from the linearity assumption in the interpolation. Various optimization schemes specifically tailored to the video coding framework are presented to further improve the performance. To accommodate applications where decoder complexity is a cardinal concern, a block-constrained speed-up algorithm is also proposed. Experimental results show that the main approach and optimization methods yield significant coding gains across a diverse set of video sequences. Further experiments focus on the trade-off between performance and complexity, and demonstrate that the proposed speed-up algorithm offers complexity reduction by a large factor while maintaining most of the performance gains. Bohan Li 0006, Jingning Han, Yaowu Xu, Kenneth Rose |
IEEE Trans. Image Process. | 3 |
| 2019 | A Multi-Pass Coding Mode Search Framework For AV1 Encoder OptimizationabstractThe AV1 codec recently released by the Alliance of Open Media provides nearly 30% BDrate reduction over its predecessor VP9. It substantially extends the available coding block sizes and supports a wide range of prediction modes. There are also a large variety of transform kernel types and sizes. The combination provides an extremely wide range of flexible coding options. To translate such flexibility into compression efficiency, the encoder needs to conduct an extensive search over the space of coding modes. Optimization of the encoder complexity and compression efficiency trade-off is critical to productionizing AV1. Many research efforts have been devoted to devising feature space based pruning methods ranging from decision rules based on some simple observations to more complex neural network models. A multi-pass coding mode search framework is proposed in this work to provide a structural approach to reduce the search volume. It decomposes the original high dimensional space search into cascaded stages of lower dimensional space searches. To retain a near optimal search result, the scheme departs from conventional dimension reduction approach in which one retains a single winner at each stage, and uses that winner for the next stage (dimension). Instead, this framework retains a subset of the states that are the most likely winners at each stage, which are then fed into the next stage to find the next subset of winners. The subset size at each stage is determined by the likelihood that the optimal route will be captured in the current stage. Changing this likelihood parameter tunes the encoder for speed and compression performance trade-off. This framework can integrate with most existing feature based methods at its various stages. The framework provides 60% encoding time reduction at the expense of 0.6% compression loss in libaom AV1 encoder. Ching-Han Chiang, Jingning Han, Yaowu Xu |
DCC | 3 |
| 2019 | Machine Learning Accelerated Partition Search for Video EncodingabstractWith more complex partitioning structures in recent generations of video coding standards, the computation complexity of video encoder for partition block size search has been increasing drastically. To expedite the overall encoding process, it is desired to make faster partitioning decisions without much compression performance degradation. In this paper, we propose a multi-scale multi-stage machine learning(ML) based framework to accelerate partition block size search. The framework includes a collection of ML models, each dedicated to make a simple decision for a particular block size at a particular stage during the partitioning rate-distortion optimization(RDO) process. The ML models can predict whether the RD evaluation of certain partition block sizes can be skipped, saving unnecessary computation in the encoder. The proposed approach is implemented and tested on VP9 with the open source library libvpx. Significant encoding speed improvement has been observed with neglectable compression performance regression. The framework and methodology can be easily applied to other video codecs and implementations as well. Hui Su, Chi-Yo Tsai, Yunqing Wang, Yaowu Xu |
ICIP | 4 |
| 2019 | JND-based Perceptual Rate Distortion Optimization for AV1 EncoderabstractAV1 is the next-generation open video coding format, and it can achieve significant coding efficiency with novel coding tools. It supports Lagrangian rate distortion optimization (RDO) method to optimize the coding performance. However, the distortion and the Lagrangian multiplier used in RDO ignore the characteristics of human visual system (HVS), which leads to insufficiency for perceptual video coding. To solve this problem, a perceptual RDO scheme based on the Just Noticeable Distortion (JND) threshold of HVS is proposed. The JND for each pixel is first measured according to three perceptual features: luminance adaptation, masking effects and structure sensitivity. Based on the observation that the regions with smaller distortion visibility thresholds are more sensitive to HVS, a JND-based Lagrangian multiplier is derived to adaptively adjust the rate-distortion (RD) performance for each coding block. Experiments demonstrate that the proposed method can achieve an average SSIM-based -3.93% BD-Rate saving compared with the original AV1 encoder, which effectively improve the coding performance. Li Song 0001, Rong Xie 0004, Jingning Han, Yaowu Xu |
PCS | 5 |
| 2018 | Co-located Reference Frame Interpolation Using Optical Flow Estimation for Video CompressionabstractThe hierarchical coding structure that supports bi-directional motion compensated prediction is commonly used for video compression efficiency. Conventional approach directly seeks the reference pixel block from each individual reference frame and use it or its linear combinations for prediction. It largely ignores the motion information between these reference frames. To fully utilize all the information from the bi-directional reference frames, this work builds a per-pixel motion field that connects the two-sided reference frames using optical flow estimation. A reference frame is then interpolated at the current frame location. This collocated reference frame effectively accounts for the true motion trajectories in the video signal including both translational and the more complex non-translational motion models, which are beyond the capability of the conventional block-based motion compensated prediction. The scheme is experimentally shown to provide substantial compression performance gains. A number of optimization designs are proposed to make the codec complexity feasible while largely maintaining the coding performance. Bohan Li 0006, Jingning Han, Yaowu Xu |
DCC | 3 |
| 2018 | Efficient AV1 Video Coding Using a Multi-layer FrameworkabstractThis paper proposes a multi-layer multi-reference prediction framework for effective video compression. Current AOM/AV1 baseline uses three reference frames for the inter prediction of each video frame. This paper first presents a new coding tool that extends the total number of reference frames in both forward and backward prediction directions. A multi-layer framework is then described, which suggests the encoder design and places different reference frames within one Golden Frame (GF) group to different layers. The multi-layer framework leverages the existing coding tools in the AV1 baseline, including the tool of "show_existing_frame" and the reference frame buffer update module of a wide flexibility. The use of extended ALTREF_FRAMEs is proposed, and multiple ALTREF_FRAME candidates are selected and widely spaced within one GF group. ALTREF_FRAME is a constructed, no-show reference obtained through temporal filtering of a look-ahead frame. In the multi-layer structure, one reference frame may serve different roles for the encoding of different frames through the virtual index manipulation. The experimental results have been collected over several video test sets of various resolutions and characteristics both texture- and motion-wise, which demonstrate that the proposed approach achieves a consistent coding gain compared to the AV1 baseline. For instance, using PSNR as the distortion metric, an average bitrate saving of 5.57+% in BDRate is obtained for the CIF-level resolution set, some of which has a gain of up to 13+%, and 4.47% on average for the VGA-level resolution set, some of which up to 18+%. Zoe Liu, Debargha Mukherjee, Jingning Han, Paul Wilkins, Yaowu Xu, Kenneth Rose |
DCC | 6 |
| 2018 | A Motion Vector Entropy Coding Scheme Based on Motion Field Referencing for Video CompressionabstractVideo codec exploits the temporal correlations in video signal through block-based motion compensated prediction. The motion vector associated with each prediction unit needs to be coded in the bit-stream. A differential coding scheme that employs the motion information from spatial neighbors and collocated blocks in the reference frames to predict the current motion vector is commonly used. Its efficacy is largely limited to track consistent or slow motion activities. A linear projection model is proposed in this work to create a motion field estimation that is capable to capture motion trajectory with high velocity. The resulting motion field motion vectors (MFMV) are fed into a dynamic motion vector referencing system as candidates in addition to those obtained from the spatial neighboring blocks. It allows the codec to closely track complex motion activities that the spatial neighbors or collocated motion vector referencing system usually fail to keep synchronized with. The MFMV system improves the prediction quality of the motion vectors and substantially reduces the energy in the difference motion vector for entropy coding, which translates into considerable compression performance improvements, especially for video sequences that contain complex motion activities. A number of design considerations to make the computation efficiency in both hardware and software platforms practical for production are discussed. Jingning Han, Yaowu Xu, Jim Bankoski |
ICIP | 4 |
| 2018 | A Hybrid Weighted Compound Motion Compensated Prediction for Video CompressionabstractCompound motion compensated prediction that combines reconstructed reference blocks to exploit the temporal correlation is a major component in the hierarchical coding scheme. A uniform combination that applies equal weights to reference blocks regardless of distances towards the current frame is widely employed in mainstream codecs. Linear distance weighted combination, while reflecting the temporal correlation, is likely to ignore the quantization noise factor and hence degrade the prediction quality. This work builds on the premise that the compound prediction mode effectively embeds two functionalities - exploiting temporal correlation in the video signal and canceling the quantization noise from reference blocks. A modified distance weighting scheme is introduced to optimize the trade-off between these two factors. It quantizes the weights to limit the minimum contribution from both reference blocks for noise cancellation. We further introduces a hybrid scheme allowing the codec to switch between the proposed distance weighted compound mode and the averaging mode to provide more flexibility for the trade-off between temporal correlation and noise cancellation. The scheme is implemented in the AV1 codec as part of the syntax definition. It is experimentally demonstrated to provide on average 1.5% compression gains across a wide range of test sets. Jingning Han, Yaowu Xu |
PCS | 3 |
| 2018 | An Overview of Core Coding Tools in the AV1 Video CodecabstractAV1 is an emerging open-source and royalty-free video compression format, which is jointly developed and finalized in early 2018 by the Alliance for Open Media (AOMedia) industry consortium. The main goal of AV1 development is to achieve substantial compression gain over state-of-the-art codecs while maintaining practical decoding complexity and hardware feasibility. This paper provides a brief technical overview of key coding techniques in AV1 along with preliminary compression performance comparison against VP9 and HEVC. Yue Chen 0040, Debargha Mukherjee, Jingning Han, Adrian Grange, Yaowu Xu, Zoe Liu, Sarah Parker, Hui Su, Urvang Joshi, Ching-Han Chiang, Yunqing Wang, Paul Wilkins, Jim Bankoski, Luc N. Trudeau, Nathan E. Egge, Jean-Marc Valin, Thomas Davies 0002, Steinar Midtskogen, Andrey Norkin, Peter De Rivaz |
PCS | 5 |
| 2017 | A constrained adaptive scan order approach to transform coefficient entropy codingabstractTransform coefficient coding is a key module in modern video compression systems. Typically, a block of the quantized coefficients are processed in a pre-defined zig-zag order, starting from DC and sweeping through low frequency positions to high frequency ones. Correlation between magnitudes of adjacent coefficients is exploited via context based probability models to improve compression efficiency. Such scheme is premised on the assumption that spatial transforms compact energy towards lower frequency coefficients, and the scan pattern that follows a descending order of the likelihood of coefficients being non-zero provides more accurate probability modeling. However, a pre-defined zig-zag pattern that is agnostic to signal statistics may not be optimal. This work proposes an adaptive approach to generate scan pattern dynamically. Unlike prior attempts that directly sort a 2-D array of coefficient positions according to the appearance frequency of non-zero levels only, the proposed scheme employs a topological sort that also fully accounts for the spatial constraints due to the context dependency in entropy coding. A streamlined framework is designed for processing both intra and inter prediction residuals. This generic approach is experimentally shown to provide consistent coding performance gains across a wide range of test settings. Ching-Han Chiang, Jingning Han, Yaowu Xu |
ICASSP | 3 |
| 2017 | Adaptive interpolation filter scheme in AV1abstractVideo codecs heavily depend on sub-pixel level motion compensation to achieve superior compression performance. Interpolation filters with both anti-aliasing and denoising properties play a critical role in producing high quality prediction at sub-pixel positions. Prior research has developed many adaptive filtering schemes to improve the prediction precision for compression gains. On the other hand, such filtering operations require intense computation and may lead to scattered cache footprints, therefore, account for a major portion of the overall decoding cost in both software and hardware implementations. An adaptive interpolation filtering scheme is proposed in this work to optimize the trade off between prediction quality and decoding performance. It employs a separable model and selects filter kernels independently for horizontal and vertical directions to better capture statistical variations. In order to obtain sharper transition and reduce the ripple effect in the passband in frequency domain, a 12-tap filter is introduced in conjunction with a complimentary operation design that minimizes its impact on the decoding performance. The scheme achieves on average 1.3% coding gains across a wide range of test settings, with fairly limited additional hardware cost. Ching-Han Chiang, Jingning Han, Stan Vitvitskyy, Debargha Mukherjee, Yaowu Xu |
ICIP | 5 |
| 2017 | A level-map approach to transform coefficient codingabstractTransform coding is widely used in the video and image codec to largely remove the spatial correlation. The magnitude of transform coefficient is weakly correlated to a number of factors, including its frequency band, the neighboring coefficient magnitudes, luma/chroma planes, etc. To exploit such correlations for efficient entropy coding, one would build a probability model conditioned on the available contexts. However, the interaction of these factors creates a high dimensional space, a direct use of which would easily fall into the over-fitting problem. How to construct a compact context set which effectively captures the underlying correlations remains a major challenge in video and image compression. Prior research work primarily relies on bucketizing the previously coded coefficients into a small number of categories as the context model for next coefficient. Certain information loss is inevitable due to the classification process. To fully exploit the available context in a limited model space, a level map approach is proposed in this work. It decomposes the coding of coefficient magnitudes into consecutive runs of binary map coding, each corresponds to whether a coefficient is equal to or greater than the given level. Under the Markov assumption across the levels, nearly all the reference symbols available to each level map can be approximated as binary random variables. It hence allows the context model to account for all the surrounding coefficients information provided by the lower level maps, while retaining a reasonably compact size. Experimental evidence demonstrates that the proposed coding scheme provides considerable compression performance gains consistently over a large test settings. Jingning Han, Ching-Han Chiang, Yaowu Xu |
ICIP | 3 |
| 2016 | A staircase transform coding scheme for screen content video codingabstractDemand for screen content videos that contain computer generated text and graphics is growing. They are very different from natural videos, because they include much sharper edge transitions and very repetitive patterns. On this type of material, the efficacy of the conventional discrete cosine transform (DCT) is questionable because it relies on the assumption that a Gauss-Markov model leads to a base-band signal. However, the assumption may not hold true for screen content material. This work exploits a class of staircase transforms. Unlike the DCT whose bases are samplings of sinusoidal functions, the staircase transforms have their bases sampled from staircase functions, which better approximate the sharp transitions often encountered in the context of screen content. The staircase transform is integrated into a hybrid transform coding scheme, in conjunction with DCT. It is experimentally shown that the proposed approach provides an average of 2.9% compression performance gains in terms of BD-rate reduction. A perceptual comparison further demonstrates that the use of staircase transform achieves substantial reduction in ringing artifact due to the Gibbs phenomenon. Jingning Han, Yaowu Xu, Jim Bankoski |
ICIP | 3 |
| 2016 | A dynamic motion vector referencing scheme for video codingabstractVideo codecs exploit temporal redundancy in video signals, through the use of motion compensated prediction, to achieve superior compression performance. The coding of motion vectors takes a large portion of the total rate cost. Prior research utilizes the spatial and temporal correlation of the motion field to improve the coding efficiency of the motion information. It typically constructs a candidate pool composed of a fixed number of reference motion vectors and allows the codec to select and reuse the one that best approximates the motion of the current block. This largely disconnects the entropy coding process from the block's motion information, and throws out any information related to motion consistency, leading to sub-optimal coding performance. An alternative motion vector referencing scheme is proposed in this work to fully accommodate the dynamic nature of the motion field. It adaptively extends or shortens the candidate list according to the actual number of available reference motion vectors. The associated probability model accounts for the likelihood that an individual motion vector candidate is used. A complementary motion vector candidate ranking system is also presented here. It is experimentally shown that the proposed scheme achieves about 1.6% compression performance gains on a wide range of test clips. Jingning Han, Yaowu Xu, Jim Bankoski |
ICIP | 2 |
| 2015 | An estimation-theoretic approach to video denoiseingabstractA novel denoising scheme is proposed to fully exploit the spatio-temporal correlations of the video signal for efficient enhancement. Unlike conventional pixel domain approaches that directly connect motion compensated reference pixels and spatially neighboring pixels to build statistical models for noise filtering, this work first removes spatial correlations by applying transformations to both pixel blocks and performs estimation in the frequency domain. It is premised on the realization that the precise nature of temporal dependencies, which is entirely masked in the pixel domain by the statistics of the dominant low frequency components, emerges after signal decomposition and varies considerably across the spectrum. We derive an optimal non-linear estimator that accounts for both motion compensated reference and the noisy observations to resemble the original video signal per transform coefficient. It departs from other transform domain approaches that employ linear filters over a sizable reference set to reduce the uncertainty due to the random noise term. Instead it jointly exploits this precise statistical property appeared in the transform domain and the noise probability model in an estimation-theoretic framework that works on a compact support region. Experimental results provide evidence for substantial denoising performance improvement. Jingning Han, Timothy Kopp, Yaowu Xu |
ICIP | 3 |
| 2013 | A butterfly structured design of the hybrid transform coding schemeabstractThe hybrid transform coding scheme that alternates amongst the asymmetric discrete sine transform (ADST) and the discrete cosine transform (DCT) depending on the boundary prediction conditions, is an efficient tool for video and image compression. It optimally exploits the statistical characteristics of prediction residual, thereby achieving significant coding performance gains over the conventional DCT-based approach. A practical concern lies in the intrinsic conflict between transform kernels of ADST and DCT, which prevents a butterfly structured implementation for parallel computing. Hence the hybrid transform coding scheme has to rely on matrix multiplication, which presents a speed-up barrier due to under-utilization of the hardware, especially for larger block sizes. In this work, we devise a novel ADST-like transform whose kernel is consistent with that of DCT, thereby enabling butterfly structured computation flow, while largely retaining the performance advantages of hybrid transform coding scheme in terms of compression efficiency. A prototype implementation of the proposed butterfly structured hybrid transform coding scheme is available in the VP9 codec repository. Jingning Han, Yaowu Xu, Debargha Mukherjee |
PCS | 2 |
| 2013 | The latest open-source video codec VP9 - An overview and preliminary resultsabstractGoogle has recently finalized a next generation open-source video codec called VP9, as part of the libvpx repository of the WebM project (http://www.webmproject.org/). Starting from the VP8 video codec released by Google in 2010 as the baseline, various enhancements and new tools were added, resulting in the next-generation VP9 bit-stream. This paper provides a brief technical overview of VP9 along with comparisons with other state-of-the-art video codecs H.264/AVC and HEVC on standard test sets. Results show VP9 to be quite competitive with mainstream state-of-the-art codecs. Debargha Mukherjee, Jim Bankoski, Adrian Grange, Jingning Han, John Koleszar, Paul Wilkins, Yaowu Xu, Ronald Bultje |
PCS | 7 |
| 2011 | Technical overview of VP8, an open source video codec for the webabstractVP8 is an open source video compression format supported by a consortium of technology companies. This paper provides a technical overview of the format, with an emphasis on its unique features. The paper also discusses how these features benefit VP8 in achieving high compression efficiency and low decoding complexity at the same time. Jim Bankoski, Paul Wilkins, Yaowu Xu |
ICME | 3 |
| 2005 | Object Recognition by Partial Shape Matching Guided SearchabstractWe retrieve samples from large image/video databases by means of learning by example. We propose a fast partial shape matching guided 2D object recognition algorithm to significantly accelerate recognition/matching of full (non-occluded) or partially occluded objects. The significant increase in speed comes from the fact that the search space is reduced to only those combinations of regions in the neighborhood of potential partial matches, as opposed to all combinations of regions as was done in our prior work (Xu et al. (2003)). Theoretical calculations and experimental results are provided to demonstrate the effectiveness of the proposed algorithm on real images. Eli Saber, Yaowu Xu, A. Murat Tekalp |
ICASSP (2) | 2 |
| 2005 | Partial shape recognition by sub-matrix matching for partial matching guided image labeling
Eli Saber, Yaowu Xu, A. Murat Tekalp |
Pattern Recognit. | 2 |
| 2004 | Semantic object segmentation by dynamic learning from multiple examplesabstractWe present a novel "dynamic learning" approach for an intelligent image database system to automatically improve object segmentation and labeling without user intervention, as new examples become available, for object-based indexing. The proposed approach is an extension of our earlier work on "learning by example", which addressed labeling of similar objects in a set of database images based on a single example (Saber et al. (2003)). It utilizes multiple example object templates to improve the accuracy of existing object segmentations and labels. We also propose to use Normalized Area of Symmetric Differences (NASD) as the similarity metric in "dynamic learning", due to its robustness to boundary noise that results from automatic image segmentation. The performance of the dynamic learning concept is demonstrated by experimental results. Yaowu Xu, Eli Saber, A. Murat Tekalp |
ICASSP (3) | 1 |
| 2004 | Dynamic learning from multiple examples for semantic object segmentation and search
Yaowu Xu, Eli Saber, A. Murat Tekalp |
Comput. Vis. Image Underst. | 1 |
| 2003 | Object-based image labeling through learning by example and multi-level segmentation
Yaowu Xu, Pinar Duygulu, Eli Saber, A. Murat Tekalp, Fatos T. Yarman-Vural |
Pattern Recognit. | 1 |
| 2003 | Object segmentation and labeling by learning from examplesabstractWe propose a system that employs low-level image segmentation followed by color and two-dimensional (2-D) shape matching to automatically group those low-level segments into objects based on their similarity to a set of example object templates presented by the user. A hierarchical content tree data structure is used for each database image to store matching combinations of low-level regions as objects. The system automatically initializes the content tree with only "elementary nodes" representing homogeneous low-level regions. The "learning" phase refers to labeling of combinations of low-level regions that have resulted in successful color and/or 2-D shape matches with the example template(s). These combinations are labeled as "object nodes" in the hierarchical content tree. Once learning is performed, the speed of second-time retrieval of learned objects in the database increases significantly. The learning step can be performed off-line provided that example objects are given in the form of user interest profiles. Experimental results are presented to demonstrate the effectiveness of the proposed system with hierarchical content tree representation and learning by color and 2-D shape matching on collections of car and face images. Yaowu Xu, Eli Saber, A. Murat Tekalp |
IEEE Trans. Image Process. | 1 |
| 2000 | Object based image retrieval based on multi-level segmentationabstractCurrently, image retrieval systems are based on low-level features of color, texture and shape, not on the semantic descriptions that are common to humans, such as objects, people, and place. In order to narrow down the gap between the low level and semantic level, object-based content analysis, which segments the semantically meaningful objects of images, is an essential step. In this study, we propose a learning process in order to perform effective automatic off-line analysis on a multi-level segmented image stack. Meaningful objects are extracted given certain user search patterns and interest profiles. Color and/or shape information of the objects is stored in the hierarchical content representations of the images. This information is utilized by a hierarchical matching scheme to improve the retrieval speed in the subsequent searches. Yaowu Xu, Pinar Duygulu, Eli Saber, A. Murat Tekalp, Fatos T. Yarman-Vural |
ICASSP | 1 |
| 2000 | Image Retrieval Through Shape Matching of Partially Occluded Objects Using Hierarchical Content DescriptionabstractThis paper proposes a new contour based shape matching approach capable of recognizing partially occluded objects in images. The process can be divided into the following steps: region formation from low level color image segmentation, B-spline filtering and feature point extraction, correspondence determination, least square estimation of affine transformation parameters, and similarity measuring. Once similarity is established, the information is retained using a hierarchical content description scheme, enabling expedient object based image retrieval at a later time. Yaowu Xu, Eli Saber, A. Murat Tekalp |
ICIP | 1 |
| 1999 | Object Formation by Learning in Visual Databases Using Hierarchical Content DescriptionabstractThis paper proposes a self-learning content-based image indexing and retrieval system that employs a hierarchical content representation (consisting of objects and regions) and a hierarchical content matching method for effective and efficient image/object retrieval. The “learning” behavior is enabled by our proposed hierarchical content representation which allows easy storage of combinations of regions that have resulted in successful matches to objects of interest as determined by user search patterns and profiles. The learning step effectively performs an automatic off-line analysis of database images into meaningful objects. Once the learning phase is complete, the speed of shape based retrieval of the learned objects in the database increases significantly. Experimental results are presented to show the effectiveness of the proposed hierarchical content representation, hierarchical matching, and the learning behavior on collections of car images. Yaowu Xu, Eli Saber, A. Murat Tekalp |
ICIP (2) | 1 |