Onur G. Guleryuz

dblp:52/4220 · DBLP profile ↗
← Back
50ranked-venue papers
22as first author
9since 2021 · last 2025
0009-0007-8637-7181ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 20 first-author · 9 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 1 since 2021Theory of computation · 3 · 2 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Joint Optimization of Primary and Secondary Transforms Using Rate-Distortion Optimized Transform Design
abstract
Data-dependent transforms are increasingly being incorporated into next-generation video coding systems such as AVM, a codec under development by the Alliance for Open Media (AOM), and VVC. To circumvent the computational complexities associated with implementing non-separable data-dependent transforms, combinations of separable primary transforms and non-separable secondary transforms have been studied and integrated into video coding standards. These codecs often utilize rate-distortion optimized transforms (RDOT) to ensure that the new transforms complement existing transforms like the DCT and the ADST. In this work, we propose an optimization framework for jointly designing primary and secondary transforms from data through a rate-distortion optimized clustering. Primary transforms are assumed to follow a path-graph model, while secondary transforms are non-separable. We empirically evaluate our proposed approach using AVM residual data and demonstrate that 1) the joint clustering method achieves lower total RD cost in the RDOT design framework, and 2) jointly optimized separable path-graph transforms (SPGT) provide better coding efficiency compared to separable KLTs obtained from the same data.
Darukeesan Pakiyarajah, Eduardo Pavez, Antonio Ortega, Debargha Mukherjee, Onur G. Guleryuz, Keng-Shih Lu, Kruthika Koratti Sivakumar
ICIP5
2024 Standard Compatible Efficient Video Coding with Jointly Optimized Neural Wrappers
abstract
We present a standard-compatible video coding scheme with end-to-end optimized neural wrapper over standard video codecs that achieves significant rate-distortion (R-D) performance gains and is still efficient in decoding. We train a pair of pre- and post-processor using a differential JPEG proxy. The pre-processor applies a learned transform to the video and downsamples the video by a factor of 2. It generates a bottleneck video to be coded by a standard codec as a YUV sequence. The post-processor takes the decoded bottleneck video, does the inverse transform, and upsamples it to the original resolution. We follow the design in [1] , where we configure downsample using a layer of strided convolution. We optimize the post-processor for efficiency by replacing convolutions with kernel size larger than 1×1 to depth-wise convolutions [2] .
Yueyu Hu, Onur G. Guleryuz, Debargha Mukherjee, Yao Wang 0001
DCC3
2024 Standard Compliant Video Coding Using Low Complexity, Switchable Neural Wrappers
abstract
The proliferation of high resolution videos posts great storage and bandwidth pressure on cloud video services, driving the development of next-generation video codecs. Despite great progress made in neural video coding, existing approaches are still far from economical deployment considering the complexity and rate-distortion performance tradeoff. To clear the roadblocks for neural video coding, in this paper we propose a new framework featuring standard compatibility, high performance, and low decoding complexity. We employ a set of jointly optimized neural pre and post-processors, wrapping a standard video codec, to encode videos at different resolutions. The rate-distorion optimal downsampling ratio is signaled to the decoder at the per-sequence level for each target rate. We design a low complexity neural post-processor architecture that can handle different upsampling ratios. The change of resolution exploits the spatial redundancy in high-resolution videos, while the neural wrapper further achieves rate-distortion performance improvement through end-to-end optimization with a codec proxy. Our light-weight post-processor architecture has a complexity of 516 MACs / pixel, and achieves 9.3% BD-Rate reduction over VVC on the UVG dataset, and $6.4 \%$ on AOM CTC Class A1. Our approach has the potential to further advance the performance of the latest video coding standards using neural processing with minimal added complexity.
Yueyu Hu, Onur G. Guleryuz, Debargha Mukherjee, Yao Wang 0001
ICIP3
2023 Sandwiched Video Compression: Efficiently Extending the Reach of Standard Codecs with Neural Wrappers
abstract
We propose sandwiched video compression – a video compression system that wraps neural networks around a standard video codec. The sandwich framework consists of a neural pre- and post-processor with a standard video codec between them. The networks are trained jointly to optimize a rate-distortion loss function with the goal of significantly improving over the standard codec in various compression scenarios. End-to-end training in this setting requires a differentiable proxy for the standard video codec, which incorporates temporal processing with motion compensation, inter/intra mode decisions, and in-loop filtering. We propose differentiable approximations to key video codec components and demonstrate that, in addition to providing meaningful compression improvements over the standard codec, the neural codes of the sandwich lead to significantly better rate-distortion performance in two important scenarios. When transporting high-resolution video via low-resolution HEVC, the sandwich system obtains 6.5 dB improvements over standard HEVC. More importantly, using the well-known perceptual similarity metric, LPIPS, we observe 30% improvements in rate at the same quality over HEVC. Last but not least, we show that pre- and post-processors formed by very modestly-parameterized, light-weight networks can closely approximate these results.
Berivan Isik, Onur G. Guleryuz, Danhang Tang, Jonathan Taylor 0001, Philip A. Chou
ICIP2
2022 Non-Separable Filtering with Side-Information and Contextually-Designed Filters for Next Generation Video Codecs
abstract
We propose two new in-loop filtering tools to augment the loop restoration process in AV1. Targeting challenging future compression scenarios that will be faced by the emerging AV2 standard, the proposed tools are designed to improve performance at both high and low bit-rates. Our non-separable Wiener filtering proposal aims to increase quality especially over directional features and textures in the decoded picture with the help of side-information at higher rates. The proposed pixel-adaptive filter complements at lower rates by improving performance without side-information relying solely on finely characterized pixel contexts. Simulation results show the rate-distortion efficacy of the proposed tools.
Onur G. Guleryuz, Debargha Mukherjee, Yue Chen 0040, Keng-Shih Lu, Urvang Joshi
ICIP1
2022 Switchable CNN-Based Same-Resolution and Super-Resolution In-Loop Restoration for Next Generation Video Codecs
abstract
We present a common framework for in-loop same-resolution and super-resolution restoration for incorporation into a next-generation video codec. Building on the in-loop filtering pipeline in the AV1 video codec from the Alliance for Open Media (AOM), we first enhance it to support symmetric spatial down-up scaling with better down and upscaling filters, followed by adding switchable frame-level CNNs to restore the reconstructed frames. Furthermore, the architectures for the CNNs used are constrained to be very simple with a relatively small number of parameters and multiply-add operations per decoded pixel to make them practically feasible in hardware. Preliminary results are presented on test sets being used by AOM for their ongoing next-generation video codec (AV2) development effort.
Urvang Joshi, Yue Chen 0040, Debargha Mukherjee, Onur G. Guleryuz, Shan Li 0001, In Suk Chong
ICIP4
2022 Intra Prediction of Regular and Near-Regular Textures Via Graph-Based Inpainting
abstract
Intra prediction is an important technique to improve coding efficiency by exploiting the spatial redundancy present in typical video sequences. In video coding standards such as H.264/AVC, HEVC and VVC, directional predictors are utilized to generate prediction along a single direction within a block to be coded. However, these predictors fail to generate an accurate prediction when the block contains complex patterns such as periodic textures. In this paper, we propose a graph-based inpainting method that can handle both regular and near-regular textures. The proposed inpainting method utilizes a total variation model associated with the Laplacian matrix of a graph, whose edge weights are a function of pixel patch distance. We evaluate the performance of our proposed method as an additional prediction mode combined with the H.264/AVC coding standard. Experimental results show that the proposed method can significantly outperform H.264/AVC predictors in areas with high frequency periodic patterns.
Wen-Yang Lu, Eduardo Pavez, Antonio Ortega, Debargha Mukherjee, Onur G. Guleryuz, Keng-Shih Lu
ICIP5
2022 Sandwiched Image Compression: Increasing the resolution and dynamic range of standard codecs
abstract
Given a standard image codec, we compress images that may have higher resolution and/or higher bit depth than allowed in the codec's specifications, by sandwiching the standard codec between a neural pre-processor (before the standard encoder) and a neural post-processor (after the standard decoder). Using a differentiable proxy for the the standard codec, we design the neural pre-and post-processors to transport the high resolution (super-resolution, SR) or high bit depth (high dynamic range, HDR) images as lower resolution and lower bit depth images. The neural processors accomplish this with spatially coded modulation, which acts as watermarks to preserve the important image detail during compression. Experiments show that compared to conventional methods of transmitting high resolution or high bit depth through lower resolution or lower bit depth codecs, our sandwich architecture gains ~9dB for SR images and $\sim$3dB for HDR images at the same rate over large test sets. We also observe significant gains in visual quality.
Onur G. Guleryuz, Philip A. Chou, Hugues Hoppe, Danhang Tang, Ruofei Du, Philip Davidson, Sean Ryan Fanello
PCS1
2021 Sandwiched Image Compression: Wrapping Neural Networks Around A Standard Codec
abstract
We sandwich a standard image codec between two neural networks: a preprocessor that outputs neural codes, and a postprocessor that reconstructs the image. The neural codes are compressed as ordinary images by the standard codec. Using differentiable proxies for both rate and distortion, we develop a rate-distortion optimization framework that trains the networks to generate neural codes that are efficiently compressible as images. This architecture not only improves rate-distortion performance for ordinary RGB images, but also enables efficient compression of alternative image types (such as normal maps of computer graphics) using standard image codecs. Results demonstrate the effectiveness and flexibility of neural processing in mapping a variety of input data modalities to the rigid structure of standard codecs. A surprising result is that the rate-distortion-optimized neural processing seamlessly learns to transport color images using a single-channel (grayscale) codec.
Onur G. Guleryuz, Philip A. Chou, Hugues Hoppe, Danhang Tang, Ruofei Du, Philip Davidson, Sean Ryan Fanello
ICIP1
2020 Deep Implicit Volume Compression
abstract
We describe a novel approach for compressing truncated signed distance fields (TSDF) stored in 3D voxel grids, and their corresponding textures. To compress the TSDF, our method relies on a block-based neural network architecture trained end-to-end, achieving state-of-the-art rate-distortion trade-off. To prevent topological errors, we losslessly com- press the signs of the TSDF, which also upper bounds the reconstruction error by the voxel size. To compress the corresponding texture, we designed a fast block-based UV parameterization, generating coherent texture maps that can be effectively compressed using existing video compression algorithms. We demonstrate the performance of our algo- rithms on two 4D performance capture datasets, reducing bitrate by 66% for the same distortion, or alternatively re- ducing the distortion by 50% for the same bitrate, compared to the state-of-the-art.
Danhang Tang, Philip A. Chou, Christian Häne, Mingsong Dou, Sean Ryan Fanello, Jonathan Taylor 0001, Philip L. Davidson, Onur G. Guleryuz, Yinda Zhang 0001, Shahram Izadi, Andrea Tagliasacchi, Sofien Bouaziz, Cem Keskin
CVPR9
2018 Fast Lifting for 3D Hand Pose Estimation in AR/VR Applications
abstract
We introduce a simple model for the human hand skeleton that is geared toward estimating 3D hand poses from 2D keypoints. The estimation problem arises in AR/VR scenarios where low-cost cameras are used to generate 2D views through which rich interactions with the world are desired. Starting with a noisy set of 2D hand keypoints (camera-plane coordinates of detected joints of the hand), the proposed algorithm generates 3D keypoints that are (i) compliant with human hand skeleton constraints and (ii) perspective-project down to the given 2D keypoints. Our work considers the 2D to 3D lifting problem algebraically, identifies the parts of the hand that can be lifted accurately, points out the parts that may lead to ambiguities, and proposes remedies for ambiguous cases. Most importantly, we show that the finger-tip localization errors are a good proxy for the errors at other finger joints. This observation leads to a look-up-table-based formulation that instantaneously determines finger poses without solving constrained trigonometric problems. The result is a fast algorithm running super real-time on a single core. When hand bone-lengths are unknown our technique estimates these and allows smooth AR/VR sessions where a user's hand is automatically estimated in the beginning and the rest of the session seamlessly continued. Our work provides accurate 3D results that are competitive with the state-of-the-art without requiring any 3D training data.
Onur G. Guleryuz, Christine Kaeser-Chen
ICIP1
2018 Depth from motion for smartphone AR
abstract
Augmented reality (AR) for smartphones has matured from a technology for earlier adopters, available only on select high-end phones, to one that is truly available to the general public. One of the key breakthroughs has been in low-compute methods for six degree of freedom (6DoF) tracking on phones using only the existing hardware (camera and inertial sensors). 6DoF tracking is the cornerstone of smartphone AR allowing virtual content to be precisely locked on top of the real world. However, to really give users the impression of believable AR, one requires mobile depth. Without depth, even simple effects such as a virtual object being correctly occluded by the real-world is impossible. However, requiring a mobile depth sensor would severely restrict the access to such features. In this article, we provide a novel pipeline for mobile depth that supports a wide array of mobile phones, and uses only the existing monocular color sensor. Through several technical contributions, we provide the ability to compute low latency dense depth maps using only a single CPU core of a wide range of (medium-high) mobile phones. We demonstrate the capabilities of our approach on high-level AR applications including real-time navigation and shopping.
Julien P. C. Valentin, Adarsh Kowdle, Jonathan T. Barron, Neal Wadhwa, Maksym Dzitsiuk, Michael Schoenberg, Ambrus Csaszar, Eric Turner 0001, Ivan Dryanovski, João Afonso, Jose Pascoal, Konstantine Tsotsos, Mira Leung, Mirko Schmidt, Onur G. Guleryuz, Sameh Khamis, Vladimir Tankovich, Sean Ryan Fanello, Shahram Izadi, Christoph Rhemann
ACM Trans. Graph.16
2017 Layered-givens transforms: Tunable complexity, high-performance approximation of optimal non-separable transforms
abstract
We introduce layered-Givens transforms (LGTs) which are arbitrary-dimensional, tunable-complexity, orthonormal transforms of data. LGTs are formed by layers of data permutations and Givens rotations with the number of layers controlling the overall transform's computational complexity. We propose a novel method for the design of layered-Givens transforms that approximate desired complex transforms (such as KLTs, SOTs, etc.) so that most of the performance of the approximated transform is retained at significantly reduced complexity. Key to the LGT performance is the choice of the permutations that arrange the data prior to the Givens rotations. Despite the very highly combinatorial nature of the LGT design problem, we provide an algorithm that finds the optimal parameters (permutation and Givens rotations) for each layer. Our results show that designed LGTs closely approximate desired non-separable transforms at significantly reduced complexity.
Onur G. Guleryuz, Jana Ehmann, Arash Vosoughi
ICIP2
2016 An optimization framework for combining multiple graphs
abstract
This paper introduces a novel framework for combining multiple weighted graphs into a single optimized weighted graph. In our framework, we first develop a statistical formulation for the graph combining problem with a maximum likelihood criterion, and derive its optimality conditions. We then use these conditions to formulate the deterministic graph combining problem and propose a solution. Our experimental results show that the proposed solution provides better modeling compared to the commonly used averaging method. The introduced framework has various applications in signal processing and machine learning.
Hilmi E. Egilmez, Antonio Ortega, Onur G. Guleryuz, Jana Ehmann, Sehoon Yea
ICASSP3
2016 Row-column transforms: Low-complexity approximation of optimal non-separable transforms
abstract
This paper introduces row-column transforms (RCTs) which are 2D non-separable transforms defined with the aid of a set of 1-D linear transforms and a basis ordering permutation. We propose a novel method for the design of row-column transforms that approximate desired complex transforms (such as KLTs, SOTs, etc.) so that most of the performance of the approximated transforms is retained at significantly reduced complexity. Given a non-separable block transform of interest, our method designs an RCT by (i) optimizing a set of 1-D transforms applied to rows and columns of the signal block and (ii) finding the best transform coefficient ordering permutation. Our experimental results show that optimized RCTs closely approach the compression performance of the desired non-separable transforms while retaining the computational complexity of separable transforms.
Hilmi E. Egilmez, Onur G. Guleryuz, Jana Ehmann, Sehoon Yea
ICIP2
2016 Transform-coded pel-recursive video compression
abstract
We propose an algorithm that accomplishes transform-coded, spatiotemporal, pel-recursive video compression. Traditional pel-recursive coders obtain sophisticated spatio-temporal predictions for the current pixel based on previously decoded data. The resulting per-pixel prediction errors are encoded independently so that the decoder can use previously-encoded pixels in the prediction of the current. It is well-known that pel-recursive coders significantly under-perform modern hybrid coders which use comparatively very simple predictors but transform code the prediction errors. Our algorithm combines the accurate predictors of pel-recursive techniques with transform coders so that the strengths of both approaches can be taken advantage of. In the proposed work, the decoder transform decodes the residuals for a block and then acts like a simple pel-recursive decoder for the pixels of that block. We show that the proposed algorithm can be seen as the use of a pel-recursion enabling transform at the encoder, which generates and encodes the correct transform coefficients to avoid error propagation at the decoder. A straightforward implementation of our work (implemented as an extension of HEVC that preserves independent decodability of INTER blocks) shows compression improvements over the baseline HEVC.
Jana Ehmann, Onur G. Guleryuz, Sehoon Yea
ICIP2
2015 Reduced-rank condensed filter dictionaries for inter-picture prediction
abstract
We consider the motion-compensated temporal prediction loop at the heart of modern video coders. Rather than using motion-compensated reference frame blocks directly as predictors, we incorporate their spatially-filtered versions into the prediction loop. We design adaptive filters that are geared toward successful prediction over sophisticated temporal evolutions involving lighting changes, focus changes, structured noise, and so on. The spatially and temporally varying nature of such video evolutions requires the learning and transmission of many filters, necessitating parameter reduction for compression and related applications. Unlike earlier work that tries to limit parameters by using a small set of general filters, or by restricting to symmetric filters, etc., we propose a novel parametrization of filters in terms of a set of base-filter kernels and modulation weights. Given a filter dictionary of K-tap filters, our work can be seen as providing a reduced-rank, prediction-optimal approximation of this dictionary that represents its filters with K' ≪ K parameters.
Shunyao Li, Onur G. Guleryuz, Sehoon Yea
ICASSP2
2015 Approximation and Compression With Sparse Orthonormal Transforms
abstract
We propose a new transform design method that targets the generation of compression-optimized transforms for next-generation multimedia applications. The fundamental idea behind transform compression is to exploit regularity within signals such that redundancy is minimized subject to a fidelity cost. Multimedia signals, in particular images and video, are well known to contain a diverse set of localized structures, leading to many different types of regularity and to nonstationary signal statistics. The proposed method designs sparse orthonormal transforms (SOTs) that automatically exploit regularity over different signal structures and provides an adaptation method that determines the best representation over localized regions. Unlike earlier work that is motivated by linear approximation constructs and model-based designs that are limited to specific types of signal regularity, our work uses general nonlinear approximation ideas and a data-driven setup to significantly broaden its reach. We show that our SOT designs provide a safe and principled extension of the Karhunen-Loeve transform (KLT) by reducing to the KLT on Gaussian processes and by automatically exploiting non-Gaussian statistics to significantly improve over the KLT on more general processes. We provide an algebraic optimization framework that generates optimized designs for any desired transform structure (multiresolution, block, lapped, and so on) with significantly better n -term approximation performance. For each structure, we propose a new prototype codec and test over a database of images. Simulation results show consistent increase in compression and approximation performance compared with conventional methods.
Osman Gokhan Sezer, Onur G. Guleryuz, Yücel Altunbasak
IEEE Trans. Image Process.2
2014 Non-causal encoding of predictively coded samples
abstract
We study the INTRA-prediction operation employed in hybrid video coders where sample values to be compressed are predicted using previously-decoded context values and the prediction errors are transform coded. Given the context values this procedure can be argued to be rate-distortion optimal for Gaussian signals. Yet natural images and video contain many structures that do not readily fit into Gaussian signal assumptions. Targeting such structures we propose a technique that predicts each sample using the context values and samples that are jointly transform coded with the predicted sample. We show that this joint, non-causal encoding can be represented with a nonorthogonal transform whose form and parameters we derive. We augment the HEVC standard with our work and show significant compression improvements on images/video that contain directional structures.
Onur G. Guleryuz, Amir Said, Sehoon Yea
ICIP1
2014 Improving hybrid coding via control of quantization errors in the spatial and frequency domains
abstract
We propose new techniques to improve hybrid coding of images and video, without increasing decoder complexity, by controlling quantization effects in spatial and frequency domains, and finding optimal quantized coefficients via optimization. Additional compression is achieved by finding sets of optimal parameters, determined using training techniques. Experimental results confirm that the new method is able to provide better compression, both in objective measures, like signal-to-noise ratios, and also subjective, by reducing visually annoying blocking and banding artifacts.
Amir Said, Onur G. Guleryuz, Sehoon Yea
ICIP2
2012 Visual conditioning for augmented-reality-assisted video conferencing
abstract
Typical video conferencing scenarios bring together individuals from disparate environments. Unless one commits to expensive tele-presence rooms, conferences involving many individuals result in a cacophony of visuals and backgrounds. Ideally one would like to separate participant visuals from their respective environments and render them over visually pleasing backgrounds that enhance immersion for all. Yet available image/video segmentation techniques are limited and result in significant artifacts even with recently popular commodity depth sensors. In this paper we present a technique that accomplishes robust and visually pleasing rendering of segmented participants over adaptively-designed virtual backgrounds. Our method works by determining virtual backgrounds that match and highlight participant visuals and uses directional textures to hide segmentation artifacts due to noisy segmentation boundaries, missing regions, etc. Taking advantage of simple computations and look-up-tables, our work leads to fast, real-time implementations that can run on mobile and other computationally-limited platforms.
Onur G. Guleryuz, Ton Kalker
MMSP1
2011 Spatial Sparsity-Induced Prediction (SIP) for Images and Video: A Simple Way to Reject Structured Interference
abstract
We propose a prediction technique that is geared toward forming successful estimates of a signal based on a correlated anchor signal that is contaminated with complex interference. The corruption in the anchor signal involves intensity modulations, linear distortions, structured interference, clutter, and noise just to name a few. The proposed setup reflects nontrivial prediction scenarios involving images and video frames where statistically related data is rendered ineffective for traditional methods due to cross-fades, blends, clutter, brightness variations, focus changes, and other complex transitions. Rather than trying to solve a difficult estimation problem involving nonstationary signal statistics, we obtain simple predictors in linear transform domain where the underlying signals are assumed to be sparse. We show that these simple predictors achieve surprisingly good performance and seamlessly allow successful predictions even under complicated cases. None of the interference parameters are estimated as our algorithm provides completely blind and automated operation. We provide a general formulation that allows for nonlinearities in the prediction loop and we consider prediction optimal decompositions. Beyond an extensive set of results on prediction and registration, the proposed method is also implemented to operate inside a state-of-the-art compression codec and results show significant improvements on scenes that are difficult to encode using traditional prediction techniques.
Gang Hua 0003, Onur G. Guleryuz
IEEE Trans. Image Process.2
2010 A sparsity-distortion-optimized multiscale representation of geometry
abstract
This paper describes the construction of a new multiresolutional decomposition with applications to image compression. The proposed method designs sparsity-distortion-optimized orthonormal transforms applied in wavelet domain to arrive at a multiresolutional representation which we term the Sparse Multiresolutional Transform (SMT). Our optimization operates over sub-bands of given orientation and exploits the inter-scale and intra-scale dependencies of wavelet co-efficients over image singularities. The resulting SMT is substantially sparser than the wavelet transform and leads to compaction that can be exploited by well-known coefficient coders. Our construction deviates from the literature, which mainly focuses on model-based methods, by offering a data-driven optimization of wavelet representations. Simulation experiments show that the proposed method consistently offers better performance compared to the original wavelet-representation and can reach up to 1dB improvements within state-of-the-art coefficient coders.
Osman Gokhan Sezer, Yücel Altunbasak, Onur G. Guleryuz
ICIP3
2009 Complexity regularized pattern matching
abstract
We propose a technique that finds optimized descriptors for pattern matching applications. We formulate the pattern matching problem as the search of a pattern library for vectors defined in a query manifold. Our approach trades off the computational complexity involved in the search with matching accuracy by representing the query manifold with its complexity-dependent approximations. This is done in an optimal way so that a user with a given complexity budget accomplishes the optimal matching performance for that budget. Our work can be seen as defining a covering around the query manifold with the aid of the derived descriptors. The higher the allowed computational complexity, the tighter the covering, and the more accurate the match. Our formulation results in sparse descriptors which naturally emerge as the optimal solutions. The proposed descriptors are adaptively optimized for the particular search problem so that application-specific simplifications are taken full advantage of. Thanks to our algebraic approach, the presented formulation is general and can readily be applied to many different types of signals in addition to images and video.
Jana Zujovic, Onur G. Guleryuz
ICIP2
2008 On optimal royalty costs for video compression
abstract
Modern video compression includes a mature set of codecs, tools, and techniques that correspond to a variety of coding efficiencies and royalty costs in video delivery. In early work we showed that there can be significant benefits to employing a royalty-aware encoder that jointly optimizes rate, distortion, and royalty cost by choosing among several codecs having different royalty costs [1]. In this paper we assume the existence of such an encoder and concentrate on how royalty costs should be assigned to benefit a system of users, video service providers, and tool/codec owners. Our work operates on the intersection of rate-distortion principles with basic concepts from economics and tries to address issues that are rapidly becoming very relevant in media delivery. Using a simple model that captures user preferences of video quality levels (given costs associated with these levels), we formulate the assignment of royalty costs as an optimization problem and solve it for optimal costs for several scenarios. The main trade-offs that govern our results are the cost of bandwidth and previously unachievable quality levels (if any) that a state-of-the-art codec gives access to. We show that content-based cost assignment not only increases user utilization but also the royalty earnings for tool owners compared to content-unaware pricing policies.
Ankur Saxena, Onur G. Guleryuz, M. Reha Civanlar
ICIP2
2008 Sparse orthonormal transforms for image compression
abstract
We propose a block-based transform optimization and associated image compression technique that exploits regularity along directional image singularities. Unlike established work, directionality comes about as a byproduct of the proposed optimization rather than a built in constraint. Our work classifies image blocks and uses transforms that are optimal for each class, thereby decomposing image information into classification and transform coefficient information. The transforms are optimized using a set of training images. Our algebraic framework allows straightforward extension to non-block transforms, allowing us to also design sparse lapped transforms that exploit geometric regularity. We use an EZW/SPIHT like entropy coder to encode the transform coefficients to show that our block and lapped designs have competitive rate-distortion performance. Our work can be seen as nonlinear approximation optimized transform coding of images subject to structural constraints on transform basis functions.
Osman Gokhan Sezer, Oztan Harmanci, Onur G. Guleryuz
ICIP3
2007 Spatial Sparsity Induced Temporal Prediction for Hybrid Video Compression
abstract
In this paper we propose a new motion compensated prediction technique that enables successful predictive encoding during fades, blended scenes, temporally decorrelated noise, and many other temporal evolutions which force predictors used in traditional hybrid video coders to fail. We model reference frame blocks to be used in motion compensated prediction as consisting of two superimposed parts: one part that is relevant for prediction and another part that is not relevant. By performing prediction in a domain where the video frames are spatially sparse, our work allows the automatic isolation of the prediction-relevant parts. These are then used to enable better prediction than would be possible otherwise. Our sparsity induced prediction algorithm (SIP) generates successful predictors by exploiting the non-convex structure of the sets that natural images and video frames lie in. Correctly determining this non-convexity through sparse representations allows better performance in hybrid video codecs equipped with the proposed work
Gang Hua 0003, Onur G. Guleryuz
DCC2
2007 Royalty Cost Based Optimization for Video Compression
abstract
A video compression standard incorporates many tools and technologies which must be licensed by systems that deploy the standard. The licensing determines the royalty costs that must be paid to the holders of intellectual property on the respective tools. With current abundance of well understood and effective video compression tools, one can imagine the formation of cross cutting tool libraries with tools drawn from different video compression standards. This allows dynamic selection from a large pool of tools, having potentially overlapping functionality, when encoding individual video sequences. In this paper we examine the royalty cost aspect of the scenario where video is encoded using a library of royalty bearing tools by considering encoding that jointly optimizes rate, distortion, and royalty cost. We provide a system that optimizes video delivery under various licensing conditions imposed on tool intellectual property. We present an example of royalty based encoding (using assumed royalty costs) to show the merit of the proposed framework.
Emrah Akyol, Onur G. Guleryuz, M. Reha Civanlar
ICIP (1)2
2007 Weighted Averaging for Denoising With Overcomplete Dictionaries
abstract
We consider the scenario where additive, independent, and identically distributed (i.i.d) noise in an image is removed using an overcomplete set of linear transforms and thresholding. Rather than the standard approach, where one obtains the denoised signal by ad hoc averaging of the denoised estimates provided by denoising with each of the transforms, we formulate the optimal combination as a conditional linear estimation problem and solve it for optimal estimates. Our approach is independent of the utilized transforms and the thresholding scheme, and as we illustrate using oracle-based denoisers, it extends established work by exploiting a separate degree of freedom that is, in general, not reachable using previous techniques. Our derivation of the optimal estimates specifically relies on the assumption that the utilized transforms provide sparse decompositions. At the same time, our work is robust as it does not require any assumptions about image statistics beyond sparsity. Unlike existing work, which tries to devise ever more sophisticated transforms and thresholding algorithms to deal with the myriad types of image singularities, our work uses basic tools to obtain very high performance on singularities by taking better advantage of the sparsity that surrounds them. With well-established transforms, we obtain results that are competitive with state-of-the-art methods.
Onur G. Guleryuz
IEEE Trans. Image Process.1
2006 Image Compression with a Geometrical Entropy Coder
abstract
In this paper we propose a geometrical entropy coder to compress images with geometry. Our proposed scheme works on quantized coefficients of the discrete wavelet transform (DWT). The geometry inherent in those coefficients is exploited by means of directional, predictive coding. We accomplish directional prediction with the aid of the coefficients of the undecimated wavelet transform (UDWT), which are estimated from the available DWT ones both at the encoder and decoder. Following directional prediction, a simple entropy coder takes care of the remaining redundancy. The resulting codec attains competitive performance.
Onur G. Guleryuz, Arthur L. da Cunha
ICIP1
2006 Nonlinear approximation based image recovery using adaptive sparse reconstructions and iterated denoising-part I: theory
abstract
We study the robust estimation of missing regions in images and video using adaptive, sparse reconstructions. Our primary application is on missing regions of pixels containing textures, edges, and other image features that are not readily handled by prevalent estimation and recovery algorithms. We assume that we are given a linear transform that is expected to provide sparse decompositions over missing regions such that a portion of the transform coefficients over missing regions are zero or close to zero. We adaptively determine these small magnitude coefficients through thresholding, establish sparsity constraints, and estimate missing regions in images using information surrounding these regions. Unlike prevalent algorithms, our approach does not necessitate any complex preconditioning, segmentation, or edge detection steps, and it can be written as a sequence of denoising operations. We show that the region types we can effectively estimate in a mean-squared error sense are those for which the given transform provides a close approximation using sparse nonlinear approximants. We show the nature of the constructed estimators and how these estimators relate to the utilized transform and its sparsity over regions of interest. The developed estimation framework is general, and can readily be applied to other nonstationary signals with a suitable choice of linear transforms. Part I discusses fundamental issues, and Part II is devoted to adaptive algorithms with extensive simulation examples that demonstrate the power of the proposed techniques.
Onur G. Guleryuz
IEEE Trans. Image Process.1
2006 Nonlinear approximation based image recovery using adaptive sparse reconstructions and iterated denoising-part II: adaptive algorithms
abstract
We combine the main ideas introduced in Part I with adaptive techniques to arrive at a powerful algorithm that estimates missing data in nonstationary signals. The proposed approach operates automatically based on a chosen linear transform that is expected to provide sparse decompositions over missing regions such that a portion of the transform coefficients over missing regions are zero or close to zero. Unlike prevalent algorithms, our method does not necessitate any complex preconditioning, segmentation, or edge detection steps, and it can be written as a progression of denoising operations. We show that constructing estimates based on nonlinear approximants is fundamentally a nonconvex problem and we propose a progressive algorithm that is designed to deal with this issue directly. The algorithm is applied to images through an extensive set of simulation examples, primarily on missing regions containing textures, edges, and other image features that are not readily handled by established estimation and recovery methods. We discuss the properties required of good transforms, and in conjunction, show the types of regions over which well-known transforms provide good predictors. We further discuss extensions of the algorithm where the utilized transforms are also chosen adaptively, where unpredictable signal components in the progressions are identified and not predicted, and where the prediction scenario is more general.
Onur G. Guleryuz
IEEE Trans. Image Process.1
2006 Linear, Worst-Case Estimators for Denoising Quantization Noise in Transform Coded Images
abstract
Transform-coded images exhibit distortions that fall outside of the assumptions of traditional denoising techniques. In this paper, we use tools from robust signal processing to construct linear, worst-case estimators for the denoising of transform compressed images. We show that while standard denoising is fundamentally determined by statistical models for images alone, the distortions induced by transform coding are heavily dependent on the structure of the transform used. Our method, thus, uses simple models for the image and for the quantization error, with the latter capturing the transform dependency. Based on these models, we derive optimal, linear estimators of the original image that are optimal in the mean-squared error sense for the worst-case cross correlation between the original and the quantization error. Our construction is transform agnostic and is applicable to transforms from block discrete cosine transforms to wavelets. Furthermore, our approach is applicable to different types of image statistics and can also serve as an optimization tool for the design of transforms/quantizers. Through the interaction of the source and quantizer models, our work provides useful insights and is instrumental in identifying and removing quantization artifacts from general signals coded with general transforms. As we decouple the modeling and processing steps, we allow for the construction of many different types of estimators depending on the desired sophistication and available computational complexity. In the low end of this spectrum, our lookup table based estimator, which can be deployed in low complexity environments, provides competitive PSNR values with some of the best results in the literature.
Onur G. Guleryuz
IEEE Trans. Image Process.1
2005 A nonlinear loop filter for quantization noise removal in hybrid video compression
abstract
We propose a high-performance, nonlinear, loop filter that reduces quantization noise over video frames composed of locally uniform regions (smooth, high frequency, texture, etc.) separated by singularities. Unlike earlier work, the designed filter is not limited to block based transform coders and provides a robust solution for intra as well as differentially encoded frames. Our formulation based on sparse decompositions allows us to take advantage of the spatial dependencies inherent in video frames while incorporating temporal dependencies through the information provided by previously decoded frames. The proposed filter results in 10% improvements in rate (at typical distortions) in combination with significant visual quality improvements, especially around singularities.
Onur G. Guleryuz
ICIP (2)1
2004 Predicting Wavelet Coefficients Over Edges Using Estimates Based on Nonlinear Approximants
abstract
It is well-known that wavelet transforms provide sparse decompositions over many types of image regions but not over image singularities/edges that manifest themselves along curves. It is now widely accepted that, on 2D piecewise smooth signals, wavelet compression performance is dominated by coefficients over edges. Research in this area has focused on two tracks, each suffering from issues related to translation invariance. Methods that directly model high order coefficient dependencies over edges have to combat aliasing issues, and new transforms that have been designed lose their full strength if they are not used in a translation invariant fashion. In this paper we combine these approaches and use translation invariant, overcomplete representations to predict wavelet edge coefficients. By starting with the lowest frequency band of an l level wavelet decomposition, we reliably estimate missing higher frequency coefficients over piecewise smooth signals. Unlike existing techniques, our approach does not model edges directly but implicitly obtains boundaries by aggressively determining regions where the utilized translation invariant decomposition is sparse.
Onur G. Guleryuz
Data Compression Conference1
2004 Pixel recovery via el minimization in the wavelet domain
Ivan W. Selesnick, Richard M. Van Slyke, Onur G. Guleryuz
ICIP3
2004 Globally optimal wavelet-based motion estimation using interscale edge and occlusion models
abstract
We propose a non-iterative, globally optimal dense motion field estimation technique based on a multiresolutional probability model. We consider the field to be estimated in terms of its wavelet coefficients and carry out the estimation in the field’s wavelet transform domain. Our approach models interscale dependencies of the wavelet coefficients and allows for smooth, edge, and occluded regions in the field. We obtain segmentations of the field and our results show that the field estimates yield accurate depictions of scene motion. The globally optimal nature of our estimation framework allows it to be applicable in scenes exhibiting large motion and in settings of ill-posed motion. Hence, our algorithms can also be used to determine accurate initializations for optical flow type estimation techniques, which use more sophisticated models but can only obtain locally optimal solutions that are heavily dependent on initial conditions. The performance is illustrated on several examples.
Levent Sendur, Onur G. Guleryuz
VCIP2
2003 Nonlinear approximation based image recovery using adaptive sparse reconstructions
abstract
We study the robust estimation of missing regions in images and video using adaptive, sparse reconstructions. We assume that we are given a linear transform that is expected to provide sparse decompositions over missing regions such that a portion of the transform coefficients over missing regions are zero or close to zero. We adaptively determine these small magnitude coefficients through thresholding, establish sparsity constraints, and estimate missing regions in images using information surrounding these regions. We show that the region types we can effectively estimate in a mean squared error sense are those for which the given transform provides a close approximation using nonlinear approximation. We show the nature of the constructed estimators and how these estimators relate to the utilized transform and its sparsity over regions of interest. For images the developed algorithms are applicable over locally uniform (smooth, high frequency, texture, etc.) regions separated by edges or edge-like singularities. However, the developed estimation framework is general, and can readily be applied to other nonstationary signals with a suitable choice of linear transforms. Equations are derived and extensive simulation examples included.
Onur G. Guleryuz
ICIP (1)1
2002 Iterated Denoising for Image Recovery
abstract
We propose an algorithm for image recovery where completely lost blocks in an image/video-frame are recovered using spatial information surrounding these blocks. Our primary application is on lost regions of pixels containing textures, edges and other image features that are not readily handled by prevalent recovery and error concealment algorithms. The proposed algorithm is based on the iterative application of a generic denoising algorithm and it does not necessitate any complex preconditioning, segmentation, or edge detection steps. Utilizing locally sparse linear transforms and overcomplete denoising, we obtain good PSNR performance in the recovery of such regions. In addition to results on image recovery, the paper provides further insights into the usefulness of popular transforms like wavelets, wavelet packets, discrete cosine transform (DCT) and complex wavelets in providing sparse image representations.
Onur G. Guleryuz
DCC1
2002 Fast text/graphics resolution improvement using wavelet based denoising and chain-code table lookup
abstract
We propose a fast text/graphics resolution improvement algorithm with boundary parameterization and wavelet based denoising. Given input images containing labeled text/graphics objects, local boundary segments on each object are traced, parameterized, smoothed, and subsequently rerendered at the desired resolution. All of the critical operations are delegated to lookup tables, resulting in an algorithm that is fast, computationally inexpensive, with low memory requirements. A very flexible framework is proposed which can be utilized in a variety of text/graphics applications requiring resolution improvement.
Onur G. Guleryuz, Anoop Bhattacharjya
ICIP (1)1
2002 On the importance of combining wavelet-based nonlinear approximation with coding strategies
abstract
This paper provides a mathematical analysis of transform compression in its relationship to linear and nonlinear approximation theory. Contrasting linear and nonlinear approximation spaces, we show that there are interesting classes of functions/random processes which are much more compactly represented by wavelet-based nonlinear approximation. These classes include locally smooth signals that have singularities, and provide a model for many signals encountered in practice, in particular for images. However, we also show that nonlinear approximation results do not always translate to efficient compress on strategies in a rate-distortion sense. Based on this observation, we construct compression techniques and formulate the family of functions/stochastic processes for which they provide efficient descriptions in a rate-distortion sense. We show that this family invariably leads to Besov spaces, yielding a natural relationship among Besov smoothness, linear/nonlinear approximation order, and compression performance in a rate-distortion sense. The designed compression techniques show similarities to modern high-performance transform codecs, allowing us to establish relevant rate-distortion estimates and identify performance limits.
Albert Cohen 0002, Ingrid Daubechies, Onur G. Guleryuz, Michael T. Orchard
IEEE Trans. Inf. Theory3
2002 Information-theoretic inequalities for contoured probability distributions
abstract
We show that for a special class of probability distributions that we call contoured distributions, information-theoretic invariants and inequalities are equivalent to geometric invariants and inequalities of bodies in Euclidean space associated with the distributions. Using this, we obtain characterizations of contoured distributions with extremal Shannon and Renyi entropy. We also obtain a new reverse information-theoretic inequality for contoured distributions.
Onur G. Guleryuz, Erwin Lutwak, Deane Yang, Gaoyong Zhang
IEEE Trans. Inf. Theory1
2001 On the DPCM compression of Gaussian autoregressive sequences
abstract
Differential pulse-coded modulation (DPCM) encoding of Gaussian autoregressive (AR) sequences is considered. It is pointed out that DPCM is rate-distortion inefficient at low bit rates. Simple filtering modifications are proposed and incorporated into DPCM. A rate-distortion optimization framework that results in optimal filters is presented. It is shown that the designed filters take advantage of "less significant" process spectral components in order to achieve superior rate-distortion performance. Design equations are derived, issues related to optimization and complexity addressed. It is shown that simple DPCM systems with the proposed modifications significantly outperform their standard counterparts.
Onur G. Guleryuz, Michael T. Orchard
IEEE Trans. Inf. Theory1
2000 Maximizing uniform translational motion: Motion estimation with the Haar transform and dynamic programming
abstract
This paper proposes a new motion estimation framework based on localized linear transforms, multiresolutional probability models and dynamic programming. We incorporate localized linear transforms (specifically wavelets) into motion estimation by parameterizing the motion fields to be estimated in terms of their localized linear transform coefficients. In terms of these coefficients, we propose a simple multiresolutional probability model that captures the possible local smoothness in the field to be estimated while allowing for discontinuities and uncovered regions. Within this framework we formulate the motion estimation problem as a MAP optimization problem that can be tackled with dynamic programming to yield the globally optimal dense motion field.
Onur G. Guleryuz
ICASSP1
2000 Subspaces of Quantization Artifacts for Image Transform Compression
abstract
This paper presents a new algorithm for reducing quantization artifacts in transform coded images. For transform coded signals, we analytically show that one can determine linear subspaces where most of the quantization artifacts reside. We formulate statistical constraints that identify these subspaces and propose a simple algorithm that removes them. Our formulation is general and is not tied to a specific linear transform or quantizer. The proposed algorithm, which is non-iterative, non-redundant and computationally inexpensive, results in visually pleasing images and competitive PSNR values when compared to the best post-processing results reported in the literature.
Onur G. Guleryuz
ICIP1
1999 Dense Motion Fields for Digital Video Processing and Compression
abstract
This paper presents novel algorithms that perform motion estimation for video processing and compression. We observe that "smoothness" is a very important and intuitive property in the estimation of motion fields. It is pointed out that most current motion estimation techniques implement smoothness as a constraint, differing only in terms of the specific type of smoothness demanded from video data. This paper views smoothness as a property that is determined by the underlying video data rather than a predetermined and specific property that is imposed on video data. Instead of forcing the available video data to conform to an abstract smoothness model, we try to select the "smoothest" motion field that conforms to the available data. We propose fast and efficient techniques that determine a set of possible motion fields and that select the smoothest field from this set. Issues like quantization and embedded video compression (via embedded motion fields) are discussed.
Shunan Lin, Onur G. Guleryuz
ICIP (3)2
1999 Rate-Distortion Modeling of Binary Shape Using State Partitioning
abstract
In this paper, the rate-distortion (R-D) characteristics of binary shapes are modeled. Specifically we are interested in predicting the rate and distortion that is produced by the shape coding techniques that have been adopted into the MPEG-4 standard. The shape coding algorithm is a context-based arithmetic encoder and operates on a per block basis. Currently, there is no efficient way of estimating the rate and distortion at various levels of resolution. Consequently, we propose a model that is based on a set of parameters that can be easily extracted from the binary blocks. The parameters represent states that arise from the possible binary patterns that can occur in a small neighborhood around the current pixel. Symmetry is exploited to keep the number of states to a minimum. It is shown that the proposed model is computationally efficient and provides accurate estimates of the R-D characteristics of a binary shape.
Anthony Vetro, Huifang Sun, Yao Wang 0001, Onur G. Guleryuz
ICIP (2)4
1997 Optimized nonorthogonal transforms for image compression
abstract
The transform coding of images is analyzed from a common standpoint in order to generate a framework for the design of optimal transforms. It is argued that all transform coders are alike in the way they manipulate the data structure formed by transform coefficients. A general energy compaction measure is proposed to generate optimized transforms with desirable characteristics particularly suited to the simple transform coding operation of scalar quantization and entropy coding. It is shown that the optimal linear decoder (inverse transform) must be an optimal linear estimator, independent of the structure of the transform generating the coefficients. A formulation that sequentially optimizes the transforms is presented, and design equations and algorithms for its computation provided. The properties of the resulting transform systems are investigated. In particular, it is shown that the resulting basis are nonorthogonal and complete, producing energy compaction optimized, decorrelated transform coefficients. Quantization issues related to nonorthogonal expansion coefficients are addressed with a simple, efficient algorithm. Two implementations are discussed, and image coding examples are given. It is shown that the proposed design framework results in systems with superior energy compaction properties and excellent coding results.
Onur G. Guleryuz, Michael T. Orchard
IEEE Trans. Image Process.1
1996 Rate-Distortion Based Temporal Filtering for Video Compression
abstract
We consider the temporal DPCM loop at the heart of most modern high performance video coders. Targeting low bitrate-low complexity video applications, it is shown that DPCM is inefficient in this region. The DPCM codec is analyzed in the low bitrate region and rate-distortion optimal modifications are proposed that do not violate the low complexity requirement. The proposed modifications involve negligible added complexity at the encoder and no added complexity at the decoder and are thus compatible with standard coders and bit streams.
Onur G. Guleryuz, Michael T. Orchard
Data Compression Conference1
1996 A DCT-based embedded image coder
abstract
Since Shapiro (see ibid., vol.41, no.12, p. 445, 1993) published his work on embedded zerotree wavelet (EZW) image coding, there have been increased research activities in image coding centered around wavelets. We first point out that the wavelet transform is just one member in a family of linear transformations, and the discrete cosine transform (DCT) can also be coupled with an embedded zerotree quantizer. We then present such an image coder that outperforms any other DCT-based coder published in the literature, including that of the Joint Photographers Expert Group (JPEG). Moreover, our DCT-based embedded image coder gives higher peak signal-to-noise ratios (PSNR) than the quoted results of Shapiro's EZW coder.
Zixiang Xiong, Onur G. Guleryuz, Michael T. Orchard
IEEE Signal Process. Lett.2