David S. Taubman

dblp:t/DavidSTaubman · also David Taubman · DBLP profile ↗
← Back
182ranked-venue papers
21as first author
21since 2021 · last 2025
0000-0002-8458-6402ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 172 · 19 first-author · 19 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multi-View Disparity Estimation Using the Gradient Consistency Model
abstract
Variational approaches to disparity estimation typically use a linearised brightness constancy constraint, which only applies in smooth regions and over small distances. Accordingly, current variational approaches rely on a schedule to progressively include image data. This paper proposes the use of Gradient Consistency information to assess the validity of the linearisation; this information is used to determine the weights applied to the data term as part of an analytically inspired Gradient Consistency Model. The Gradient Consistency Model penalises the data term for view pairs that have a mismatch between the spatial gradients in the source view and the spatial gradients in the target view. Instead of relying on a tuned or learned schedule, the Gradient Consistency Model is self-scheduling, since the weights evolve as the algorithm progresses. We show that the Gradient Consistency Model outperforms standard coarse-to-fine schemes and the recently proposed progressive inclusion of views approach in both rate of convergence and accuracy.
James L. Gray, Aous Thabit Naman, David S. Taubman
IEEE Trans. Image Process.3
2024 SEA: Sign-Separated Accumulation Scheme for Resource-Efficient DNN Accelerators
abstract
Deep neural network (DNN) accelerators targeting training need to support the resource-hungry floating-point (FP) arithmetic. Typically, the additions (accumulations) are performed with higher precision compared to multiplications, resulting in a substantial proportion of resources consumed by the adders. In this paper, we aim to improve the resource efficiency of FP addition in DNN training accelerators and present a sign-separated accumulation scheme (SEA). The proposed SEA scheme accumulates the same-signed terms separately, followed by a final addition of the two oppositely-signed sub-accumulations. The separate sub-accumulations, when performed using our resource-efficient same-signed FP adders, result in significant improvement in overall resource efficiency. We present a SEA-based systolic array design, involving a novel processing element design, capable of performing FP matrix multiplications for DNN training. Our experimental results show substantial improvements in delay, area-delay product (ADP), and energy consumption, by achieving reductions of up to 19.8% in delay, 29.1% in ADP and 30.7% in energy, across various adder-multiplier combinations compared to the original designs. SEA introduces no approximations in accumulation or modifications in the rounding mechanism. We integrate the SEA-based systolic array into the open-source Gemmini [1] ecosystem for use by the broader community.
Hassaan Saadat, Haris Javaid, Hasindu Gamaarachchi, David S. Taubman, Sri Parameswaran
DATE5
2024 Efficient motion modelling with variable-sized blocks from hierarchical cuboidal partitioning
Priyabrata Karmakar, M. Manzur Murshed, Manoranjan Paul, David S. Taubman
Multim. Tools Appl.4
2024 Exploration of Learned Lifting-Based Transform Structures for Fully Scalable and Accessible Wavelet-Like Image Compression
abstract
This paper provides a comprehensive study on features and performance of different ways to incorporate neural networks into lifting-based wavelet-like transforms, within the context of fully scalable and accessible image compression. Specifically, we explore different arrangements of lifting steps, as well as various network architectures for learned lifting operators. Moreover, we examine the impact of the number of learned lifting steps, the number of channels, the number of layers and the support of kernels in each learned lifting operator. To facilitate the study, we investigate two generic training methodologies that are simultaneously appropriate to a wide variety of lifting structures considered. Experimental results ultimately suggest that retaining fixed lifting steps from the base wavelet transform is highly beneficial. Moreover, we demonstrate that employing more learned lifting steps and more layers in each learned lifting operator do not contribute strongly to the compression performance. However, benefits can be obtained by utilizing more channels in each learned lifting operator. Ultimately, the learned wavelet-like transform proposed in this paper achieves over 25% bit-rate savings compared to JPEG 2000 with compact spatial support.
Xinyue Li 0002, Aous Thabit Naman, David S. Taubman
IEEE Trans. Image Process.3
2023 JPEG Pleno Light Field Encoder with Mesh based View Warping
abstract
We introduce mesh-based view warping to the JPEG Pleno light field coding framework and replace the standardized sample-based forward warping and splatting of reference texture with mesh-based backward warping, which allows for a more disciplined interpolation of the reference texture for predicting the target view. Instead of coding depth maps with JPEG 2000, which is the default option of the JPEG Pleno framework, we employ a recent extension referred to as JPEG 2000 Part 17. This extension utilises breakpoints to describe discontinuity boundary geometry for the purpose of modifying the predict and update lifting steps in the vicinity of detected discontinuities. We directly decode breakpoints and corresponding DWT coefficients onto a mesh and describe a scheme to construct a single, consolidated mesh for a large group of views, borrowing information from multiple coded depth maps. Results show that the cumulative impact of all these modifications enable improved rate-distortion performance in comparison with the default operation of the JPEG Pleno encoder.
Yue Li 0034, Reji Mathew, David S. Taubman
ICIP3
2023 Improved Transform Structures for Learned Wavelet-Like Fully Scalable Image Compression
abstract
This paper studies features and performance of different structures for learned wavelet-like transforms in fully scalable image compression. Specifically, we explore different neural network topologies and various arrangements of lifting steps to improve the existing wavelet transform. Experimental results strongly suggest that the proposed proposal-opacity network topology, comprising a collection of linear predictions modulated by non-linear opacities, performs better than the other considered designs. Results also show that augmenting a good base wavelet transform with two learned lifting steps performs significantly better than other learned lifting structures, achieving 26.8% bit-rate savings over JPEG 2000.
Xinyue Li 0002, Aous Thabit Naman, David S. Taubman
MMSP3
2023 JPEG 2000 Extensions for Scalable Coding of Discontinuous Media
abstract
In this paper we propose novel extensions to JPEG 2000 for the coding of discontinuous media which includes piecewise smooth imagery such as depth maps and optical flows. These extensions use breakpoints to model discontinuity boundary geometry and apply a breakpoint dependent Discrete Wavelet Transform (BP-DWT) to the input imagery. The highly scalable and accessible coding features provided by the JPEG 2000 compression framework are preserved by our proposed extensions, with the breakpoint and transform components encoded as independent bit streams that can be progressively decoded. Comparative rate-distortion results are provided along with corresponding visual examples which highlight the advantages of using breakpoint representations with accompanying BD-DWT and embedded bit-plane coding. Recently our proposed extensions have been adopted and are in the process of being published as a new Part 17 to the JPEG 2000 family of coding standards.
Reji Mathew, Aous Thabit Naman, Yue Li 0034, David S. Taubman
IEEE Trans. Image Process.4
2022 Efficient Scalable 360-degree Video Compression Scheme using 3D Cuboid Partitioning
abstract
Video coding techniques minimize spatial and temporal redundancies inherent in video sequences based on non-overlapping block-based image partitioning. Due to depending on the information from already encoded neighboring blocks, these algorithms lack efficient techniques to exploit the overall global redundancies. Compared to the traditional block-based coding, the cuboid coding (2D) framework has been proven to be a more effective method of image compression that exploits global redundancy by considering homogeneous pixel correlation within a frame. In this paper, we improved the idea of 2D cuboid coding to exploit both local and global redundancy from a video sequence by adopting a three-dimensional (3D) cuboid partitioning scheme for SHVC compression improvement of 360-degree videos. The proposed method considers a group of successive frames as a 3D cuboid and recursively partitions it into sub-3D cuboids where static information over a selected GOP share the same cuboid and moving regions share new cuboids with better-defined objects. All the 3D cuboids are then encoded to create a coarse representation of the video stream. Experiments indicate that the proposed framework significantly outperforms its relevant benchmarks, notably by 17.18% (average) in BD-Rate reduction and 0.82 dB in BD-PSNR gain with respect to the standard SHVC codec.
Fariha Afsana, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
ICIP4
2022 A Neural Network Lifting Based Secondary Transform for Improved Fully Scalable Image Compression in Jpeg 2000
abstract
This paper proposes a secondary transform to improve wavelet-based image compression schemes such as JPEG 2000. The proposed approach includes two neural network steps, a high-to-low step followed by a low-to-high step. The high-to-low step suppresses aliasing in the low-pass band by using the detail bands at the same resolution, while the low-to-high step aims to further remove redundant information from the detail bands so as to achieve higher energy compaction. We employ the same neural networks for these two steps at every level of the hierarchical discrete wavelet transform (DWT) decomposition. We demonstrate that the visual quality of the LL bands at different resolutions can be dramatically enhanced, along with greatly improved coding efficiency up to 2dB compared with JPEG 2000, while preserving the full quality and resolution scalability and spatial random access features of JPEG 2000.
Xinyue Li 0002, Aous Thabit Naman, David S. Taubman
ICIP3
2022 Breakpoint Dependent Scalable Coding of Optical Flow Volume
abstract
Motion representations that describe motion flow from one single anchor frame to multiple reference frames can be useful for many tasks such as motion compensation (MC) and frame-rate up-sampling. We construct an anchored, multi-flow representation which we refer to as an Optical Flow Volume (OFV). We explore the use of recent breakpoint dependent DWT (BD-DWT) being considered as part of extensions to JPEG 2000 for coding discontinuous media. Breakpoints describe discontinuity boundary geometry and we estimate a single set of breakpoints that can be shared by all individual flows of the OFV. Additionally, as flows are anchored at a common frame, we are able to readily explore inter-flow transforms. Rate scalable results show significant rate-distortion gains for BD-DWT over the 5/3 DWT commonly used with JPEG 2000. The validity and utility of OFV coding are confirmed with accompanying MC results.
Reji Mathew, David S. Taubman
ICIP2
2022 JPEG Pleno Light Field Encoder with Breakpoint Dependent Affine Wavelet Transform for Disparity Maps
abstract
The JPEG Pleno light field encoder can perform disparity compensated view prediction for coding HDCA views. For operating in this prediction mode, disparity maps need to be communicated along with the 2D array of views. In this work we explore a new adaptive wavelet transform for coding disparity maps, known as tri-breakpoint (TriBRK) dependent DWT, that is currently being considered as part of JPEG 2000 Part-17 extensions. We show that the scalable TriBRK dependent DWT, defined on a hierarchical triangular grid, provides rate-distortion gains for coding disparity maps and improves the compression of HDCA views that rely upon the disparity information. Visual quality improvements are also observed, specifically at object boundaries of decoded views.
Reji Mathew, David S. Taubman
ICIP2
2022 Gradient Consistency Based Multi-Scale Optical Flow
abstract
Estimating optical flow and/or depth from multiple views or frames, all with multiple scales can be considered a data selection problem. We propose a gradient consistency based multi-scale optical flow technique which aims to address this data selection problem. We propose a method to assess gradient consistency and show how to use it to appropriately downweight data with poor consistency. We also introduce a spatial regularisation term that exploits optical flow acceleration between consecutive frame pairs as part of a global variational framework, coupling data between consecutive frame pairs. The proposed gradient consistency based multi-scale approach produces improved results on both real and synthetic scenes compared to a coarse to fine framework especially with limited numbers of warps.
James L. Gray, Aous Thabit Naman, David S. Taubman
MMSP3
2022 Transform Quantization for CNN Compression
abstract
In this paper, we compress convolutional neural network (CNN) weights post-training via transform quantization. Previous CNN quantization techniques tend to ignore the joint statistics of weights and activations, producing sub-optimal CNN performance at a given quantization bit-rate, or consider their joint statistics during training only and do not facilitate efficient compression of already trained CNN models. We optimally transform (decorrelate) and quantize the weights post-training using a rate-distortion framework to improve compression at any given quantization bit-rate. Transform quantization unifies quantization and dimensionality reduction (decorrelation) techniques in a single framework to facilitate low bit-rate compression of CNNs and efficient inference in the transform domain. We first introduce a theory of rate and distortion for CNN quantization, and pose optimum quantization as a rate-distortion optimization problem. We then show that this problem can be solved using optimal bit-depth allocation following decorrelation by the optimal End-to-end Learned Transform (ELT) we derive in this paper. Experiments demonstrate that transform quantization advances the state of the art in CNN compression in both retrained and non-retrained quantization scenarios. In particular, we find that transform quantization with retraining is able to compress CNN models such as AlexNet, ResNet and DenseNet to very low bit-rates (1-2 bits).
Sean I. Young, Zhe Wang 0019, David S. Taubman, Bernd Girod
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Efficient Scalable UHD/360-Video Coding by Exploiting Common Information With Cuboid-Based Partitioning
abstract
The scalable extension of High Efficiency Video Coding, SHVC can code Ultra High-Definition (UHD) video, including 360-degree video for various devices to serve a single bitstream with different display resolutions and qualities. To improve the SHVC compression efficiency, this paper proposes a novel intra and inter-frame coding scheme by first separating the common/visually important information and then applying cuboid-based variable size block partitioning and coding process for the common/visually important information in the base layer. In cuboid-based partitioning a video frame is partitioned into arbitrary shaped rectangular regions, known as cuboids, based on the distribution of relatively homogeneous pixel values. As the cuboid adopts a variable block partitioning based on the homogeneity of the data value, the partitioned blocks have better alignment with the object boundary. Moreover, in the cuboid coding process, only the partitioning tree information and a single value for each block need to be coded which takes lower number of bits and computational time compared to the traditional SHVC base layer. To verify the performance of the proposed method we embedded the proposed scheme as a base layer into the standard SHVC reference software and used several popular UHD/360-degree videos. The experimental results indicate that the proposed scalable coding strategy achieves an average of 14.04% BD-Rate reduction and 0.61 dB BD-PSNR gain for UHD/360-video compared to the operation points provided by an SHVC conforming encoder.
Fariha Afsana, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
IEEE Trans. Circuits Syst. Video Technol.4
2022 A Commonality Modeling Framework for Enhanced Video Coding Leveraging on the Cuboidal Partitioning Based Representation of Frames
abstract
Video coding algorithms attempt to minimize the significant commonality that exists within a video sequence. Each new video coding standard contains tools that can perform this task more efficiently compared to its predecessors. Modern video coding systems are block-based wherein commonality modeling is carried out only from the perspective of the block that need be coded next. In this work, we argue for a commonality modeling approach that can provide a seamless blending between global and local homogeneity information. For this purpose, at first the frame that need be coded, is recursively partitioned into rectangular regions based on the homogeneity information of the entire frame. After that each obtained rectangular region’s feature descriptor is taken to be the average value of all the pixels’ intensities encompassing the region. In this way, the proposed approach generates a coarse representation of the current frame by minimizing both global and local commonality. This coarse frame is computationally simple and has a compact representation. It attempts to preserve important structural properties of the current frame which can be viewed subjectively as well as from improved rate-distortion performance of a reference scalable HEVC coder that employs the coarse frame as a reference frame for encoding the current frame.
Ashek Ahmmed, M. Manzur Murshed, Manoranjan Paul, David S. Taubman
IEEE Trans. Multim.4
2021 Dynamic Point Cloud Compression Using A Cuboid Oriented Discrete Cosine Based Motion Model
abstract
Immersive media representation format based on point clouds has underpinned significant opportunities for extended reality applications. Point cloud in its uncompressed format require very high data rate for storage and transmission. The video based point cloud compression technique projects a dynamic point cloud into geometry and texture video sequences. The projected texture video is then coded using modern video coding standard like HEVC. Since the properties of projected texture video frames are different from traditional video frames, HEVC-based commonality modeling can be inefficient. An improved commonality modeling technique is proposed that employs discrete cosine basis oriented motion models and the domains of such models are approximated by homogeneous regions called cuboids. Experimental results show that the proposed commonality modeling technique can yield savings in bit rate of up to 4.17%.
Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
ICASSP4
2021 Human-Machine Collaborative Video Coding Through Cuboidal Partitioning
abstract
Video coding algorithms encode and decode an entire video frame while feature coding techniques only preserve and communicate the most critical information needed for a given application. This is because video coding targets human perception, while feature coding aims for machine vision tasks. Recently, attempts are being made to bridge the gap between these two domains. In this work, we propose a video coding framework by leveraging on to the commonality that exists between human vision and machine vision applications using cuboids. This is because cuboids, estimated rectangular regions over a video frame, are computationally efficient, has a compact representation and object centric. Such properties are already shown to add value to traditional video coding systems. Herein cuboidal feature descriptors are extracted from the current frame and then employed for accomplishing a machine vision task in the form of object detection. Experimental results show that a trained classifier yields superior average precision when equipped with cuboidal features oriented representation of the current test frame. Additionally, this representation costs 7% less in bit rate if the captured frames are need be communicated to a receiver.
Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
ICIP4
2021 Dynamic Point Cloud Geometry Compression using Cuboid based Commonality Modeling Framework
abstract
Point cloud in its uncompressed format require very high data rate for storage and transmission. The video based point cloud compression (V-PCC) technique projects a dynamic point cloud into geometry and texture video sequences. The projected geometry and texture video frames are then encoded using modern video coding standard like HEVC. However, HEVC encoder is unable to exploit the global commonality that exists within a geometry frame and between successive geometry frames to a greater extent. This is because in HEVC, the current frame partitioning starts from a rigid $64 \times 64$ pixels level without considering the structure of the scene need be coded. In this paper, an improved commonality modeling framework is proposed, by leveraging on cuboid-based frame partitioning, to encode point cloud geometry frames. The associated frame-partitioning scheme is based on statistical properties of the current geometry frame and therefore yields a flexible block partitioning structure composed of cuboids. Additionally, the proposed commonality modeling approach is computationally efficient and has a compact representation. Experimental results show that if the V-PCC reference encoder is augmented by the proposed commonality modeling technique, a bit rate savings of 2.71% and 4.25% are achieved for full body and upper body of human point clouds’ geometry sequences respectively.
Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
ICIP4
2021 Welsch Based Multiview Disparity Estimation
abstract
In this work, we explore disparity estimation from a high number of views. We experimentally identify occlusions as a key challenge for disparity estimation for applications with high numbers of views. In particular, occlusions can actually result in a degradation in accuracy as more views are added to a dataset. We propose the use of a Welsch loss function for the data term in a global variational framework for disparity estimation. We also propose a disciplined warping strategy and a progressive inclusion of views strategy that can reduce the need for coarse to fine strategies that discard high spatial frequency components from the early iterations. Experimental results demonstrate that the proposed approach produces superior and/or more robust estimates than other conventional variational approaches.
James L. Gray, Aous Thabit Naman, David S. Taubman
ICIP3
2021 Machine-Learning Based Secondary Transform for Improved Image Compression in JPEG2000
abstract
This paper proposes a convolutional neural network (CNN) based secondary transform for not only improved coding efficiency in the JPEG2000 image compression format, but also to produce more appealing approximation sub-bands at different resolutions. The CNN in this work exploits information in detail sub-bands to predict some of the aliasing information in the corresponding approximation or low-pass sub-band; this reduction in aliasing, although not perfect, improves the compressibility of the “cleaned” approximation subband. This process is repeated in subsequent wavelet decomposition levels to further improve coding efficiency. Experimental results show that, at high bit rates, the proposed network outperforms conventional JPEG2000 compression framework by up to 1.2 dB, especially for images with strong geometric flow.
Xinyue Li 0002, Aous Thabit Naman, David S. Taubman
ICIP3
2021 Scalable Coding Of Motion And Depth Fields With Shared Breakpoints
abstract
A new breakpoint adaptive DWT, referred to as tri-break, is currently being considered by standardization efforts in relation to JPEG 2000 Part 17 extensions. We first provide a summary of the tri-break transform and then explore its performance for coding motion fields. Experimental results show that significant gains can be achieved for coding piecewise smooth motion flows by employing the tri-break transform. We demonstrate the feasibility of utilising a common set of breakpoints for compressing depth maps and motion fields anchored at the same frame. We also extend prior work to decode motion vector fields directly onto a triangular mesh, enabling operations such as view warping to be defined on triangular cells which can be efficiently processed by GPU based architectures.
Reji Mathew, Yue Li 0034, David S. Taubman
ICIP3
2020 Non-Line-of-Sight Surface Reconstruction Using the Directional Light-Cone Transform
abstract
We propose a joint albedo-normal approach to non-line-of-sight (NLOS) surface reconstruction using the directional light-cone transform (D-LCT). While current NLOS imaging methods reconstruct either the albedo or surface normals of the hidden scene, the two quantities provide complementary information of the scene, so an efficient method to estimate both simultaneously is desirable. We formulate the recovery of the two quantities as a vector deconvolution problem, and solve it via Cholesky-Wiener decomposition. We demonstrate that surfaces fitted non-parametrically using our recovered normals are more accurate than those produced with NLOS surface reconstruction methods recently proposed, and are 1,000 times faster to compute than using inverse rendering.
Sean I. Young, David B. Lindell, Bernd Girod, David S. Taubman, Gordon Wetzstein
CVPR4
2020 Augmenting JPEG2000 With Wavelet Coefficient Prediction
abstract
This paper presents a novel modification to the Discrete Wavelet Transform used in the JPEG2000 standard. The modification utilises a learning-based method to predict the detail coefficients from the approximation coefficients at each level of the transform, allowing for higher compression rates while maintaining most of JPEG2000's appealing features. A coefficient prediction method based on single-image super resolution literature has been implemented and tested to assess the potential of the compression system. We have observed that the accuracy of the prediction is highly dependent on the content of the test image, with certain features being highly predictable and others not predictable at all. In the best performing test images, we see a quality gain of up to 1.5dB or equivalently a reduction in bit-rate of up to 14% over the JPEG2000 standard.
Antoni Dimitriadis, David S. Taubman
ICIP2
2020 Adaptive Secondary Transform For Improved Image Coding Efficiency In JPEG2000
abstract
This paper proposes a secondary transform for wavelet based image compression, whose aim is to predict detail sub-bands from the corresponding low-pass sub-band so as to improve energy compaction. The proposed prediction strategy exploits the orientation of local image features to untangle aliasing components in the low-pass sub-band, and achieves prediction by projecting the cleaned synthesized low-pass sub-band back into the wavelet domain. At low bit rates, the proposed scheme can considerably reduce distortion in a JPEG 2000 based compression framework, especially for images with strong geometric flow.
Xinyue Li 0002, Aous Thabit Naman, David S. Taubman
ICIP3
2020 Encoding High-Throughput Jpeg2000 (Htj2k) Images On A Gpu
abstract
High-Throughput JPEG2000 (HTJ2K) is a new addition to the JPEG2000 suite of coding tools; it has been recently approved as Part-15 of the JPEG2000 standard, and the JPH file extension has been designated for it. The HTJ2K employs a new “fast” block coder that can achieve higher encoding and decoding throughput than a conventional JPEG2000 (C-J2K) encoder. The higher throughput is achieved because the HTJ2K codec processes wavelet coefficients in a smaller number of steps than C-J2K. Moreover, the HTJ2K block coder is also more amenable to parallelizable high-speed software and hardware implementations. The HTJ2K retains most of the features and capabilities of JPEG2000, and it also supports lossless transcoding between HTJ2K and already compressed C-J2K images. Quality scalability however is more limited than C-J2K. In a recent work, we presented preliminary performance results for decoding HTJ2K images on a GPU. In this work, we present a GPU-based HTJ2K encoder; we also present early encoding results for such an encoder, showing that it is possible to encode 4K4:4:4 HDR videos at more than 70 frames per second (fps) on a low-end card, while a high-end GPU can encode such videos at more than 400 fps.
Aous Thabit Naman, David S. Taubman
ICIP2
2020 Efficient Low Bit-Rate Intra-Frame Coding using Common Information for 360-degree Video
abstract
With the growth of video technologies, super-resolution videos, including 360-degree immersive video has become a reality due to exciting applications such as augmented/virtual/mixed reality for better interaction and a wide-angle user-view experience of a scene compared to traditional video with narrow-focused viewing angle. The new generation video contents are bandwidth-intensive in nature due to high resolution and demand high bit rate as well as low latency delivery requirements that pose challenges in solving the bottleneck of transmission and storage burdens. There is limited optimisation space in traditional video coding schemes for improving video coding efficiency in intra-frame due to the fixed size of processing block. This paper presents a new approach for improving intra-frame coding especially at low bit rate video transmission for 360-degree video for lossy mode of HEVC. Prior to using traditional HEVC intra-prediction, this approach exploits the global redundancy of entire frame by extracting common important information using multi-level discrete wavelet transformation. This paper demonstrates that the proposed method considering only low frequency information of a frame and encoding this can outperform the HEVC standard at low bit rates. The experimental results indicate that the proposed intra-frame coding strategy achieves an average of 54.07% BD-rate reduction and 2.84 dB BD-PSNR gain for low bit rate scenario compared to the HEVC. It also achieves a significant improvement in encoding time reduction of about 66.84% on an average. Moreover, this finding also demonstrates that the existing HEVC block partitioning can be applied in the transform domain for better exploitation of information concentration as we applied HEVC on wavelet frequency domain.
Fariha Afsana, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
MMSP4
2020 A Coarse Representation of Frames Oriented Video Coding By Leveraging Cuboidal Partitioning of Image Data
abstract
Video coding algorithms attempt to minimize the significant commonality that exists within a video sequence. Each new video coding standard contains tools that can perform this task more efficiently compared to its predecessors. In this work, we form a coarse representation of the current frame by minimizing commonality within that frame while preserving important structural properties of the frame. The building blocks of this coarse representation are rectangular regions called cuboids, which are computationally simple and has a compact description. Then we propose to employ the coarse frame as an additional source for predictive coding of the current frame. Experimental results show an improvement in bit rate savings over a reference codec for HEVC, with minor increase in the codec computational complexity.
Ashek Ahmmed, Manoranjan Paul, M. Manzur Murshed, David S. Taubman
MMSP4
2020 Scalable Mesh Representation for Depth from Breakpoint-Adaptive Wavelet Coding
abstract
A highly scalable and compact representation of depth data is required in many applications, and it is especially critical for plenoptic multiview image compression frameworks that use depth information for novel view synthesis and interview prediction. Efficiently coding depth data can be difficult as it contains sharp discontinuities. Breakpoint-adaptive discrete wavelet transforms (BPA-DWT) currently being standardized as part of JPEG 2000 Part-17 extensions have been found suitable for coding spatial media with hard discontinuities. In this paper, we explore a modification to the original BPA-DWT by replacing the traditional constant extrapolation strategy with the newly proposed affine extrapolation for reconstructing depth data in the vicinity of discontinuities. We also present a depth reconstruction scheme that can directly decode the BPA-DWT coefficients and breakpoints onto a compact and scalable mesh-based representation which has many potential benefits over the sample-based description. For performing depth compensated view prediction, our proposed triangular mesh representation of the depth data is a natural fit for modern graphics architectures.
Yue Li 0034, Reji Mathew, David S. Taubman
MMSP3
2020 Rate-Distortion Driven Decomposition of Multiview Imagery to Diffuse and Specular Components
abstract
In this work, we propose an overcomplete representation of multiview imagery for the purpose of compression. We present a rate-distortion (R-D) driven approach to decompose multiview datasets into two additive parts which can be interpreted as diffuse and specular content. We choose distinct and different sparsifying transforms for the diffuse and specular components and employ an R-D inspired measure as our optimization cost function to drive the decomposition based solely on compressibility. We first describe a framework which performs data separation in a registered domain to avoid the complexity of warping between views. Then a more comprehensive approach is proposed to separate specular data progressively from coordinates of multiple reference views. Experimental results show a coding gain of up to 0.6 dB for synthetic datasets and up to 0.9 dB for real datasets.
Maryam Haghighat, Reji Mathew, David S. Taubman
IEEE Trans. Image Process.3
2020 Gaussian Lifting for Fast Bilateral and Nonlocal Means Filtering
abstract
Recently, many fast implementations of the bilateral and the nonlocal filters were proposed based on lattice and vector quantization, e.g. clustering, in higher dimensions. However, these approaches can still be inefficient owing to the complexities in the resampling process or in filtering the high-dimensional resampled signal. In contrast, simply scalar resampling the high-dimensional signal after decorrelation presents the opportunity to filter signals using multi-rate signal processing techniques. Cis work proposes the Gaussian lifting framework for efficient and accurate bilateral and nonlocal means filtering, appealing to the similarities between separable wavelet transforms and Gaussian pyramids. Accurately implementing the filter is important not only for image processing applications, but also for a number of recently proposed bilateralregularized inverse problems, where the accuracy of the solutions depends ultimately on an accurate filter implementation. We show that our Gaussian lifting approach filters images more accurately and efficiently across many filter scales. Adaptive lifting schemes for bilateral and nonlocal means filtering are also explored.
Sean I. Young, Bernd Girod, David S. Taubman
IEEE Trans. Image Process.3
2020 Fast Optical Flow Extraction From Compressed Video
abstract
We propose the fast optical flow extractor, a filtering method that recovers artifact-free optical flow fields from HEVCcompressed video. To extract accurate optical flow fields, we form a regularized optimization problem that considers the smoothness of the solution and the pixelwise confidence weights of an artifactridden HEVC motion field. Solving such an optimization problem is slow, so we first convert the problem into a confidence-weighted filtering task. By leveraging the already-available HEVC motion parameters, we achieve a 100-fold speed-up in the running times compared to similar methods, while producing subpixel-accurate flow estimates. Je fast optical flow extractor is useful when video frames are already available in coded formats. Our method is not specific to a coder, and works with motion fields from video coders such as H.264/AVC and HEVC.
Sean I. Young, Bernd Girod, David S. Taubman
IEEE Trans. Image Process.3
2020 Graph Laplacian Regularization for Robust Optical Flow Estimation
abstract
This paper proposes graph Laplacian regularization for robust estimation of optical flow. First, we analyze the spectral properties of dense graph Laplacians and show that dense graphs achieve a better trade-off between preserving flow discontinuities and filtering noise, compared with the usual Laplacian. Using this analysis, we then propose a robust optical flow estimation method based on Gaussian graph Laplacians. We revisit the framework of iteratively reweighted least-squares from the perspective of graph edge reweighting, and employ the Welsch loss function to preserve flow discontinuities and handle occlusions. Our experiments using the Middlebury and MPI-Sintel optical flow datasets demonstrate the robustness and the efficiency of our proposed approach.
Sean I. Young, Aous Thabit Naman, David S. Taubman
IEEE Trans. Image Process.3
2019 Solving Vision Problems via Filtering
abstract
We propose a new, filtering approach for solving a large number of regularized inverse problems commonly found in computer vision. Traditionally, such problems are solved by finding the solution to the system of equations that expresses the first-order optimality conditions of the problem. This can be slow if the system of equations is dense due to the use of nonlocal regularization, necessitating iterative solvers such as successive over-relaxation or conjugate gradients. In this paper, we show that similar solutions can be obtained more easily via filtering, obviating the need to solve a potentially dense system of equations using slow iterative methods. Our filtered solutions are very similar to the true ones, but often up to 10 times faster to compute.
Sean I. Young, Aous Thabit Naman, Bernd Girod, David S. Taubman
ICCV4
2019 Rate-Distortion Driven Separation of Diffuse and Specular Components in Multiview Imagery
abstract
In this work we explore an overcomplete representation of multiview imagery for the purpose of compression. We present a rate-distortion (R-D) driven approach to decompose multiview datasets into two additive parts which can be interpreted as being the diffuse and specular components. We apply different transforms to each component such that the compressibility of input data is improved. We describe a framework which performs the R-D optimized separation in a registered domain to avoid the complexity of warping between views. Experimental results highlight the benefits of the proposed source separation approach in the context of compression.
Maryam Haghighat, Reji Mathew, David S. Taubman
ICIP3
2019 WaSP Encoder with Breakpoint Adaptive DWT Coding of Disparity Maps
abstract
An encoder architecture that employs view warping and sparse prediction (WaSP) has recently been proposed by the JPEG Pleno standards activity for the compression of light field imagery. The proposed WaSP encoder utilises disparity information to exploit the correlation that exists between multiple views and therefore requires disparity data to be communicated. The WaSP framework currently employs JPEG2000 for the coding of disparity maps. In this work we explore the impact of introducing breakpoint adaptive DWT (BPA-DWT) coding of disparity maps. While prior work has shown promising results for the coding of depth maps with breakpoints, its impact on multi-view compression where the depth or disparity data forms part of the communicated side information has not been previously explored. Our investigations show important gains in RD performance coupled with improvements in the visual quality of decoded views by the introduction of BPA-DWT coding of disparity maps.
Reji Mathew, David S. Taubman
ICIP2
2019 Decoding High-Throughput Jpeg2000 (HTJ2K) On A G
abstract
High-throughput JPEG2000 (HTJ2K), also known as JPEG 2000 Part 15, is the most recent addition to the JPEG2000 suite of coding tools. Ee file extension JPH has been designated for compressed images employing this new part of the standard. Eis new part describes a "fast" block coder for the JPEG 2000 format, while retaining most other JPEG2000 features and capabilities intact. Ee HTJ2K block coder is amenable to parallelizable high-speed encoding and decoding implementations; moreover, it is designed to allow lossless transcoding of already compressed JPEG2000 images that employ the regular block coder. HTJ2K supports the scalability options available in the JPEG2000 format except for quality scalability, which is available only to a limited extent. Eis work gives a high-level overview of this new block coder; we also present preliminary performance results for a GPU implementation. We show that a low-end GPU can decode 4K 4:4:4 12-bit videos at more than 60 frames per second (fps) while a high-end GPU can decode 8K HDR videos at more than 120 fps.
Aous Thabit Naman, David S. Taubman
ICIP2
2019 High Throughput Block Coding in the HTJ2K Compression Standard
abstract
This paper describes the block coding algorithm that underpins the new High Throughput JPEG 2000 (HTJ2K) standard. The objective of HTJ2K is to overcome the computational complexity of the original block coding algorithm, by providing a drop-in replacement that preserves as much of the JPEG 2000 feature set as possible, while allowing reversible transcoding to/from the original format. We show how the new standard achieves these goals, with high coding efficiency, and extremely high throughput in software.
David S. Taubman, Aous Thabit Naman, Reji Mathew
ICIP1
2019 Consistent Disparity Synthesis for Inter-View Prediction in Lightfield Compression
abstract
For efficient compression of lightfields that involve many views, it has been found preferable to explicitly communicate disparity/depth information at only a small subset of the view locations. In this study, we focus solely on inter-view prediction, which is fundamental to multi-view imagery compression, and itself depends upon the synthesis of disparity at new view locations. Current HDCA standardization activities consider a framework known as WaSP, that hierarchically predicts views, independently synthesizing the required disparity maps at the reference views for each prediction step. A potentially better approach is to progressively construct a unified multi-layered base-model for consistent disparity synthesis across many views. This paper improves significantly upon an existing base-model approach, demonstrating superior performance to WaSP. More generally, the paper investigates the implications of texture warping and disparity synthesis methods.
Yue Li 0034, Reji Mathew, Dominic Rüfenacht, Aous Thabit Naman, David S. Taubman
PCS5
2019 Efficient Delivery of Very High Dynamic Range Compressed Imagery by Dynamic-Range-of-Interest
abstract
JPEG 2000 allows scenes to be encoded in a highly scalable and accessible manner, so that only the content that is relevant to a region or resolution of interest need be transmitted and decoded. This paper extends this property to allow efficient access into compressed images that have a very high dynamic range, based on a dynamic-range-of-interest. Optimized re-prioritization of the encoded content is used to stream imagery based on a potentially dynamic set of viewing conditions that implicitly identify the visual significance of content. We propose and validate a framework for doing this, based on a single compressed representation, a sparse set of display-independent luminance statistics and a novel algorithm for inferring the display-dependent significance of code-block distortions.
David S. Taubman
PCS2
2019 Temporal Frame Interpolation With Motion-Divergence-Guided Occlusion Handling
abstract
We present a high-quality temporal frame interpolation (TFI) method that employs piecewise-smooth motion and handles disoccluded regions using the observation that motion discontinuities travel with the foreground object. We derive a “motion discontinuity” likelihood map from the divergence of a motion field between the input frames. Motion which is modeled at the reference frame is mapped to the target frame using a cellular-affine mapping strategy - a process during which regions of disocclusion are readily observed. This information is then used to guide the occlusion-aware, bidirectional FI process. Furthermore, we propose two computationally inexpensive texture optimizations that selectively improve the quality of the interpolated frames in regions around moving objects. The scheme produces very high-quality interpolated frames and outperforms current high-quality state-of-the-art TFI schemes by 2-2.5 dB; the method works with a very low-complexity motion estimation scheme and runs orders of magnitudes faster than its competitors.
Dominic Rüfenacht, Reji Mathew, David S. Taubman
IEEE Trans. Circuits Syst. Video Technol.3
2019 Illumination Estimation and Compensation of Low Frame Rate Video Sequences for Wavelet-Based Video Compression
abstract
In this paper, we are interested in the compression of image sets or video with considerable changes in illumination. We develop a framework to decompose frames into illumination fields and texture in order to achieve sparser representations of frames which is beneficial for compression. Illumination variations or contrast ratio factors among frames are described by a full resolution multiplicative field. First, we propose a Lifting-based Illumination Adaptive Transform (LIAT) framework which incorporates illumination compensation to temporal wavelet transforms. We estimate a full resolution illumination field, taking heed of its spatial sparsity by a rate-distortion (R-D) driven framework. An affine mesh model is also developed as a point of comparison. We find the operational coding cost of the subband frames by modeling a typical t + 2D wavelet video coding system. While our general findings on R-D optimization are applicable to a range of coding frameworks, in this paper, we report results based on employing JPEG 2000 coding tools. The experimental results highlight the benefits of the proposed R-D driven illumination estimation and compensation in comparison with alternative scalable coding methods and non-scalable coding schemes of AVC and HEVC employing weighted prediction.
Maryam Haghighat, Reji Mathew, Aous Thabit Naman, David S. Taubman
IEEE Trans. Image Process.4
2019 Base-Anchored Model for Highly Scalable and Accessible Compression of Multiview Imagery
abstract
We present a compression scheme for multiview imagery that facilitates high scalability and accessibility of the compressed content. Our scheme relies upon constructing at a single base view, a disparity model for a group of views, and then utilizing this base-anchored model to infer disparity at all views belonging to the group. We employ a hierarchical disparity-compensated inter-view transform where the corresponding analysis and synthesis filters are applied along the geometric flows defined by the base-anchored disparity model. The output of this inter-view transform along with the disparity information is subjected to spatial wavelet transforms and embedded block-based coding. Rate-distortion results reveal superior performance to the x.265 anchor chosen by the JPEG Pleno standards activity for the coding of multiview imagery captured by high-density camera arrays.
Dominic Rüfenacht, Aous Thabit Naman, Reji Mathew, David S. Taubman
IEEE Trans. Image Process.4
2019 COGL: Coefficient Graph Laplacians for Optimized JPEG Image Decoding
abstract
We address the problem of decoding joint photographic experts group (JPEG)-encoded images with less visual artifacts. We view the decoding task as an ill-posed inverse problem and find a regularized solution using a convex, graph Laplacian-regularized model. Since the resulting problem is non-smooth and entails non-local regularization, we use fast high-dimensional Gaussian filtering techniques with the proximal gradient descent method to solve our convex problem efficiently. Our patch-based "coefficient graph" is better suited than the traditional pixel-based ones for regularizing smooth non-stationary signals such as natural images and relates directly to classic non-local means de-noising of images. We also extend our graph along the temporal dimension to handle the decoding of M-JPEG-encoded video. Despite the minimalistic nature of our convex problem, it produces decoded images with similar quality to other more complex, state-of-the-art methods while being up to five times faster. We also expound on the relationship between our method and the classic ANCE method, reinterpreting ANCE from a graph-based regularization perspective.
Sean I. Young, Aous Thabit Naman, David S. Taubman
IEEE Trans. Image Process.3
2018 Rate-Distortion Optimized Illumination Estimation for Wavelet-Based Video Coding
abstract
We propose a rate-distortion optimized framework for estimating illumination changes (lighting variations, fade in/out effects) in a highly scalable coding system. Illumination variations are realized using multiplicative factors in the image domain and are estimated considering the coding cost of the illumination field and input frames which are first subject to a temporal Lifting-based Illumination Adaptive Transform (LIAT). The coding cost is modelled by an ℓ1-norm optimization problem which is derived to approximate a quadratic-log function which emerges from rate-distortion considerations. The optimization problem is solved using ADMM. The proposed solution works the same or better than a mesh-based approach proposed in prior work, where sparsity was controlled by explicitly choosing mesh parameters. In the compression-inspired formulation presented here, sparsity is discovered automatically through the solution of a convex program that depends only on a target rate-distortion operating point.
Maryam Haghighat, Reji Mathew, Aous Thabit Naman, Sean I. Young, David S. Taubman
ICASSP5
2018 Enhanced Homogeneous Motion Discovery Oriented Prediction for Key Intermediate Frames
abstract
Conventional video compression systems use motion model to approximate the geometry of moving object boundaries. Motion model can be relieved from describing discontinuities in the underlying motion field, by employing motion hint that exploits the spatial structure of reference frames to infer appropriate boundaries for the future ones. However, estimation of highly accurate motion hint is computationally demanding, in particular for high resolution video sequences. Leveraging on the advantages of homogeneous motion discovery oriented prediction, in this paper, we propose to tune the intra-domain motion uniformity for B-frames as per the frame's reference utility. Experimental results show an improved bit rate savings compared to the approach where no such selective tuning is enforced.
Ashek Ahmmed, Aous Thabit Naman, David S. Taubman
PCS3
2018 Responsive high throughput congestion control for interactive applications over SDN-enabled networks
Aous Thabit Naman, Yu Wang 0131, Hassan Habibi Gharakheili, Vijay Sivaraman, David S. Taubman
Comput. Networks5
2018 HEVC-EPIC: Fast Optical Flow Estimation From Coded Video via Edge-Preserving Interpolation
abstract
This paper presents a method leveraging coded motion information to obtain a fast, high quality motion field estimation. The method is inspired by a recent trend followed by a number of top-performing optical flow estimation schemes that first estimate a sparse set of features between two frames, and then use an edge-preserving interpolation scheme (EPIC) to obtain a piecewise-smooth motion field that respects moving object boundaries. In order to skip the time-consuming estimation of features, we propose to directly derive motion seeds from decoded HEVC block motion; we call the resulting scheme "HEVCEPIC". We propose motion seed weighting strategies that account for the fact that some motion seeds are less reliable than others. Experiments on a large variety of challenging sequences and various bit-rates show that HEVC-EPIC runs significantly faster than EPIC flow, while producing motion fields that have a slightly lower average endpoint error (A-EPE). HEVC-EPIC opens the door of seamlessly integrating HEVC motion into video analysis and enhancement tasks. When employed as input to a framerate upsampling scheme, the average Y-PSNR of the interpolated frames using HEVC-EPIC motion slightly outperforms EPIC flow across the tested bit-rates, while running an order of magnitude faster.
Dominic Rüfenacht, David S. Taubman
IEEE Trans. Image Process.2
2017 Lifting-based Illumination Adaptive Transform (LIAT) using mesh-based illumination modelling
abstract
State-of-the-art video coding techniques employ block-based illumination compensation to improve coding efficiency. In this work, we propose a Lifting-based Illumination Adaptive Transform (LIAT) to exploit temporal redundancy among frames that have illumination variations, such as the frames of low frame rate video or multi-view video. LIAT employs a mesh-based spatially affine model to represent illumination variations between two frames. In LIAT, transformed frames are jointly compressed, together with illumination information, into a layered rate-distortion optimal codestream, using the JPEG2000 format. We show that the LIAT framework significantly improves compression efficiency of temporal subband transforms for both predictive and more general transforms with predict and update steps.
Maryam Haghighat, Reji Mathew, Aous Thabit Naman, David S. Taubman
ICIP4
2017 Light-field image compression based on variational disparity estimation and motion-compensated wavelet decomposition
abstract
This paper presents a compression framework for light-field images. The main idea of our approach is exploiting the similarity across sub-aperture images extracted from light-field data to improve encoding performance. For this purpose we propose a variational optimisation approach to estimate the disparity map from light-field images and then apply it to a motion-compensated wavelet lifting scheme. Making use of JPEG2000 for coding all high-/low-pass sub-band views as well as disparity map, our approach can therefore support both lossless and lossy compression. The coding framework is tested with both synthetic and real-world light-field dataset. The experimental results demonstrate that our approach outperforms JPEG-LS and the direct application of JPEG2000 in both lossless and lossy compression scenarios.
Trung-Hieu Tran, Yousef Baroud, Zhe Wang 0008, Sven Simon 0001, David S. Taubman
ICIP5
2017 HEVC-EPIC: Edge-preserving interpolation of coded HEVC motion with applications to framerate upsampling
abstract
We propose a method to obtain a high quality motion field from decoded HEVC motion. We use the block motion vectors to establish a sparse set of correspondences, and then employ an affine, edge-preserving interpolation of correspondences (EPIC) to obtain a dense optical flow. Experimental results on a variety of sequences coded at a range of QP values show that the proposed HEVC-EPIC is over five times as fast as the original EPIC flow, which uses a sophisticated correspondence estimator, while only slightly decreasing the flow accuracy. The proposed work opens the door to leveraging HEVC motion into video enhancement and analysis methods. To provide some evidence of what can be achieved, we show that when used as input to a framerate upsampling scheme, the average Y-PSNR of the interpolated frames obtained using HEVC-EPIC motion is slightly lower (0.2dB) than when original EPIC flow is used, with hardly any visible differences.
Dominic Rüfenacht, David S. Taubman
ICME2
2017 Leveraging decoded HEVC motion for fast, high quality optical flow estimation
abstract
We propose a method of improving the quality of decoded HEVC motion fields attached to B-frames, in order to make them more suitable for video analysis and enhancement tasks. We use decoded HEVC motion vectors as a sparse set of motion "seeds", which guide an edge-preserving affine interpolation of coded motion (HEVC-EPIC) in order to obtain a much more physical representation of the scene motion. We further propose HEVC-EPIC-BI, which adds a bidirectional motion completion step that leverages the fact that regions which are occluded in one direction are usually visible in the other. The use of decoded motion allows us to avoid the time-consuming estimation of "seeds". Experiments on a large variety of synthetic sequences show that compared to a state-of-the-art "seed-based" optical flow estimator, the computational complexity can be reduced by 80%, while incurring no increase at in average EPE at higher bit-rates, and a slight increase of 0.09 at low bit-rates.
Dominic Rüfenacht, David S. Taubman
MMSP2
2017 Progressive Dictionary Learning With Hierarchical Predictive Structure for Low Bit-Rate Scalable Video Coding
abstract
Dictionary learning has emerged as a promising alternative to the conventional hybrid coding framework. However, the rigid structure of sequential training and prediction degrades its performance in scalable video coding. This paper proposes a progressive dictionary learning framework with hierarchical predictive structure for scalable video coding, especially in low bitrate region. For pyramidal layers, sparse representation based on spatio-temporal dictionary is adopted to improve the coding efficiency of enhancement layers with a guarantee of reconstruction performance. The overcomplete dictionary is trained to adaptively capture local structures along motion trajectories as well as exploit the correlations between the neighboring layers of resolutions. Furthermore, progressive dictionary learning is developed to enable the scalability in temporal domain and restrict the error propagation in a closed-loop predictor. Under the hierarchical predictive structure, online learning is leveraged to guarantee the training and prediction performance with an improved convergence rate. To accommodate with the state-of-the-art scalable extension of H.264/AVC and latest High Efficiency Video Coding (HEVC), standardized codec cores are utilized to encode the base and enhancement layers. Experimental results show that the proposed method outperforms the latest scalable extension of HEVC and HEVC simulcast over extensive test sequences with various resolutions.
Wenrui Dai, Yangmei Shen, Hongkai Xiong, Xiaoqian Jiang, Junni Zou, David S. Taubman
IEEE Trans. Image Process.6
2016 Efficient action recognition from compressed depth maps
abstract
We propose an efficient action recognition scheme based solely on compressed depth maps. Each depth map is coded by a recently proposed scalable encoder that employs multi-scale breakpoints and an adaptive discrete wavelet transform (DWT). DWT coefficients describe smooth variations in depth while breakpoints communicate sharp boundaries. Both of these attributes are extracted from the bit-stream and utilized to construct features which are subject to a classification scheme for human action recognition. By extracting features from the compressed bit-stream computational complexity is significantly reduced thereby making the proposed scheme suitable for real-time applications. A L2-regularized collaborative representation classifier is employed for classification. The proposed scheme is computationally more efficient when compared with conventional approaches. Experimental results on the MSR 3D action dataset validate the effectiveness and efficiency of our proposed scheme.
Jie Miao, Xiaoyi Jia, Reji Mathew, Xiangmin Xu 0001, David S. Taubman, Chunmei Qing
ICIP5
2016 Optimization and compression of geometry discontinuities for graph-based representation of piecewise smooth media
abstract
Earlier research has shown the efficacy of using geometry discontinuities, encoded using “breakpoints,” in improving the coding efficiency of piecewise smooth media, such as depth maps and motion flows. This work proposes a new structure for encoding these breakpoints that is more suited for piecewise affine media, such as affine motion flows. Here, we choose to employ belief propagation over a graphical model to discover a set of breakpoints that is rate-distortion optimal. Similar to earlier works, the discovered breakpoints are employed in a breakpoint adaptive (BPA) wavelet decomposition of the media under consideration; the resulting BPA wavelet coefficients and the proposed “affine” breakpoints are then compressed using the JPEG2000 format, in a way that provides spatial and quality scalability as well as accessibility. We show that the coding cost of affine breakpoints is moderate, and that, for coding motion flows, the coding efficiency of the proposed approach is comparable to H.264.
Aous Thabit Naman, David S. Taubman, Reji Mathew
ICIP2
2016 Optimizing block-coded motion parameters with block-partition graphs
abstract
We address the problem of optimizing block-coded motion parameters for use inside typical motion-compensating video encoders. We cast the given discrete problem as a nonsmooth nonconvex optimization problem which is defined over some graph, and solve it using the split primal-dual hybrid gradient algorithm. Although computational efficiency is not the main focus of this paper, an efficient, parallelized implementation of our proposed approach can be used as a way of performing rate-distortion optimal motion estimation in video encoders such as those following the H.264 or HEVC standard. Results from our experiments highlight the degree of sub-optimality demonstrated by motion parameters that have been computed by H.264 block matching algorithms.
Sean I. Young, Reji Mathew, David S. Taubman
ICIP3
2016 Temporally consistent high frame-rate upsampling with motion sparsification
abstract
This paper continues our work on occlusion-aware temporal frame interpolation (TFI) that employs piecewise-smooth motion with sharp motion boundaries. In this work, we propose a triangular mesh sparsification algorithm, which allows to trade off computational complexity with reconstruction quality. Furthermore, we propose a method to create a background motion layer in regions that get disoccluded between the two reference frames, which is used to get temporally consistent interpolations among frames interpolated between the two reference frames. Experimental results on a large data set show the proposed mesh sparsification is able to reduce the processing time by 75%, with a minor drop in PSNR of 0.02 dB. The proposed TFI scheme outperforms various state-of-the-art TFI methods in terms of quality of the interpolated frames, while having the lowest processing times. Further experiments on challenging synthetic sequences highlight the temporal consistency in traditionally difficult regions of disocclusion.
Dominic Rüfenacht, David S. Taubman
MMSP2
2016 Homogeneous motion discovery oriented reference frame for high efficiency video coding
abstract
Traditional video coding uses the motion model to approximate geometric boundaries of moving objects where motion discontinuities occur. Motion hints based inter-frame prediction paradigm moves away from this redundant approach and employs an innovative framework consisting of motion hint fields that are continuous and invertible, at least, over their respective domains. However, estimation of motion hint is computationally demanding, in particular for high resolution video sequences. In this paper, we propose to discover motion models and their associated masks over the current frame and then use these models and masks to form a prediction of the current frame. The prediction process is computationally simpler and experimental results show that a savings in bit rate of 2.3% is achievable over standalone HEVC if this predicted frame is used as an additional reference frame.
Ashek Ahmmed, David S. Taubman, Aous Thabit Naman, Mark R. Pickering
PCS2
2016 Higher-order motion models for temporal frame interpolation with applications to video coding
abstract
We have recently proposed a motion-centric temporal frame interpolation (TFI) method, called BAM-TFI, which is able to produce high quality interpolated frames under a constant velocity assumption. However, for objects that do not follow constant velocity motion, the predictions, although credible, will differ from the “true” target frames, leading to high prediction residuals. In this paper, we show how higher-order motion models can be incorporated into the BAM-TFI scheme to interpolate frames that better predict the target frames. This opens up the door to a seamless integration of TFI with a video coding scheme. Comparisons on a variety of both synthetic and natural video sequences highlight the benefits of a second-order motion model. We further integrate the proposed TFI scheme into HEVC; preliminary comparisons with HEVC show promising results.
Dominic Rüfenacht, Reji Mathew, David S. Taubman
PCS3
2016 Optimized decoding of JPEG images based on generalized graph Laplacians
abstract
We address the problem of optimizing the decoding of JPEG-compressed images, employing an approach based on the “generalized” graph Laplacian, a higher-order generalization of the usual graph Laplacian. The optimal decoding problem is formulated as a non-smooth but convex problem over a graph, and solved via the alternating directions method of multipliers. While similar, graph-based optimized decoding techniques exist, the use of our generalized graph Laplacian enables better recovery of the original smooth image from the transform coefficients, especially those coded at low bit-rates. Experimental results highlight the performance of our proposed approach in terms of reconstructed distortion and visual quality. Comparisons with the method based on the usual graph Laplacian and other state-of-the-art methods are also given.
Sean I. Young, Aous Thabit Naman, Reji Mathew, David S. Taubman
PCS4
2016 A Novel Motion Field Anchoring Paradigm for Highly Scalable Wavelet-Based Video Coding
abstract
Existing video coders anchor motion fields at frames that are to be predicted. In this paper, we demonstrate how changing the anchoring of motion fields to reference frames has some important advantages over conventional anchoring. We work with piecewise-smooth motion fields, and use breakpoints to signal discontinuities at moving object boundaries. We show how discontinuity information can be used to resolve double mappings arising when motion is warped from reference to target frames. We present an analytical model that allows to determine weights for texture, motion, and breakpoints to guide the rate-allocation for scalable encoding. Compared with the conventional way of anchoring motion fields, the proposed scheme requires fewer bits for the coding of motion; furthermore, the reconstructed video frames contain fewer ghosting artefacts. The experimental results show the superior performance compared with the traditional anchoring, and demonstrate the high scalability attributes of the proposed method.
Dominic Rüfenacht, Reji Mathew, David S. Taubman
IEEE Trans. Image Process.3
2016 Motion Estimation Based on Mutual Information and Adaptive Multi-Scale Thresholding
abstract
This paper proposes a new method of calculating a matching metric for motion estimation. The proposed method splits the information in the source images into multiple scale and orientation subbands, reduces the subband values to a binary representation via an adaptive thresholding algorithm, and uses mutual information to model the similarity of corresponding square windows in each image. A moving window strategy is applied to recover a dense estimated motion field whose properties are explored. The proposed matching metric is a sum of mutual information scores across space, scale, and orientation. This facilitates the exploitation of information diversity in the source images. Experimental comparisons are performed amongst several related approaches, revealing that the proposed matching metric is better able to exploit information diversity, generating more accurate motion fields.
Rui Xu 0030, David S. Taubman, Aous Thabit Naman
IEEE Trans. Image Process.2
2015 Rate-distortion optimized optical flow estimation
abstract
We propose a general framework for rate-distortion optimized estimation of optical flow, taking heed of the optimality conditions when both the residue and the motion coefficients are quantized and coded. The framework is particularly well suited for wavelet estimation of motion, where the estimated coefficients can additionally be coded in a scalable fashion. We also give a summary of two majorization-minimization type of optimization algorithms that can be used to solve our non-linear and non-convex rate-distortion optimization problem. We show experimentally that the proposed scheme produces motion descriptions that are superior to those produced by other optical flow algorithms, in the rate and distortion sense.
Sean I. Young, David S. Taubman
ICIP2
2015 Motion blur modelling for hierarchically anchored motion with discontinuities
abstract
We have previously proposed a scheme for representing motion with motion discontinuities which has beneficial properties in terms of compactness (efficiency) and scalability. This so-called BIHA scheme has applications in video coding as well as temporal frame interpolation. In both cases, modelling of motion discontinuities has proven to be valuable. In these earlier works, we have ignored effects of motion blur, which can result in artificial sharp transitions of texture information at moving object boundaries. In this paper, we extend the BIHA framework to account for motion blur. Experimental results show significant improvements over the original BIHA scheme in texture blending regions, resulting in more visually pleasing predictions, as well as better rate-distortion performance.
Dominic Rüfenacht, Reji Mathew, David S. Taubman
MMSP3
2015 Motion hints mode for macroblock coding in bi-predictive slices
abstract
Recent advances in motion modelling have largely focused on careful partitioning of motion blocks in the vicinity of object boundaries. The need for such fine partitioning can be avoided by using motion hints which provide a global description of motion over specific domains. Experimental results show that, with a hybrid setting, more than 50% of the motion discontinuity macroblocks are coded using the motion hints mode in low bit rate cases. The use of this mode leads to a gain of prediction PSNR of 1.11 dB, or equivalently 17.05% savings in bit rate, when compared to the H.264/AVC reference and considering both low and high bit rate applications.
Ashek Ahmmed, Md. Jahangir Alam 0005, Aous Thabit Naman, Mark R. Pickering, David S. Taubman
PCS5
2015 Optimization of optical flow for scalable coding
abstract
Optical flow or dense motion field representations provide an alternative to block based schemes and are capable of describing smooth motion flows without introducing any artificial block boundaries. However dense motion fields have not been widely adopted for video coding applications principally due to the prohibitive cost of communicating such detailed motion representations. In this paper our focus is on estimating dense motion fields subject to R-D constraints derived from wavelet based coding of the motion data. We develop a probabilistic framework for estimating dense motion fields that is capable of evaluating the tradeoff between sparsity in the transformed domain and motion compensated distortion. R-D results show that we are able to improve upon traditional optical flow descriptions.
Reji Mathew, Sean I. Young, David S. Taubman
PCS3
2015 Bidirectional, occlusion-aware temporal frame interpolation in a highly scalable video setting
abstract
We present a bidirectional, occlusion-aware temporal frame interpolation (BOA-TFI) scheme that builds upon our recently proposed highly scalable video coding scheme. Unlike previous TFI methods, our scheme attempts to put “correct” information in problematic regions around moving objects. From a “parent” motion field between two existing reference frames, we compose motion from both reference frames to the target frame. These motion fields, together with motion discontinuity information, are then warped to the target frame - a process during which we discover valuable information about disocclusions, which we then use to guide the bidirectional prediction of the interpolated frame. The scheme can be used in any state-of-the-art codec, but is most beneficial if used in conjunction with a highly scalable video coder. Evaluation of the method on synthetic data allows us to shine a light on problematic regions around moving object boundaries, which has not been the focus of previous frame interpolation methods. The proposed frame interpolation method yields credible results, and compares favourably to current state-of-the-art frame interpolation methods.
Dominic Rüfenacht, Reji Mathew, David S. Taubman
PCS3
2015 Spatial induction policies for scalable depth coding
abstract
An edge-adaptive depth coding proposed in [1] uses a multi-resolution field of breakpoints to encode edges in a scene's geometry. In order to remain efficient, this scheme uses spatial induction to infer the location of breakpoints from the surrounding geometry where possible. This intelligence means only a subset of the breakpoints need to be encoded. The original proposal employs a simple policy, using linear interpolation between known points to induce breakpoints. Whilst this is effective, a more sophisticated policy could better exploit the available breakpoint data in synthesising edges. This paper presents our findings regarding the potential benefits of higher order spatial induction policies (SIPs).
Mitchell S. Ward, David S. Taubman, Reji Mathew
PCS2
2015 Motion estimation with accurate boundaries
abstract
This paper investigates several techniques that increase the accuracy of motion boundaries in estimated motion fields of a local dense estimation scheme. In particular, we examine two matching metrics, one is MSE in the image domain and the other one is a recently proposed multiresolution metric that has been shown to produce more accurate motion boundaries. We also examine several different edge-preserving filters. The edge-aware moving average filter, proposed in this paper, takes an input image and the result of an edge detection algorithm, and outputs an image that is smooth except at the detected edges. Compared to the adoption of edge-preserving filters, we find that matching metrics play a more important role in estimating accurate and compressible motion fields. Nevertheless, the proposed filter may provide further improvements in the accuracy of the motion boundaries. These findings can be very useful for a number of recently proposed scalable interactive video coding schemes.
Rui Xu 0030, Aous Thabit Naman, Reji Mathew, Dominic Rüfenacht, David S. Taubman
PCS5
2014 Overlapping motion hints with polynomial motion for video communication
abstract
In a recent work, we propose the use of motion hints for communicating motion. A motion hint describes motion that is accurate (describes the actual motion) for only a region inside a domain associated with that motion hint; it is the job of the client or decoder to decide the exact region of applicability (ROA). The motion described by a motion hint is invertible and global; that is, it allows the prediction of the ROA associated with a motion hint from any frame that has that hint. Motion hints are applicable to closed-loop prediction, but they are more useful in open-loop prediction scenarios, such as remote browsing of surveillance footage, communicated by a JPIP server, which is the focus of this work. This work proposes a probabilistic multi-scale framework to identifying the ROA; the framework is applicable to multiple overlapping motion hints. The proposed approach is localized (and therefore amenable to parallel processing) and robust to noise, quantization, and changes in contrast. We show that motion hints can be used for real video sequences, and we also present results for the case of three overlapping motion hints (one background and two overlapping foregrounds).
Aous Thabit Naman, David S. Taubman, Rui Xu 0030
ICIP2
2014 Hierarchical anchoring of motion fields for fully scalable video coding
abstract
Traditional video codecs anchor motion fields in the frame that is to be predicted, which is natural in a non-scalable context. In this paper, we propose a hierarchical anchoring of motion fields at reference frames, which allows to “reuse” them at finer temporal levels - a very desirable property for temporal scalability. The main challenge using this approach is that the motion fields need to be warped to the target frames, leading to disocclusions and motion folding in the warped motion fields. We show how to resolve motion folding ambiguities that occur in the vicinity of moving object boundaries by using breakpoint fields that have recently been proposed for the scalable coding of motion. During the motion field warping process, we obtain disocclusion and folding maps on-the-fly, which are used to control the temporal update step of the Haar wavelet. Results on synthetic data show that the proposed hierarchical anchoring scheme outperforms the traditional way of anchoring motion fields.
Dominic Rüfenacht, Reji Mathew, David S. Taubman
ICIP3
2014 Block motion matching on directional subbands with interband suppression
abstract
We propose a novel matching metric for dense block-based true-motion estimation. Block-based matching scores are calculated on directional detail bands of a steerable pyramid after conversion to a 2-bit representation. The proposed non-linear transform involves a novel inter-band suppression mechanism so that matching scores can be accumulated across resolutions and directions. We show that edge features of different strength that occupy similar locations can be isolated in different directional bands, after which the 2-bit transformation effectively equalises their contrast. This considerably improves the robustness of motion estimation procedure when compared to other matching metrics, including our prior work and the commonly used MSE.
Rui Xu 0030, David S. Taubman, Aous Thabit Naman
ICIP2
2014 Bidirectional hierarchical anchoring of motion fields for scalable video coding
abstract
The ability to predict motion fields at finer temporal scales from coarser ones is a very desirable property for temporal scalability. This is at best very difficult in current state-of-the-art video codecs (i.e., H.264, HEVC), where motion fields are anchored in the frame that is to be predicted (target frame). In this paper, we propose to anchor motion fields in the reference frames. We show how from only one fully coded motion field at the coarsest temporal level as well as breakpoints which signal discontinuities in the motion field, we are able to reliably predict motion fields used at finer temporal levels. This significantly reduces the cost for coding the motion fields. Results on synthetic data show improved rate-distortion (R-D) performance and superior scalability, when compared to the traditional way of anchoring motion fields.
Dominic Rüfenacht, Reji Mathew, David S. Taubman
MMSP3
2014 Embedded coding of optical flow fields for scalable video compression
abstract
An embedded coding scheme for dense motion (optical flow) fields is proposed. Such a scheme is particularly useful in scalable video compression where one must compensate for inter-frame motion at various visual qualities and resolutions. However, the high cost of coding such fields has often made this option prohibitive. Using our previously developed `breakpoint'-adaptive wavelet transform, we show that it is possible to code dense motion fields efficiently while simultaneously endowing the coded motion representation with embedded resolution and quality scalability attributes. Performance comparisons with the traditional non-scalable block-based model are also made and presented with the aid of a modified H.264/AVC JM reference encoder.
Sean I. Young, Reji Mathew, David S. Taubman
MMSP3
2014 Flexible Synthesis of Video Frames Based on Motion Hints
abstract
In this paper, we propose the use of "motion hints" to produce interframe predictions. A motion hint is a loose and global description of motion that can be communicated using metadata; it describes a continuous and invertible motion model over multiple frames, spatially overlapping other motion hints. A motion hint provides a reasonably accurate description of motion but only a loose description of where it is applicable; it is the task of the client to identify the exact locations where this motion model is applicable. The focus of this paper is a probabilistic multiscale approach to identifying these locations of applicability; the method is robust to noise, quantization, and contrast changes. The proposed approach employs the Laplacian pyramid; it generates motion hint probabilities from observations at each scale of the pyramid. These probabilities are then combined across the scales of the pyramid starting from the coarsest scale. The computational cost of the approach is reasonable, and only the neighborhood of a pixel is employed to determine a motion hint probability, which makes parallel implementation feasible. This paper also elaborates on how motion hint probabilities are exploited in generating interframe predictions. The scheme of this paper is applicable to closed-loop prediction, but it is more useful in open-loop prediction scenarios, such as using prediction in conjunction with remote browsing of surveillance footage, communicated by a JPEG2000 Interactive Protocol (JPIP) server. We show that the interframe predictions obtained using the proposed approach are good both visually and in terms of PSNR.
Aous Thabit Naman, David S. Taubman
IEEE Trans. Image Process.2
2014 Nonlinear Transform for Robust Dense Block-Based Motion Estimation
abstract
We present a noniterative multiresolution motion estimation strategy, involving block-based comparisons in each detail band of a Laplacian pyramid. A novel matching score is developed and analyzed. The proposed matching score is based on a class of nonlinear transformations of Laplacian detail bands, yielding 1-bit or 2-bit representations. The matching score is evaluated in a dense full-search motion estimation setting, with synthetic video frames and an optical flow data set. Together with a strategy for combining the matching scores across resolutions, the proposed method is shown to produce smoother and more robust estimates than mean square error (MSE) in each detail band and combined. It tolerates more of nontranslational motion, such as rotation, validating the analysis, while providing much better localization of the motion discontinuities. We also provide an efficient implementation of the motion estimation strategy and show that the computational complexity of the approach is closely related to the traditional MSE block-based full-search motion estimation procedure.
Rui Xu 0030, David S. Taubman, Aous Thabit Naman
IEEE Trans. Image Process.2
2013 A soft measure for identifying structure from randomness in images
abstract
This paper presents a novel measure for identifying strong structure features, such as edges, from randomness, such as regions predominated by noise, within an image. The proposed structural measure is localized in space and scale; for a given scale, it gives values close to one in the vicinity of strong structures and close to zero in regions predominated by noise. The proposed structural measure is a primitive operation that can be used in a wide variety of image analysis techniques to identify regions which has structure; for example, motion estimation is more meaningful in structured regions than in regions filled with noise. The first innovation in this work is in converting an image into a ternary feature map that are rather resistant to noise and changes in illumination. The second is the structural measure, which is derived from the degree of non-uniformity amongst the magnitudes of the DFT coefficients obtained over a small window within the ternary maps. In this work, we show that the proposed structural measure is robust and gives a good indication of the strength of structure when compared to alternate strategies; moreover, we show that the computational cost of the proposed structural measure is reasonable.
Aous Thabit Naman, David S. Taubman
ICIP2
2013 Inter-frame prediction using motion hints
abstract
We recently proposed a novel approach that employs motion hints for inter-frame prediction. Motion hints are a loose and global description of motion communicated as metadata; they specify motion but they leave it to the client/decoder to find the exact locations where motion is applicable. This work proposes a multi-scale approach for identifying these exact locations, which are then used with the available reference frames to generate an inter-frame prediction. The proposed approach is localized and robust to noise and illumination changes. The scheme of this work is applicable to close-loop prediction, but it is more useful in open-loop prediction scenarios, such as using prediction in conjunction with remote browsing of surveillance footage, communicated by a JPIP server. We show that, with reasonably accurate motion, it is possible to produce good inter-frame predictions visually and in terms of PSNR.
Aous Thabit Naman, Rui Xu 0030, David S. Taubman
ICIP3
2013 Robust dense block-based motion estimation using a 2-bit transform on a Laplacian pyramid
abstract
We propose a non-iterative multi-resolution motion estimation strategy, involving block-based comparisons in each detail band of a Laplacian pyramid. A novel matching score is developed, together with a novel strategy for combining the matching scores across resolutions. The proposed matching score is based on a non-linear transformation of the Laplacian detail subbands, yielding a 2-bit representation that also has computational advantages. In a dense motion estimation setting, with synthetic content, the proposed method is shown to produce smoother and more robust estimates than MSE, including a multi-resolution MSE matching strategy, while providing much better localization of the motion discontinuities.
Rui Xu 0030, David S. Taubman
ICIP2
2013 Motion segmentation initialization strategies for bi-directional inter-frame prediction
abstract
Experimental results and the latest standards have proved that segmentation based video coding systems can outperform the traditional block-based video coding systems. However, this approach requires the simultaneous estimation of both the shape and motion of moving objects in a video scene. In most of the cases neither the shape nor the motion are known initially. Another critical aspect of this tightly-coupled relationship is that inaccurate motion estimation may cause poor segmentation and erroneous segmentation may negatively impact motion estimation. While some of the existing approaches require user intervention and some use clues such as depth, colour or occlusion to separate the foreground from the background, we propose to use motion reliability information for this purpose. This is because the ingredients necessary for the calculation of motion reliability are the by-product of block-based motion estimation and compensation between the reference frames. Therefore, they require very little or no increase in the computational overhead. In this paper, we explore several motion segmentation initialization strategies based on motion reliability. The performances of these initialization approaches are investigated, in terms of the PSNR, for the predicted inter-frames.
Ashek Ahmmed, Rui Xu 0030, Aous Thabit Naman, Md. Jahangir Alam 0005, Mark R. Pickering, David S. Taubman
MMSP6
2013 Motion hints based inter-frame prediction for hybrid video coding
abstract
Experimental results and the latest standards have proved video coding systems with the ability to adapt the size and shape of the motion estimation area to the objects in the scene can outperform the traditional block-based video coding systems. In this paper, a segmentation-based coding strategy that employs bi-directional motion hints for interframe prediction is proposed. The appealing thing about motion hints is that they are continuous and invertible, even though the observed motion field for a frame will be discontinuous and non-invertible. The proposed scheme outperforms the rate-distortion performance of H.264/AVC reference by 1.1 dB and a bit rebate of 26.6% is achieved.
Ashek Ahmmed, Md. Jahangir Alam 0005, Mark R. Pickering, Rui Xu 0030, Aous Thabit Naman, David S. Taubman
PCS6
2013 Robust sum of Linear-Log Squared Differences distortion measure and its applications
abstract
A robust distortion measure known as the Sum of Linear-Log Squared Differences (SLLSD) derived analytically from the R-D optimality conditions for compression is proposed with particular applications in motion estimation (ME) and coding. When used in the context of ME, this new measure is shown both to improve motion robustness (accuracy) and lower the Lagrangian cost (coding) unlike the typically employed Sum of Squared Differences (SSD) measure. Our measure is defined in the image domain and does not necessitate or assume the use of a particular transform. Its relationship to other M-estimators is briefly discussed and coding related performance results presented with the aid of a suitably modified H.264/AVC JM reference encoder.
Sean I. Young, Reji Mathew, David S. Taubman
PCS3
2013 Augmented Active Surface Model for the Recovery of Small Structures in CT
abstract
This paper devises an augmented active surface model for the recovery of small structures in a low resolution and high noise setting, where the role of regularization is especially important. The emphasis here is on evaluating performance using real clinical computed tomography (CT) data with comparisons made to an objective ground truth acquired using micro-CT. In this paper, we show that the application of conventional active contour methods to small objects leads to non-optimal results because of the inherent properties of the energy terms and their interactions with one another. We show that the blind use of a gradient magnitude based energy performs poorly at these object scales and that the point spread function (PSF) is a critical factor that needs to be accounted for. We propose a new model that augments the external energy with prior knowledge by incorporating the PSF and the assumption of reasonably constant underlying CT numbers.
Andrew Philip Bradshaw, David S. Taubman, Michael J. Todd, John S. Magnussen, G. Michael Halmagyi
IEEE Trans. Image Process.2
2013 Scalable Coding of Depth Maps With R-D Optimized Embedding
abstract
Recent work on depth map compression has revealed the importance of incorporating a description of discontinuity boundary geometry into the compression scheme. We propose a novel compression strategy for depth maps that incorporates geometry information while achieving the goals of scalability and embedded representation. Our scheme involves two separate image pyramid structures, one for breakpoints and the other for sub-band samples produced by a breakpoint-adaptive transform. Breakpoints capture geometric attributes, and are amenable to scalable coding. We develop a rate-distortion optimization framework for determining the presence and precision of breakpoints in the pyramid representation. We employ a variation of the EBCOT scheme to produce embedded bit-streams for both the breakpoint and sub-band data. Compared to JPEG 2000, our proposed scheme enables the same the scalability features while achieving substantially improved rate-distortion performance at the higher bit-rate range and comparable performance at the lower rates.
Reji Mathew, David S. Taubman, Pietro Zanuttigh
IEEE Trans. Image Process.2
2013 PET Protection Optimization for Streaming Scalable Videos With Multiple Transmissions
abstract
This paper investigates priority encoding transmission (PET) protection for streaming scalably compressed video streams over erasure channels, for the scenarios where a small number of retransmissions are allowed. In principle, the optimal protection depends not only on the importance of each stream element, but also on the expected channel behavior. By formulating a collection of hypotheses concerning its own behavior in future transmissions, limited-retransmission PET (LR-PET) effectively constructs channel codes spanning multiple transmission slots and thus offers better protection efficiency than the original PET. As the number of transmission opportunities increases, the optimization for LR-PET becomes very challenging because the number of hypothetical retransmission paths increases exponentially. As a key contribution, this paper develops a method to derive the effective recovery-probability versus redundancy-rate characteristic for the LR-PET procedure with any number of transmission opportunities. This significantly accelerates the protection assignment procedure in the original LR-PET with only two transmissions, and also makes a quick and optimal protection assignment feasible for scenarios where more transmissions are possible. This paper also gives a concrete proof to the redundancy embedding property of the channel codes formed by LR-PET, which allows for a decoupled optimization for sequentially dependent source elements with convex utility-length characteristic. This essentially justifies the source-independent construction of the protection convex hull for LR-PET.
Ruiqin Xiong, David S. Taubman, Vijay Sivaraman
IEEE Trans. Image Process.2
2012 Highly Scalable Coding of Depth Maps with Arc Breakpoints
abstract
Recent work highlights the importance of incorporating geometry information into the compression of depth maps. For many applications, features such as resolution scalability and embedded coding are also highly desirable. JPEG 2000 offers these scalability features but suffers from poor compression performance in the vicinity of strong discontinuities. We propose a novel compression strategy for depth maps that incorporates geometry information while retaining the highly scalable coding properties of JPEG 2000. Our scheme involves two separate image pyramid structures, one for arc breakpoints and other for sub-band samples produced by a breakpoint-adaptive transform. Breakpoints capture geometric attributes and are also amenable to scalable coding. We develop an R-D optimization framework for the breakpoint data. We also use a variation of the EBCOT scheme to produce embedded bit-streams for both the breakpoint and sub-band data, allowing them to be independently and incrementally sequenced based on R-D considerations.
Reji Mathew, Pietro Zanuttigh, David S. Taubman
DCC3
2012 Scalable depth maps with R-D optimized embedding
abstract
Recent work has highlighted the importance of incorporating geometry information into the compression of depth maps. In prior approaches however the geometry information is not resolution scalable nor amenable to embedded coding. In this paper we propose a novel compression strategy for depth maps that incorporates geometry information while achieving the goals of scalability and embedded representation. Our scheme involves two separate image pyramid structures, one for breakpoints and other for sub-band samples produced by a breakpoint-adaptive transform. Breakpoints capture geometric attributes and are amenable to scalable coding. We develop an R-D optimization framework for the breakpoint data. We also use a variation of the EBCOT scheme to produce embedded bit-streams for both the breakpoint and sub-band data, allowing them to be independently and incrementally sequenced based on R-D considerations.
Reji Mathew, David S. Taubman, Pietro Zanuttigh
MMSP2
2012 Multithreaded processing paradigms for JPEG2000
abstract
We provide an overview of techniques that are used in a highly efficient software implementation of the JPEG2000 compression standard. We describe a novel multithreaded processing paradigm that achieves near perfect parallel processing scalability, at least over the 8 logical processors for which results are reported. Both thread scalability and overall throughput are shown to be far better than has previously been reported in the literature. For example, we demonstrate realtime processing of 2K and 4K cinema content at 244Mbits/s on a laptop computer. We also demonstrate 9/7 DWT throughputs on a single CPU core with comparable performance to a parallel multi-GPU solution. Finally, we provide results demonstrating the throughput and compression performance of a new “fast mode” that is the subject of an amendment to the standard.
David S. Taubman
MMSP1
2011 Coupled distributed arithmetic coding
abstract
In this paper, we propose a novel scheme of coupled distributed arithmetic coding to overcome the de-synchronization problem caused by causal decoding in existing distributed arithmetic coding system. Simulation results show that decoding performance is significantly improved and longer sequences outperform shorter sequences using this approach.
David S. Taubman
ICIP2
2011 Efficient communication of video using metadata
abstract
In traditional video coding schemes, motion information is tightly coupled to the prediction strategy. In this preliminary work, we depart from this model by utilizing metadata to convey motion information to the client; in particular, metadata conveys crude boundaries of objects together with motion information for these objects. Here, we are interested in applications where metadata itself carries semantics that the client is interested in, such as tracking information in surveillance applications. To keep things simple, we focus on the case where we have a single object to track. Therefore, we model each frame as a background region with a foreground region/object, enclosing each region by a quadrilateral that identifies it. The foreground quadrilateral does not follow the exact boundaries of the foreground object; it leaves the task of identifying these boundaries to the client. The advantages of metadata is that it provides a global representation of motion, which allows predicting a given object from potentially all the frames that contain that object. The approach is applicable in fully open loop systems such as in the case of the JPEG2000-Based Scalable Interactive Video (JSIV) paradigm. In this work, we present the concepts behind the proposed approach and detail the modifications introduced to the JSIV server and client policies, presenting some promising preliminary results.
Aous Thabit Naman, Duncan Edwards, David S. Taubman
ICIP3
2011 Scalable Modeling of Motion and Boundary Geometry With Quad-Tree Node Merging
abstract
Quad-tree structures are often used to model motion between frames of a video sequence. However, a fundamental limitation of the quad-tree structure is that it can only capture horizontal and vertical edge discontinuities at dyadically related locations. To address this limitation, recent work has focused on the introduction of geometry information to nodes of tree structured motion representations. In this paper we explore modeling boundary geometry and motion with separate quad-tree structures; thereby enabling each attribute to be refined separately. Recent work into quad-tree representations has also highlighted the benefits of leaf merging. We extend the leaf merging paradigm to incorporate both geometry and motion attributes. We also explore resolution scalability of the merged dual tree representation and present experimental results which demonstrate both rate-distortion and scabalility performance. Theoretical investigations conducted in this paper reveal that to achieve optimal rate-distortion behavior, quad-tree motion models need to incorporate both geometry information and node merging.
Reji Mathew, David S. Taubman
IEEE Trans. Circuits Syst. Video Technol.2
2011 A Filtering Approach to Edge Preserving MAP Estimation of Images
abstract
The authors present a computationally efficient technique for maximum a posteriori (MAP) estimation of images in the presence of both blur and noise. The image is divided into statistically independent regions. Each region is modelled with a WSS Gaussian prior. Classical Wiener filter theory is used to generate a set of convex sets in the solution space, with the solution to the MAP estimation problem lying at the intersection of these sets. The proposed algorithm uses an underlying segmentation of the image, and a means of determining the segmentation and refining it are described. The algorithm is suitable for a range of image restoration problems, as it provides a computationally efficient means to deal with the shortcomings of Wiener filtering without sacrificing the computational simplicity of the filtering approach. The algorithm is also of interest from a theoretical viewpoint as it provides a continuum of solutions between Wiener filtering and Inverse filtering depending upon the segmentation used. We do not attempt to show here that the proposed method is the best general approach to the image reconstruction problem. However, related work referenced herein shows excellent performance in the specific problem of demosaicing.
David Humphrey, David S. Taubman
IEEE Trans. Image Process.2
2011 JPEG2000-Based Scalable Interactive Video (JSIV)
abstract
We propose a novel paradigm for interactive video streaming and we coin the term JPEG2000-based scalable interactive video (JSIV) for it. JSIV utilizes JPEG2000 to independently compress the original video sequence frames and provide for quality and spatial resolution scalability. To exploit interframe redundancy, JSIV utilizes prediction and conditional replenishment of code-blocks aided by a server policy that optimally selects the number of quality layer for each code-block transmitted and a client policy that makes most of the received (distorted) frames. It is also possible for JSIV to employ motion compensation; however, we leave this topic to future work. To optimally solve the server transmission problem, a Lagrangian-style rate-distortion optimization procedure is employed. In JSIV, a wide variety of frame prediction arrangements can be employed including hierarchical B-frames of the scalable video coding (SVC) extension of the H.264/AVC standard. JSIV provides considerably better interactivity compared to existing schemes and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. Experimental results for surveillance footage, which does not suffer from the absence of motion compensation, show that JSIV's performance is comparable to that of SVC in some usage scenarios while JSIV performs better in others.
Aous Thabit Naman, David S. Taubman
IEEE Trans. Image Process.2
2011 JPEG2000-Based Scalable Interactive Video (JSIV) With Motion Compensation
abstract
In a recent work, the authors proposed a novel paradigm for interactive video streaming and coined the term JPEG2000-Based Scalable Interactive Video (JSIV) for it. In this work, we investigate JSIV when motion compensation is employed to improve prediction, something that was intentionally left out in our earlier treatment. JSIV relies on three concepts: storing the video sequence as independent JPEG2000 frames to provide quality and spatial resolution scalability, prediction and conditional replenishment of code-blocks to exploit inter-frame redundancy, and loosely coupled server and client policies in which a server optimally selects the number of quality layers for each code-block transmitted and a client makes the most of the received (distorted) frames. In JSIV, the server transmission problem is optimally solved using Lagrangian-style rate-distortion optimization. The flexibility of JSIV enables us to employ a wide variety of frame prediction arrangements, including hierarchical B-frames. JSIV provides considerably better interactivity compared with existing schemes and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. Experimental results show that JSIV's performance is inferior to that of SVC in conventional streaming applications while JSIV performs better in interactive browsing applications.
Aous Thabit Naman, David S. Taubman
IEEE Trans. Image Process.2
2010 Distributed source coding based on punctured conditional arithmetic codes
abstract
In this paper, we describe properties of punctured conditional arithmetic code for distributed source coding. A novel scheme using hierarchical interleaving has been developed to overcome non-stationarity in side channel statistics. The simulation results show that this proposed approach has excellent sphere-packing properties and outperforms the one without interleaving.
David S. Taubman
ICIP2
2010 Predictor selection using quantization intervals in JPEG2000-Based Scalable Interactive Video (JSIV)
abstract
The authors have recently introduced the JPEG2000-Based Scalable Interactive Video (JSIV) paradigm. JSIV relies on JPEG2000 format for providing scalability and accessibility, and on motion compensation and conditional replenishment to exploit temporal redundancy. JSIV can provide considerably better interactivity compared to existing video streaming practices, and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. This work extends our previous work by providing server and client policies that can exploit the client's knowledge about the quantization intervals of received samples in selecting a favorable predictor in dyadic hierarchical B-frame arrangement that does not employ motion compensation. We also demonstrate the flexibility of the JSIV paradigm by showing an improved client policy working with a non-improved server policy without any negative impact on reconstructed video.
Aous Thabit Naman, David S. Taubman
ICIP2
2010 Quad-Tree Motion Modeling With Leaf Merging
abstract
In this paper, we are concerned with the modeling of motion between frames of a video sequence. Typically, it is not possible to represent the motion between frames by a single model and therefore a quad-tree structure is often employed where smaller, variable size regions or blocks are allowed to take on separate motion models. Previous work into quad-tree representations has demonstrated the sub-optimal performance of quad-trees where the dependency between neighboring leaf nodes with different parents is not exploited. Leaf merging has been proposed to rectify this performance loss as it allows joint coding and optimization of related nodes. In this paper, we describe how the merging step can be incorporated into quad-tree motion representations for a range of motion modeling contexts. In particular, we study the impact of rate-distortion optimized merging for two motion coding schemes, these being spatially predictive coding, as used in H.264, and hierarchical coding. We present experimental results which demonstrate that node merging can provide significant gains for both the hierarchical and spatial prediction schemes. Interestingly, experimental results also show that in the presence of merging, the rate-distortion performance of hierarchical coding is comparable to that of spatial prediction. We pursue the case of hierarchical coding further in this paper, introducing polynomial motion models to the quad-tree representation and exploring resolution scalability of the merged quad-tree structure. We also present a theoretical study of the impact of leaf merging in modeling motion, identifying the inherent advantages of merging which give rise to a more efficient description of frame motion.
Reji Mathew, David S. Taubman
IEEE Trans. Circuits Syst. Video Technol.2
2010 Optimal PET Protection for Streaming Scalably Compressed Video Streams With Limited Retransmission Based on Incomplete Feedback
abstract
For streaming scalably compressed video streams over unreliable networks, Limited-Retransmission Priority Encoding Transmission (LR-PET) outperforms PET remarkably since the opportunity to retransmit is fully exploited by hypothesizing the possible future retransmission behavior before the retransmission really occurs. For the retransmission to be efficient in such a scheme, it is critical to get adequate acknowledgment from a previous transmission before deciding what data to retransmit. However, in many scenarios, the presence of a stochastic packet delay process results in frequent late acknowledgements, while imperfect feedback channels can impair the server's knowledge of what the client has received. This paper proposes an extended LR-PET scheme, which optimizes PET-protection of transmitted bitstreams, recognizing that the received feedback information is likely to be incomplete. Similar to the original LR-PET, the behavior of future retransmissions is hypothesized in the optimization objective of each transmission opportunity. As the key contribution, we develop a method to efficiently derive the effective recovery probability versus redundancy rate characteristic for the extended LR-PET communication process. This significantly simplifies the ultimate protection assignment procedure. This paper also demonstrates the advantage of the proposed strategy over several alternative strategies.
Ruiqin Xiong, David S. Taubman, Vijay Sivaraman
IEEE Trans. Image Process.2
2009 A content-adaptive wavelet-like transform for aliasing suppression in image and video compression
abstract
In the application of resolution scalable video coding, pure wavelet transforms offer attractive compression performance yet suffer from poor frequency rolloff, resulting in noticeable aliasing artifacts. In our previous work, we proposed a wavelet-based scheme which exploits spectral properties of natural images to achieve good antialiasing performance with a critically sampled transform. However, the scheme suffers from ringing artifacts and reduced energy compaction, which results in compression on par with Laplacian pyramids. In this paper we extend this scheme to be spatially adaptive. We find this achieves improvements over the non-adaptive method of up to 0.5 dB at 1 bpp, substantially reduces incidence of ringing artifacts and maintains the original property of aliasing suppression.
Jonathan Gan, David S. Taubman
ICIP2
2009 Joint scalable modeling of motion and boundary geometry with quad-tree node merging
abstract
Quad-tree structures are often used to model motion between frames of a video sequence. In this study we are interested in the bit-rate efficiency and resolution scalability of quad-tree motion models. Publications have reported improvements to motion modeling efficiency achieved by introducing geometry information to nodes of quad-tree structures. The benefits of leaf merging to tree based representations have also been highlighted. In keeping with these findings we explore scalability in the context of merged quad-tree representations, modeling the underlying motion flow with joint geometry and motion models. We employ hierarchical coding of the merged quad-tree and ensure that the merging process retains the property of resolution scalability. We show that the performance of scalable coding can be significantly improved by incorporating a new cost objective which takes into account the possibility of scalable decoding. Experimental results reveal that these scalability improvements are achieved without significant loss in overall efficiency and with competitive performance at all resolutions.
Reji Mathew, David S. Taubman
ICIP2
2009 Rate-distortion optimized JPEG2000-based scalable interactive video (JSIV) with motion and quantization bin side-information
abstract
The authors have recently proposed a paradigm that can potentially provide for considerably better interactivity compared to existing practices and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. The proposed paradigm relies on JPEG2000 format for providing scalability, flexibility, and accessibility; and on transmitting a server-optimized selection of code-blocks and motion side-information. Motion compensation and conditional replenishment are employed to reduce needed bandwidth. This work extends the previous work by providing server and client policies that allow for a realistic implementation and by introducing the use of coarsely quantized code-blocks in improving prediction. This work introduces the concepts, formulates the policies and optimization problems, proposes solutions, and compares the performance to alternate strategies.
Aous Thabit Naman, David S. Taubman
ICIP2
2009 Optimal linear detector for spread spectrum based multidimensional signal watermarking
abstract
This paper presents a new efficient spread spectrum (SS) based watermarking technique for multidimensional signals. In this method, the Karhunen Loeve transform (KLT) is used to completely decorrelate the components of the cover signal and obtain the maximum energy compaction. In order to improve the robustness, the same secret message is embedded into all components and linear fusion is used to exploit the collaboration among the local correlation detectors. Based on a theoretic analysis, the condition for perfect reconstruction of the KLT without the original signal is established. The performance of the proposed scheme is considered for both average and optimal detector and a closed-form expression of the optimal weighting coefficients to minimize the error probability is then derived. Experimental results are provided in the context of color image watermarking, with and without additive Gaussian noise attack.
David S. Taubman
ICIP2
2009 Performance Analysis of Multi-branch Non-regenerative Relay Systems
abstract
The end-to-end performance of multi-branch dual-hop wireless communication systems with non-regenerative relays and equal gain combiner (EGC) at the destination over independent Nakagami-m fading channels is studied. We present new closed form expressions for probability distribution function (PDF) and cumulative distribution function (CDF) of end-to-end signal to noise ratio (SNR) per branch in terms of Meijer's G function. From these results, analytical formulae for the moments of the output SNR, the average overall SNR, the amount of fading and the spectral efficiency are also obtained in closed form. Instead of using moments based approach to analyze the asymptotic error performance of the system, we employ the characteristic function (CHF) method to calculate the average bit error probability (ABEP) and the outage probability for several coherent and non-coherent modulation schemes. The accuracy of the analytical formulae is verified by various numerical results and simulations.
Hieu Q. Huynh, Syed Imtiaz Husain, Jinhong Yuan, Adeel Razi, David S. Taubman
VTC Fall5
2009 Design and Analysis of System on a Chip Encoder for JPEG2000
abstract
Much work has been performed on optimizing the throughput of the block coding system within JPEG2000. However, the question remains as to whether providing parallel simple block coders provides a cheaper method of increasing throughput than complicated optimized block coders. We present the analysis and results for a system on a chip (SoC) software/hardware codesign platform, for parallel coding in JPEG2000 compression standard. We design both a simple and a high performance, optimized peripheral encoder as a hardware accelerator for the JPEG2000 SoC encoding system. The system is implemented on an Altera NIOS II processor with flexible integrated peripheral. We show that there are optimum numbers of parallel block coders and scheduling granularity per row of codeblocks, and that parallel optimized encoders outperform parallel simple encoders. We also demonstrate that the block coding system becomes work starved rather than memory blocked when many parallel coders are present, indicating a discrete wavelet transform bottleneck.
Michael Dyer, Saeid Nooshabadi, David S. Taubman
IEEE Trans. Circuits Syst. Video Technol.3
2009 Perceptual Optimization for Scalable Video Compression Based on Visual Masking Principles
abstract
This paper describes a visual optimization strategy for scalable video compression. The challenge scalable coding presents is that truncation of an embedded codestream may induce variable and highly visible distortion. To overcome the deficiencies of visually lossless coding schemes, we propose using an adaptive masking slope to model the perceptual impact of suprathreshold distortion arising from resolution and bit-rate scaling. This allows important scene structures to be better preserved. Following visual masking principles, local sensitivity to distortion is assessed within each frame. To keep the perceptual response uniform against spatiotemporal errors, we mitigate errors compounded by the motion field during temporal synthesis. Visual sensitivity weights are projected into the subband domain along motion trajectories via a process called perceptual mapping. This uses error propagation paths to capture some of the noise-shaping effects attributed to the motion-compensated transform. A key observation is that low contrast regions in the video are generally more susceptible to unmasking of quantization errors. The proposed approach raises the distortion-length slope associated with these critical regions, altering the bitstream embedding order so that visually sensitive sites may be encoded with higher fidelity. Subjective evaluation demonstrates perceptual improvement with respect to bit-rate, spatial and temporal scalability.
Raymond Leung, David S. Taubman
IEEE Trans. Circuits Syst. Video Technol.2
2008 Optimal delivery of motion JPEG2000 over JPIP with block-wise truncation of quality layers
abstract
This work addresses the transmission of motion JPEG2000 over JPIP, proposing a Variable Bit-Rate (VBR) server-driven policy aimed to minimize the overall distortion of the transmitted frames. In addition, a block-wise strategy of quality layer truncation is introduced to enhance the quality of transmitted frames containing insufficient quality layers. We pose these matters in a general framework where the server is able to deliver video on-the-fly, without needing computationally intensive processing. Experimental results suggest that the proposed VBR policy effectively minimizes the overall distortion, and that the block-wise layer truncation can be jointly combined with the VBR policy to widely enhance the quality of the transmitted frames.
Francesc Aulí Llinàs, David S. Taubman
ICIP2
2008 Active surface modeling at CT resolution limits with micro CT ground truth
abstract
B-spline active surfaces have been extensively utilized to perform automated feature extraction in medical imaging. In this study we focus our attention on the accuracy of 3D B-spline active surface modeling of very small objects. In particular we investigate different smoothing conditions, based upon the regularization of L1 and L2 norms of surface curvature. The model accuracy is directly determined by ground truth through the use of micro CT, allowing the analysis of real data instead of synthetic data. This study highlights the detrimental effect that results from expansive terms in the energy equation such as the L2 norm and the importance of the point spread function in reconstruction problems where the objects are comparable in size to the resolution of the imaging system.
Andrew Philip Bradshaw, Michael J. Todd, David S. Taubman
ICIP3
2008 Improving the resolution scalability of orientation adaptive wavelets
abstract
Oriented wavelets have attracted attention in recent times due to their superior coding performance of images containing diagonal features. We focus on a scheme which effectively aligns the wavelet transform to edges by lifting between shift-interpolated pixels. While compression is improved, poor specification of the orientation can introduce additional aliasing to the lower resolution subband. In this paper, we find that enhanced estimation of the shift field improves both the visual quality of the low resolution subband and compression of the overall image. We also apply an antialiasing transform to the packet decomposition created by the oriented wavelet, which substantially improves the lower resolution at a cost of a reduction in the coding performance.
Jonathan Gan, David S. Taubman
ICIP2
2008 Mixed content image compression by gradient field integration
abstract
Sketch based image coding decomposes an input image into a piece-wise smooth approximation image and a residual image. Image compression by gradient field integration follows this model, but differs by generating the approximation image from gradient data along edge contours and regularly sampled low resolution image data. This allows direct and efficient calculation of an approximation image which is smooth between edges. In this paper, we describe the image compression by gradient field integration approach, together with a low complexity implementation intended for near visually lossless compression of mixed content images. The implementation uses gradient field integration by scanline convolution, simple edge data extraction, DCT residual image coding and is block-based; it is suited to compression of images containing sharp edges occurring along pixel borders. Compression results are provided indicating potential gains from this method.
Peter W. M. Ilbery, David S. Taubman, Andrew P. Bradley
ICIP2
2008 Scalable video compression and spatiotemporal scalability with lifted pyramid and antialiased DWT schemes
abstract
This paper examines the effects of aliasing and investigates the extent to which different transform structures support spatiotemporal scalability. The efficacy of an open-loop spatial pyramid and antialiased DWT schemes are assessed under scalable conditions in terms of their ability to generate highly compressible quality embedded subsets at reduced resolution. The main emphasis is placed on lifting inspired structures with noise suppression or antialiasing properties. As an example, the DWT is augmented with spectral energy exchange lifting steps which disperse aliased content from critical regions of the video sequence at reduced resolution. Finally, we propose alternate ways to characterize a video compression system based on the amount of shift variance and aliasing distortion incurred in half resolution sequences.
Raymond Leung, David S. Taubman
ICIP2
2008 Joint motion and geometry modeling with quad-tree leaf merging
abstract
Quad-tree structures are often used to model motion between frames of a video sequence. However, a fundamental limitation of the quad-tree structure is that it can only capture horizontal and vertical edge discontinuities at dyadically related locations. To address this limitation recent work has focused on the introduction of geometry information to nodes of tree structured motion representations. Recent research into quad-tree structures have also demonstrated the importance of leaf merging. In this paper we create a general quadtree structure, well suited to jointly representing geometry and motion information. We then explore tree pruning and leaf merging, where the geometry information is treated as an equal partner with motion. To achieve an efficient joint representation we improve on the estimation algorithm that detects boundary geometry and introduce polynomial motion models. Experimental results show that the approach taken in this paper provides significant improvement over previous quad-tree based motion representation schemes.
Reji Mathew, David S. Taubman
ICIP2
2008 Rate-distortion optimized delivery of JPEG2000 compressed video with hierarchical motion side information
abstract
Streaming video as a sequence of JPEG2000 images provides the scalability, flexibility, and accessibility at a wide range of bit-rates that is lacking from the current motion-compensated predictive video coding standards; however, streaming this sequence requires considerably more bandwidth. The authors have recently proposed a novel approach that reduces the required bandwidth; this approach uses motion compensation and conditional replenishment of the JPEG2000 code-blocks, aided by server-optimized selection of these code-blocks. This work extends the previous work to the case of hierarchical arrangement of frames, similar to the hierarchical B-frames of the SVC scalable video coding extension of the H.264/AVC standard. We employ a Lagrangian-style rate-distortion optimization procedure to the server transmission problem and compare the performance to that of streaming individual frames and also to that of predictive video coding. The proposed approach can serve a diverse range of client requirements and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. This paper introduces the concepts, formulates the optimization problem, proposes a solution, and compares the performance to alternate strategies.
Aous Thabit Naman, David S. Taubman
ICIP2
2008 JPIP proxy server for remote browsing of JPEG2000 images
abstract
The JPEG2000 image compression standard offers scalability features in support of remote browsing applications. In particular Part 9 of the JPEG2000 standard defines a protocol called JPIP for interactivity with JPEG2000 code-streams and files. In client-server application based on JPIP, a client does not directly interact with the compressed file, but formulates requests using a simple syntax which identifies the current ldquoFocus Windowrdquo. In this kind of application particularly useful could be a proxy server, that potentially can improve the performance of the system through a better use of the network infrastructure. The aim of this work is to propose a proxy server with JPIP capabilities and shows the benefits that can be brought to remote browsing applications.
Livio Lima, David S. Taubman, Riccardo Leonardi
MMSP2
2008 Motion modeling with separate quad-tree structures for geometry and motion
abstract
Quad-tree structures are often used to model motion between frames of a video sequence. However, a fundamental limitation of the quad-tree structure is that it can only capture horizontal and vertical edge discontinuities at dyadically related locations. To address this limitation recent work has focused on the introduction of geometry information to nodes of tree structured motion representations. In this paper we explore modeling boundary geometry and motion with separate quadtree structures. Recent work into quad-tree representations have also highlighted the benefits of leaf merging. We extend the leaf merging paradigm to incorporate separate tree structures for boundary geometry and motion. To achieve an efficient joint representation we introduce polynomial motion models and piecewise linear boundary geometry to our quad-tree structures. Experimental results show that the approach taken in this paper provides significant improvement over previous quad-tree based motion representation schemes.
Reji Mathew, David S. Taubman
MMSP2
2008 Distortion estimation for optimized delivery of JPEG2000 compressed video with motion
abstract
A JPEG2000 compressed video sequence can provide better support for scalability, flexibility, and accessibility at a wider range of bit-rates than the current motion-compensated predictive video coding standards; however, it requires considerably more bandwidth to stream. The authors have recently proposed a novel approach that reduces the required bandwidth; this approach uses motion compensation and conditional replenishment of JPEG2000 code-blocks, aided by server-optimized selection of these code-blocks. The proposed approach can serve a diverse range of client requirements and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. This work extends the previous work by approximating the distortion associated with the decisions made by the server without the need to recreate the actual video sequence at the server. The proposed distortion estimation algorithm is general and can be applied to various frames arrangements. Here, we choose to employ it in a hierarchical arrangement of frames, similar to the hierarchical B-frames of the SVC scalable video coding extension of the H.264/AVC standard. We employ a Lagrangian-style rate-distortion optimization procedure to the server transmission problem and compare the performance of both distortion estimation and exact distortion calculation cases against streaming individual frames and SVC. Results obtained suggest that the distortion estimation algorithm considerably reduces the amount of calculation needed by the server without enormously degrading the performance compared to the exact distortion calculation case. This work introduces the concepts, formulates the estimation and optimization problems, proposes a solution, and compares the performance to alternate strategies.
Aous Thabit Naman, David S. Taubman
MMSP2
2008 Optimal LR-PET protection for scalable video streams over lossy channels with random delay
abstract
This paper investigates the optimal PET protection for streaming scalably compressed streams over networks where the delivery time constraints allow limited retransmissions (LR) and the communication channels exhibit both random losses and delays. A key property must be considered in this scenario is the possibility that a packet successfully arrives at the receiver in time, even if its acknowledgment is not received by the sender at certain deadlines. This paper proposes an extended LRPET scheme, namely random-delay LR-PET, in which additional streams may be sent to provide supplemental protection for the packets whose acknowledgments are still missing at a specified time after the transmission. To determine the optimal protection in each transmission opportunity, hypotheses concerning the number of acknowledged packets and the effect of future retransmission are considered. As the key contribution of this paper, we develop a method to derive the effective overall recovery probability versus redundancy characteristic, which significantly simplifies the actual protection assignment procedure. This paper also demonstrates the benefits of the optimization strategy proposed for this random-delay LR-PET scheme and the cruciality of time selection for scheduling retransmission.
Ruiqin Xiong, David S. Taubman
MMSP2
2008 Efficient Interfacing of DWT and EBCOT in JPEG2000
abstract
Discrete wavelet transform (DWT) and embedded block coder (BC) are two main modules in JPEG2000 compression system. Data transfer between the DWT and BC modules presents challenges due to difference in data format generated by the DWT and data format required by the BC module. In this paper, we investigate data transfer and storage techniques between the DWT and BC modules. We propose an efficient memory organization and data transfer schemes to reduce the data bandwidth. A VLSI architecture for the proposed data transfer system is proposed and synthesized for TSMC 0.18- process. Simulation results show that our proposed techniques result in approximately four times less bandwidth for the BC module while requiring an extra hardware cost of only 11%.
Saeid Nooshabadi, David S. Taubman
IEEE Trans. Circuits Syst. Video Technol.3
2007 Antialiasing Scalable Video with a Modulated Lifting Structure
abstract
A key problem existing in scalable video based on wavelet transforms is the incidence of aliasing at reduced spatial resolutions. The problem exists because of a fundamental limit in the design of wavelet filters. However, typical video frames are observed to have a strong rolloff in the high frequencies before the Nyquist frequency, which we term the "deadzone". In this paper, we propose a transform that not only exploits this property, but replicates the property in successive DWT levels. The transform is created by modifying the lifting implementation of any regular discrete wavelet transform with two additional modulated lifting steps. The two steps act to transfer aliased content from the low pass band into the deadzone of the high pass band. The proposed transform still retains the property of perfect reconstruction. Visual confirmation shows that the transform performs well in antialiasing at half resolution. However, currently the modulated lifting transform exacts a significant coding penalty relative to baseline DWT compression.
Jonathan Gan, David S. Taubman
ICASSP (1)2
2007 Analysis of Multiple Parallel Block Coding in JPEG2000
abstract
We present the analysis and results for a system on a chip (SoC) software/hardware codesign platform, for parallel coding in JPEG2000 compression standard. We show that there are optimum numbers of parallel block coders and scheduling granularity per row of codeblocks. The system was implemented on an Altera NIOS II processor with flexible integrated peripheral.
Michael Dyer, Saeid Nooshabadi, David S. Taubman
ICIP (5)3
2007 Non-Separable Wavelet-Like Lifting Structure for Image and Video Compression with Aliasing Suppression
abstract
Wavelet-based approaches to resolution scalable video are significantly hindered due to the phenomenon of aliasing at reduced spatial resolutions. Improvements to the wavelet filters themselves have limited success, because they are fundamentally constrained by the perfect reconstruction conditions. We propose a packet lifting structure that is capable of perfect reconstruction, yet achieves superior antialiasing performance to improved wavelet designs. While the aliasing removal process incurs some loss of overall compression performance, this is offset by increased compressibility of the lower resolution information. Preliminary results suggest performance competitive with pyramid schemes for achieving similar goals.
Jonathan Gan, David S. Taubman
ICIP (6)2
2007 Motion Modeling with Geometry and Quad-tree Leaf Merging
abstract
Quad-tree structures are often used to model motion between frames of a video sequence. However, a fundamental limitation of the quadtree structure is that it can only capture horizontal and vertical edge discontinuities at dyadically related locations. In this paper we seek to address this limitation by introducing geometry information into the nodes of a pruned quad-tree. We start with a typical optimally pruned quad-tree where each node is allowed to model motion. Then for each node in the tree, we consider augmenting the node's motion model with a linear geometry model. Experimental results show that the introduction of geometry information improves the performance of quad-trees in representing motion. Recent work into quad-tree representations have highlighted the benefits of leaf merging. In this paper we extend the leaf merging paradigm to incorporate both geometry and motion information, allowing the creation of regions that have both motion and geometry attributes.
Reji Mathew, David S. Taubman
ICIP (2)2
2007 A Novel Paradigm for Optimized Scalable Video Transmission Based on JPEG2000 with Motion
abstract
A novel paradigm for optimized scalable video streaming is presented. The paradigm proposes transmission of motion vectors and selected code-blocks of the JPEG 2000 representation of each new frame, instead of frame differences as in existing methods. These code-blocks are selected to achieve the highest MSE. This paradigm overcomes the flexibility and accessibility limitations imposed by predictive motion compensated video by relying on the JPEG 2000 stream features for spatial scalability and on motion compensation and server-optimized conditional replenishment for temporal redundancy reduction. It is expected that real-time and interactive applications, such as teleconferencing and surveillance, would benefit most from this paradigm. This paper introduces the paradigm, formulates an optimization procedure for one simple case where it can be applied and compares its performance with alternate strategies.
Aous Thabit Naman, David S. Taubman
ICIP (5)2
2007 Interpolation Specific Resolution Synthesis
abstract
We introduce a new approach to resolution synthesis which is specifically matched to the interpolation problem. This is achieved by explicitly aligning classification and subsequent interpolation activities with image models conducive to interpolation. Our image models are pre-determined rather than discovered and represent a range of edges of arbitrary orientation, profile and relative position. We demonstrate superior interpolation outcomes compared to statistical classification based resolution synthesis.
Ramez Yoakeim, David S. Taubman
ICIP (4)2
2007 Efficient Data Transfer Techniques and VLSI architecture for DWT-Block Coder Integration of JPEG2000 Encoder
abstract
JPEG2000 is a new image compression standard known for its rich set of features, impressive compression performance as well as its complexity for efficient hardware implementation. Discrete Wavelet Transform (DWT) and embedded Block Coder (BC) are two main modules in JPEG2000 compression system. Data transfer between the DWT and BC modules presents challenges due to difference in data format generated by the DWT and data format accepted by the BC module. This paper investigates data transfer and storage techniques between the DWT and BC modules. The paper proposed an efficient memory organization and a data transfer scheme to reduce the data bandwidth. A VLSI architecture for the proposed data transfer (DT) system is proposed and synthesized for TSMC 0.18μm process. Simulation results show that the proposed techniques result in an aggregate reduction in the bandwidth requirement by a factor of four for the BC module while incurring an extra hardware cost of only 5%.
Saeid Nooshabadi, David S. Taubman
ISCAS3
2006 Near-Optimal Low-Cost Distortion Estimation Technique for JPEG2000 Encoder
abstract
Optimal rate-control is a very important feature of JPEG2000 which allows simple truncation of compressed bit-stream to achieve best image quality at a given target bit-rate. Accurate distortion estimation with respect to allowed bit-stream truncation points, is essential for rate-control performance. In this paper, we address the issues involved in accurate distortion estimation for hardware oriented implementation of JPEG2000 encoding systems with generic block coding capabilities. We propose a novel hardware friendly distortion estimation technique. Rate control based on the proposed technique results in only an average 0.02 dB PSNR degradation with respect to the optimal distortion estimation approach used in the software implementations of JPEG2000. This is the best performance reported in comparison to existing techniques. The proposed technique requires only an additional 4096 bits per block coder which is 80% less than the memory requirements of optimal approach
David S. Taubman, Saeid Nooshabadi
ICASSP (3)2
2006 LR-PET Optimization Strategy for Protection of Scalable Video with Unreliable Acknowledgement
abstract
The priority encoding transmission (PET) framework may be leveraged to exploit both unequal error protection and limited retransmission, for RD optimized delivery of streaming media. The resulting paradigm, known as LR-PET, has been shown to offer significant advantages over a variety of other approaches. This paper extends the existing LR-PET framework to allow for unreliable acknowledgement, providing both an efficient policy for retransmission, as well as an efficient optimization algorithm. The extension is non-obvious, yet retains the complexity benefits of LR-PET. Experimental results confirm that the extended LR-PET procedure also retains the protection benefits of LR-PET with the expected dependence on the acknowledgement probability.
Marco Durigon, David S. Taubman
ICIP2
2006 Minimizing the Perceptual Impact of Visual Distortion in Scalable Wavelet Compressed Video
abstract
This paper considers the efficacy of a class of intra-channel contrast masking models and proposes a simple extension to increase their effectiveness in capturing perceptual diversity in complex stimuli. "Perceptual diversity" encapsulates the idea that our perceptual response to visual distortion varies, depending not only on where quantization errors occur, but also how they coincide with structural elements that constitute a video frame. We offer a fresh perspective on perceptual modeling and identify factors that limit visual optimization performance. We find that the perceptual impact of suprathreshold visual distortion cannot be fully assessed based on contrast alone, hi this work, we introduce a context-adaptive masking slope to complement the basic functions of our perceptual model. This versatile feature allows visual sensitivity to be emphasized, or suppressed, in a manner which reflects the perceptual significance of local distortion in a video. Finally, we propose a visual optimization strategy for embedded video coding based on a perceptual mapping and distortion scaling approach. This technique mitigates the noise shaping effects due to motion-compensated temporal synthesis.
Raymond Leung, David S. Taubman
ICIP2
2006 Hierarchical and Polynomial Motion Modeling with Quad-Tree Leaf Merging
abstract
Recent work into quad-tree representations has commented on the sub-optimal performance of quad-trees due to not exploiting the dependency between neighboring leaf nodes with different parents. Leaf merging therefore has been proposed to rectify this performance loss. In the context of quad-tree motion models, the performance of leaf merging, was recently demonstrated in which an optimally pruned H.264 motion model was taken as the starting point for subsequent merging steps. In this paper we explore the performance of leaf merging over a wider range of conditions by starting with a general tree structure where each node is allowed to represent motion using either polynomial models or a single translational vector. We also explore two cases of motion prediction methods, these being spatial prediction and hierarchical prediction. Experimental results show the benefit of merging across these broad conditions. In comparison to the merged H.264 model reported in the previous work, substantial gains are evident. Furthermore we explore applying the merging principles to branch nodes of a quad-tree to achieve efficient hierarchical motion representation.
Reji Mathew, David S. Taubman
ICIP2
2006 Spatially Continuous Orientation Adaptive Discrete Packet Wavelet Decomposition for Image Compression
abstract
In this paper, we propose an orientation adaptive discrete wavelet transform (DWT) with perfect reconstruction. The proposed transform utilizes the lifting structure to effectively orient the 2D-DWT bases in the direction of local image features. A shifting operator is employed within each lifting step to align spatial geometric features along the vertical or horizontal directions. The proposed oriented transform generates a scalable representation for the image and the orientation information. To approximate the asymptotically optimal rate-distortion performance of a piecewise regular function more closely, we adopt a packet wavelet decomposition. The experimental results obtained by implementing the proposed transform in a JPEG2000 codec illustrate superior compression performance for the oriented transform with more than 2.5 dB improvement for highly oriented natural images. More importantly, even at the same PSNR, the proposed scheme reduces the visual appearance of the Gibbs-like artifacts significantly, considerably improving the visual quality of the reconstructed image.
Nagita Mehrseresht, David S. Taubman
ICIP2
2006 Localized Distortion Estimation from Already Compressed JPEG2000 Images
abstract
We consider the estimation of local residual distortion from an already compressed image, without access to the original. We are motivated by applications in which multiple compressed views of a scene are combined to obtain a single best rendering. However, other applications in image analysis and reconstruction could doubtless benefit from reliable local distortion estimates. We show that prediction of the residual distortion in JPEG2000 code-blocks is tantamount to extrapolating distortion-rate slope information. It also appears that the extrapolation problem can be rendered more reliable by the provision of one additional global parameter for each subband. Results based on our proposed method suggest that estimation errors on the order of 1 to 3 dB can be obtained on average, over a wide range of bit-rates; these results appear to hold for a range of images, block sizes and quality layer densities.
David S. Taubman
ICIP1
2006 Server Policies For Interactive Transmission Of 3D Scenes
abstract
We consider an interactive client-server application for remote browsing of 3D scenes. The information about texture and geometry is available at server side in the form of scalably compressed images and depthmaps, corresponding to a multitude of original image views. Image and depth components are both open to augmentation as more content becomes available. During the interactive browsing experience, the server allocates the available bandwidth between the delivery of new elements from the various original view bit-streams and new elements from the original geometry bit-streams. We propose a rate-distortion criterion to decide the best transmission policy for the server, since the best solution is not always to send the nearest original view image to the one which the client is rendering. We also outline how the JPIP standard for interactive transmission of JPEG2000 images can be exploited for remote exploration of 3D scenes
Pietro Zanuttigh, Nicola Brusco, David S. Taubman, Guido M. Cortelazzo
MMSP3
2006 Interactive representation of still and dynamic scenes
Guido M. Cortelazzo, David S. Taubman
Signal Process. Image Commun.2
2006 A novel framework for the interactive transmission of 3D scenes
Pietro Zanuttigh, Nicola Brusco, David S. Taubman, Guido M. Cortelazzo
Signal Process. Image Commun.3
2006 Realizing Low-Cost High-Throughput General-Purpose Block Encoder for JPEG2000
abstract
The block coder, which is a key module in the JPEG2000 image compression system, presents challenges for realization of a high-throughput, low-hardware-cost VLSI architecture. Though efficient architectures have been proposed for a block coder operating in specific modes, existing generic block coder architectures have low throughput versus hardware cost performance. In this paper, we present a low-cost, high-throughput VLSI architecture for a generic block coder. Concurrent symbol processing (CSP) is used to improve throughput of the block coder's submodules, the bit plane coder (BPC) and arithmetic coder (AC). The proposed BPC processes one stripe-column/clock-cycle during every coding pass and generates up to 10 context-data (CxD) pairs/clock-cycle. The proposed AC processes two CxD/clock-cycles. Throughput is then further increased by using column speedup and novel run-mode skipping techniques at the BPC module. Hardware cost for the proposed block coder is reduced by using an optimal two-subbank BPC memory architecture. Additionally, image statistics are used to choose efficient configuration parameters for the VLSI architecture. The proposed block coder is implemented on Altera stratix FPGA and TSMC ASIC 0.18-mum platforms. Implementation results show that our block coder has average throughputs of 16.23 and 73.42 Msamples/s, respectively, on the FPGA and ASIC platforms. The block-coder test chip has 22515 gates and 2.33 mm2chip area. In comparison with similar existing architectures, it has the highest throughput versus hardware cost performance
Saeid Nooshabadi, David S. Taubman, Michael Dyer
IEEE Trans. Circuits Syst. Video Technol.3
2006 A flexible structure for fully scalable motion-compensated 3-D DWT with emphasis on the impact of spatial scalability
abstract
We investigate the implications of the conventional "t+2-D" motion-compensated (MC) three-dimensional (3-D) discrete wavelet/subband transform structure for spatial scalability and propose a novel flexible structure for fully scalable video compression. In this structure, any number of levels of "pretemporal" spatial wavelet decomposition are performed on the original full resolution frames, followed by MC temporal decomposition of the subbands within each spatial resolution level. Further levels of "posttemporal" spatial decomposition may be performed on the spatiotemporal subbands to provide additional levels of spatial scalability and energy compaction. This structure allows us to trade energy compaction against the potential for artifacts at reduced spatial resolutions. More importantly, the structure permits extensive study of the interaction between spatial aliasing, scalability and energy compaction. We show that where the motion model fails, the "t+2-D" structure inevitably produces misaligned spatial aliasing artifacts in reduced resolution sequences. These artifacts can be removed by using pretemporal spatial decomposition. On the other hand, we also show that the "t+2-D" structure necessarily maximizes compression efficiency. We propose different schemes to minimize the loss of compression efficiency associated with pretemporal spatial decomposition.
Nagita Mehrseresht, David S. Taubman
IEEE Trans. Image Process.2
2006 An efficient content-adaptive motion-compensated 3-D DWT with enhanced spatial and temporal scalability
abstract
We propose a novel, content adaptive method for motion-compensated three-dimensional wavelet transformation (MC 3-D DWT) of video. The proposed method overcomes problems of ghosting and nonaligned aliasing artifacts which can arise in regions of motion model failure, when the video is reconstructed at reduced temporal or spatial resolutions. Previous MC 3-D DWT structures either take the form of MC temporal DWT followed by a spatial transform ("t+2D"), or perform the spatial transform first ("2D + t"), limiting the spatial frequencies which can be jointly compensated in the temporal transform, and hence limiting the compression efficiency. When the motion model fails, the "t + 2D" structure causes nonaligned aliasing artifacts in reduced spatial resolution sequences. Essentially, the proposed transform continuously adapts itself between the "t + 2D" and "2D + t" structures, based on information available within the compressed bit stream. Ghosting artifacts may also appear in reduced frame-rate sequences due to temporal low-pass filtering along invalid motion trajectories. To avoid the ghosting artifacts, we continuously select between different low-pass temporal filters, based on the estimated accuracy of the motion model. Experimental results indicate that the proposed adaptive transform preserves high compression efficiency while substantially improving the quality of reduced spatial and temporal resolution sequences.
Nagita Mehrseresht, David S. Taubman
IEEE Trans. Image Process.2
2005 Efficient scheduling schemes for real-time traffic in wireless networks
abstract
In this paper, we study the problem of scheduling real-time traffic in wireless TDMA channels. Specifically, we propose scheduling schemes that dramatically improve the delay performance of individual flows, while at the same time provide fairness and quality of service (QoS) guarantees. Previous works in this area have overlooked one or more of these aspects, so our study is both novel and valuable. Our first scheduling scheme assigns equal delay to all flows. However, it turns out that insisting on equal delays can adversely affect network utilization. We propose a second, related scheduling scheme which improves the network utilization by exploiting the time-varying nature of wireless channel conditions, while at the same time minimizing delay violations. We observe that the impact of these two schemes on delay and network utilization depend strongly on the channel characteristics. Interestingly, our proposed equal-delay scheme appears to perform well over a wide range of channel characteristics, with the added advantages of its simplicity for implementation and admitting an analyzable admission control
Emily Lee, David S. Taubman
GLOBECOM2
2005 On the benefits of leaf merging in quad-tree motion models
abstract
In the field of image coding, it has recently been shown that quad-tree schemes for representing geometric image features are unable to obtain the optimal exponentially decaying rate-distortion behaviour - a problem which can be rectified by following R-D optimal tree pruning with a leaf merging step. Inspired by such results, we note that quad-tree based motion representations for video compression suffer from the same fundamental shortcomings, which can again be overcome by leaf merging. Based on these observations, we describe a non-iterative low complexity extension to the existing motion model used by H.264, demonstrating significant improvements in overall compression performance. At low to moderate bit-rates, our experimental implementation yields reductions of 10% or more in overall bit-rate, relative to H.264. Perhaps more importantly, this paper reinforces the importance of leaf merging as a general tool for quad-tree based modeling schemes.
Raffaele De Forni, David S. Taubman
ICIP (2)2
2005 Filter based MAP estimation of images with integrated segmentation
abstract
We present a computationally efficient technique for MAP estimation of images in the presence of both blur and noise. The method uses a piecewise stationary Gaussian prior with the segmentation incorporated in a natural way. To generate the solution we apply the Wiener filter to the data after first subtracting out the influence of the surrounding regions. For any particular segmentation, the method gives rise to a linear system which can be solved using the successive over relaxation (SOR) iterative method. We incorporate the segmentation as an extra, nonlinear step, at each point in the SOR method. The resulting combined method gives a good solution after just a few iterations. The proposed method has wide applicability to inverse imaging problems, and examples are provided showing application to the demosaicking problem.
David Humphrey, David S. Taubman
ICIP (1)2
2005 Perceptual mappings for visual quality enhancement in scalable video compression
abstract
This paper presents a new framework for achieving superior visual quality in scalable video compression. In contrast with adaptive quantization strategies, our approach compensates for motion modeling deficiencies and alleviates temporal artifacts without needing to modify the decoder. The proposed technique works on the principle of post-compression distortion scaling and consists of two key steps. Perceptual analysis identifies regions in the frame, which are visually sensitive. Spatial mappings then incorporate the non-linear effects of the motion field, projecting the sensitivity information back into the subband domain along motion trajectories. The derived sensitivity maps are used to affect the codestream embedding order when quality layers are formed. The overall objective is to raise the distortion-length slope associated with low contrast regions, which are particularly susceptible to the unmasking of artifacts, so that they will be encoded with higher fidelity. This technique substantially improves the visual quality of structural elements that otherwise suffer significant degradation.
Raymond Leung, David S. Taubman
ICIP (2)2
2005 Impact of motion on the random access efficiency of scalable compressed video
abstract
To achieve high access efficiency in scalable video compression, the embedded codestream should facilitate efficient viewing through a window with compact support in space and time. The ability to interact with a region of interest in any spatio-temporal subband is limited by the granularity of the code blocks chosen during compression, which is linked to coding performance. Previous studies have investigated the optimal code block dimensions with respect to coding efficiency and random accessibility without considering the impact of motion. This paper further examines how the random access cost is exacerbated by motion. The loss in performance due to motion-compensation is considered from a random access perspective, in the context of a non-separable spatio-temporal subband transform. The methods are general and the results are pertinent to a variety of interactive video browsing applications.
Raymond Leung, David S. Taubman
ICIP (3)2
2005 Greedy non-linear approximation of the plenoptic function for interactive transmission of 3D scenes
abstract
We consider an interactive browsing environment, with greedy optimization of a current view, conditioned on the availability of previously transmitted information for other (possibly nearby) views, and subject to a transmission budget constraint. Texture information is available at a server in the form of scalably compressed images, corresponding to a multitude of original image views. Surface geometry is also represented at the server in a scalable fashion. At any point in the interactive browsing experience, the server must decide how to allocate transmission resources between the delivery of new elements from the various original view bit-streams and new elements from the geometry bit-stream. The proposed framework may be interpreted as a greedy strategy for non-linear approximation of the plenoptic function, since it considers both view sampling and rate-distortion criteria. We particularly elaborate upon a novel geometry- and distortion-sensitive strategy for blending the information available from different views at the client.
Pietro Zanuttigh, Nicola Brusco, David S. Taubman, Guido M. Cortelazzo
ICIP (1)3
2005 Improvements to the Intra-Coding Modes Offered by H.264
abstract
In this paper we investigate how to improve the overall prediction efficiency of intra frame coding in H.264, incorporating into the rate-distortion optimization framework several new prediction algorithms that can both fix some weaknesses of the existing 16times16 standard mode and partially fill the performance gap between available 4times4 and 16times16 strategies. Results show a good gain in performance respect to plain H.264
Luca Merello, David S. Taubman
MMSP2
2005 Optimal erasure protection for scalably compressed video streams with limited retransmission on channels with IID and bursty loss characteristics
Johnson Thie, David S. Taubman
Signal Process. Image Commun.2
2005 Transform and Embedded Coding Techniques for Maximum Efficiency and Random Accessibility in 3-D Scalable Compression
abstract
This study investigates random accessibility and efficiency enhancements in highly scalable video and volumetric compression. With the advent of interactive multimedia technology, random accessibility has emerged as an increasingly important consideration in the design and optimization process. In this paper, we assess the impact that the transform, embedded coding components, and code-block configurations have on the compression efficiency and accessibility of a scalable codestream. We develop performance bounds on techniques which exploit temporal redundancy within the confines of a feed-forward compression system. We also examine their random access properties to argue the significance of motion-adaptive subband transforms. When information-theoretic measures are used to determine the potential benefits of three-dimensional (3-D) context coding, we find that most of the coding gain is attributed to code-block extension, rather than interslice context modeling itself. To gain further insight into the tradeoffs that the coding part has to offer, we run a series of simulations to determine code-block partitioning strategies which maximize reconstruction quality and space-time localization. The LIMAT framework and EBCOT coding paradigm have laid a solid foundation for further progress in the development of highly scalable 3-D compression systems.
Raymond Leung, David S. Taubman
IEEE Trans. Image Process.2
2005 Optimal erasure protection for scalably compressed video streams with limited retransmission
abstract
This paper shows how the priority encoding transmission (PET) framework may be leveraged to exploit both unequal error protection and limited retransmission for RD-optimized delivery of streaming media. Previous work on scalable media protection with PET has largely ignored the possibility of retransmission. Conversely, the PET framework has not been harnessed by the substantial body of previous work on RD optimized hybrid forward error correction/automatic repeat request schemes. We limit our attention to sources which can be modeled as independently compressed frames (e.g., video frames), where each element in the scalable representation of each frame can be transmitted in one or both of two transmission slots. An optimization algorithm determines the level of protection which should be assigned to each element in each slot, subject to transmission bandwidth constraints. To balance the protection assigned to elements which are being transmitted for the first time with those which are being retransmitted, the proposed algorithm formulates a collection of hypotheses concerning its own behavior in future transmission slots. We show how the PET framework allows for a decoupled optimization algorithm with only modest complexity. Experimental results obtained with Motion JPEG2000 compressed video demonstrate that substantial performance benefits can be obtained using the proposed framework.
David S. Taubman, Johnson Thie
IEEE Trans. Image Process.1
2005 Optimal erasure protection strategy for scalably compressed data with tree-structured dependencies
abstract
This paper is concerned with the transmission of scalably compressed data sources over lossy channels. Specifically, this paper is concerned with packet networks or, more generally, erasure channels. Previous work has generally assumed that the source elements form linear dependencies. The contribution of this paper is an unequal erasure protection algorithm which is able to take advantage of scalable data with more general dependency structures. In particular, the proposed scheme is adapted to data with tree-structured dependencies. The source elements are allocated to clusters of packets according to their dependency structure, subject to constraints on packet size and channel code-word length. Given a packet cluster arrangement, source elements are assigned optimal channel codes subject to a constraint on the total transmission length. Experimental results confirm the benefit associated with exploiting the actual dependency structure of the data.
Johnson Thie, David S. Taubman
IEEE Trans. Image Process.2
2004 Improved throughput arithmetic coder for jpeg2000
Michael Dyer, David S. Taubman, Saeid Nooshabadi
ICIP2
2004 A filtering approach to edge preserving map estimation of images
abstract
We present a computationally efficient technique for MAP estimation of images in the presence of both blur and noise. The image is divided into statistically independent regions. Each region is modelled with a WSS Gaussian prior. Classical Wiener filter theory is used to generate a set of convex sets in the solution space, with the solution to the MAP estimation problem lying at the intersection of these sets. The proposed algorithm uses an underlying segmentation of the image, and a means of determining the segmentation and refining it are described. The algorithm is suitable for demosaicking of digital camera images and other image restoration problems.
David Humphrey, David S. Taubman
ICIP2
2004 Spatial scalability and compression efficiency within a flexible motion compensated 3D-DWT
abstract
We investigate the implications of the conventional "t+2D" MC 3D-DWT structure for spatial scalability and propose a more flexible "2D+t+2D" structure. An initial P levels of spatial wavelet decomposition are followed by T levels of motion compensated temporal decomposition, applied separately to each spatial resolution level. A further S-P levels of spatial decomposition are applied to the resulting subbands. By adjusting P, the structure allows us to trade energy compaction with the potential for artifacts at reduced spatial resolutions. This allows us to study the interaction between scalability and compression efficiency. We show that the "t+2D" structure (P=0) necessarily maximizes compression efficiency, while allowing misaligned spatial aliasing artifacts to arise at reduced resolutions. These artifacts can be removed by increasing the value of P, at an inevitable cost in compression efficiency. We show how this cost can be minimized.
Nagita Mehrseresht, David S. Taubman
ICIP2
2004 An efficient content-adaptive MC 3D-DWT with enhanced spatial and temporal scalability
abstract
In this paper we propose a novel, adaptive method for motion compensated 3D wavelet transformation (MC 3D-DWT) of video. The proposed method overcomes problems of ghosting and nonaligned aliasing artifacts which can arise in regions of motion model failure, when the video is reconstructed at reduced temporal or spatial resolutions. Previous MC 3D-DWT structures either take the form of a MC temporal DWT followed by a spatial transform ("t+2D"), or perform the spatial transform first, limiting the spatial frequencies which can be jointly compensated in the temporal transform and hence limiting the compression efficiency. Essentially, the proposed transform continuously adapts itself between these two extremes, based on information available within the compressed bit-stream. Experimental results indicate that the proposed adaptive transform significantly reduces the cost in compression efficiency required to achieve high quality spatial and temporal scalability.
Nagita Mehrseresht, David S. Taubman
ICIP2
2004 Optimal erasure protection for scalably compressed video streams with limited retransmission on channels with HD and bursty loss ciiaracteiustics
Johnson Thie, David S. Taubman
ICIP2
2004 Quantitative analysis of resolution synthesis
Ramez Yoakeim, David S. Taubman
ICIP2
2004 Highly scalable video compression with scalable motion coding
abstract
A scalable video coder cannot be equally efficient over a wide range of bit rates unless both the video data and the motion information are scalable. We propose a wavelet-based, highly scalable video compression scheme with rate-scalable motion coding. The proposed method involves the construction of quality layers for the coded sample data and a separate set of quality layers for the coded motion parameters. When the motion layers are truncated, the decoder receives a quantized version of the motion parameters used to code the sample data. The effect of motion parameter quantization on the reconstructed video distortion is described by a linear model. The optimal tradeoff between the motion and subband bit rates is determined after compression. We propose two methods to determine the optimal tradeoff, one of which explicitly utilizes the linear model. This method performs comparably to a brute force search method, reinforcing the validity of the linear model itself. Experimental results indicate that the cost of scalability is small. In addition, considerable performance improvements are observed at low bit rates, relative to lossless coding of the motion information.
Andrew J. Secker, David S. Taubman
IEEE Trans. Image Process.2
2003 Context modeling and accessibility for 3D scalable compression
abstract
Highly scalable video compression based on invertible motion adaptive lifting transforms has emerged as a promising area in image processing research and an important component in interactive multimedia technology. However, within this feed-forward framework, the potential for coding efficiency improvement and its impact on random accessibility still has not been carefully assessed. In this paper, we compare the merits of several three-dimensional context coding strategies from an information-theoretic perspective. The variation in random access cost in response to coding parameter adjustments is analyzed, for a variety of spatial and temporal configurations.
Raymond Leung, David S. Taubman
ICIP (2)2
2003 Adaptively weighted update steps in motion compensated lifting based scalable video compression
abstract
Motion-compensated temporal wavelet decomposition is a use-framework for fully scalable video compression schemes. In paper we propose a new approach to reduce the ghosting artifacts in low-pass temporal subbands; we adaptively weight the update steps according to the energy in the high-pass temporal sub-bands at the corresponding location. Experimental results show that the proposed algorithm can substantially remove ghosting from low-pass temporal frames. Importantly, at full frame-rate, the proposed algorithm has similar performance to the original motion-compensated temporal decomposition, with superior performance where the motion model fails significantly. While entirely skipping the update steps accomplishes a similar objective, we show that the proposed method for adaptively weighting the update steps has better performance, especially in the presence of additive noise. Since the compressed bit-stream is scalable, the decoder does not generally have exactly the same information which the encoder used to determine weights for the update steps. Nevertheless, the proposed method exhibits good robustness to quantization error.
Nagita Mehrseresht, David S. Taubman
ICIP (2)2
2003 Highly scalable video compression with scalable motion coding
abstract
For a scalable video coder to remain efficient over a wide range of bit-rates, some form of scalability must exist in the motion information. We propose a wavelet-based highly scalable video coder with scalable motion coding. Our proposed method involves the construction of quality layers for the coded wavelet sample data and a separate set of quality layers for scalably coded motion parameters. When the motion layers are truncated, the decoder receives a quantized version of the motion parameters used to generate the wavelet sample data. A linear model is used to infer the impact of motion quantization on reconstructed video distortion. An optimal tradeoff between the motion and subband bit-rates may then be found. Experimental results indicate that the cost of scalability is small. At low bit-rates, significant improvements are observed relative to lossless coding of the motion information.
Andrew Secker, David S. Taubman
ICIP (3)2
2003 Rate-distortion optimized interactive browsing of JPEG2000 images
abstract
This paper is concerned with remote browsing of JPEG2000 compressed imagery. A potentially interactive client identifies a region and maximum resolution of interest. The server responds by sending information from the JPEG2000 code-stream which is relevant to this region and resolution. We propose an algorithm for sequencing data from the original code-stream in such a way as to maximize the received image quality within the region of interest, at each point in the transmission. The proposed R-D optimal sequencing algorithm is demonstrated in the context of two quite different client-server paradigms, one of which is consistent with the evolving JPIP (JPEG2000 Internet protocols) standard. Performance improvements as large as 8 dB are achieved with respect to a layer progressive sequencing strategy.
David S. Taubman, René Rosenbaum
ICIP (3)1
2003 Optimal erasure protection assignment for scalable compressed images with tree-structured dependencies
abstract
This paper is concerned with the efficient transmission of scalable compressed images with complex dependency structures over lossy communication channels. Our recent work proposed a strategy for allocating source elements into clusters of packets and finding their optimal code rates. However, the previous work assumes that source elements form a simple chain of dependencies. The present paper proposes a modification to the earlier strategy to exploit the properties of scalable sources which have tree-structured dependency. The source elements are allocated to clusters of packets according to their dependency structure, subject to constraints on packet size and channel codeword length. Given a packet cluster arrangement, the proposed strategy assigns optimal code rates to the source elements, subject to a constraint on transmission length. Our experimental results suggest that the proposed strategy can outperform the earlier strategy by exploiting the dependency structure.
Johnson Thie, David S. Taubman
ICIP (3)2
2003 Image compression practices and standards for geospatial information systems
abstract
Compression technology is becoming increasingly important in geospatial information systems. In this paper we address some of the most relevant compression issues for remote sensing applications, and highlight the potential benefits of the JPEG set of standards. In particular, we review the JPEG, JPEG 2000, and JPEG-LS compression standards, and the JPIP protocol for interactive image retrieval. Finally, we discuss the use of compressed-domain processing, along with the use of flexible file formats for efficient storage and access to metadata.
Enrico Magli, David S. Taubman
IGARSS2
2003 Successive refinement of video: fundamental issues, past efforts, and new directions
David S. Taubman
VCIP1
2003 Architecture, philosophy, and performance of JPIP: internet protocol standard for JPEG2000
David S. Taubman, Robert Prandolini
VCIP1
2003 Lifting-based invertible motion adaptive transform (LIMAT) framework for highly scalable video compression
abstract
We propose a new framework for highly scalable video compression, using a lifting-based invertible motion adaptive transform (LIMAT). We use motion-compensated lifting steps to implement the temporal wavelet transform, which preserves invertibility, regardless of the motion model. By contrast, the invertibility requirement has restricted previous approaches to either block-based or global motion compensation. We show that the proposed framework effectively applies the temporal wavelet transform along a set of motion trajectories. An implementation demonstrates high coding gain from a finely embedded, scalable compressed bit-stream. Results also demonstrate the effectiveness of temporal wavelet kernels other than the simple Haar, and the benefits of complex motion modeling, using a deformable triangular mesh. These advances are either incompatible or difficult to achieve with previously proposed strategies for scalable video compression. Video sequences reconstructed at reduced frame-rates, from subsets of the compressed bit-stream, demonstrate the visually pleasing properties expected from low-pass filtering along the motion trajectories. The paper also describes a compact representation for the motion parameters, having motion overhead comparable to that of motion-compensated predictive coders. Our experimental results compare favorably to others reported in the literature, however, our principal objective is to motivate a new framework for highly scalable video compression.
Andrew J. Secker, David S. Taubman
IEEE Trans. Image Process.2
2002 Unequal protection of JPEG2000 code-streams in wireless channels
abstract
The rapid growth of wireless communication has resulted in a demand for robust transmission of compressed images over wireless channels. The challenge of robust transmission is to protect the compressed image data against loss in such a way as to maximize the received image quality. The paper addresses this problem; investigating unequal error protection of JPEG2000 compressed imagery. More particularly, the results reported in this paper provide guidance concerning the selection of JPEG2000 coding parameters and appropriate combinations of RS (Reed-Solomon) codes, for typical wireless bit error rates in the range 10/sup -4/ to 10/sup -3/.
Ambarish Natu, David S. Taubman
GLOBECOM2
2002 Scalable compression of volumetric images
abstract
An important problem in medical imaging is that of efficient volumetric image compression. In addition to compression efficiency, scalable representations which allow access to the data at various qualities up to and including lossless reproduction are particularly important. Also important is the ability to access local regions of the volumetric data set or to alter the fidelity or resolution with which a selected region is accessed. In this paper, we first consider bounds to the efficacy of three dimensional transform techniques. We then propose a particular technique, based on motion adaptive wavelet lifting, which is able to realize the goals of high compression efficiency, lossy to lossless scalability, and efficient random access, all within a single compressed data stream.
Andrew Secker, Raymond Leung, David S. Taubman
ICIP (2)3
2002 Remote browsing of JPEG2000 images
abstract
JPEG2000 is a highly scalable compression standard, allowing access to image representations with a reduced resolution, a reduced quality or confined to a spatial region of interest. As such, it is well placed to play an important role in interactive imaging applications. However, the standard itself stops short of providing guidance or specific mechanisms for exploiting its scalability in such applications. We describe the JPIK (JPEG2000 interactive with Kakadu) protocol for interactive imaging with JPEG2000. Our results suggest, somewhat surprisingly, that image tiling (dividing into independently compressed sub-images), can be detrimental to effective browsing of large compressed images over low bandwidth connections.
David S. Taubman
ICIP (1)1
2002 Highly scalable video compression using a lifting-based 3D wavelet transform with deformable mesh motion compensation
abstract
This paper continues the development of a new framework for the construction of motion-compensated wavelet transforms for highly scalable video compression. The current authors recently proposed a motion adaptive wavelet transform based on motion-compensated lifting steps. This approach overcomes several limitations of existing methods. In particular, frame warping and block displacement methods cannot efficiently exploit complex motion without sacrificing invertibility. By contrast, the motion-compensated lifting transform remains invertible regardless of the motion model. The previous work was primarily in the context of a block motion model. However, block motion models inevitably yield discontinuous motion fields, which poorly represent complex motion in real video sequences. In this paper we consider the benefits of a continuous motion field, by incorporating a deformable mesh motion model into the existing framework. Experimental results show that this leads to improved compression performance. In addition, we show that the invertibility of continuous motion fields allows greater potential for compactly representing the motion information.
Andrew J. Secker, David S. Taubman
ICIP (3)2
2002 Optimal protection assignment for scalable compressed images
abstract
This paper is concerned with the efficient transmission of scalable compressed images over lossy communication channels. Recent works have proposed several strategies for assigning optimal code-rates to elements in a scalable data stream, under the assumption that all elements are encoded onto a common group of network packets. When the size of the data to be encoded becomes large in comparison with the size of the network packets, such schemes require very long channel codes with high computational complexity. In networks with high loss, small packets are generally more desirable than long packets. This paper proposes a strategy for assigning optimal code-rates to elements of the scalable compressed data source, subject to constraints on the packet size, transmission length and code complexity. Our experimental results suggest that the proposed scheme can outperform previously proposed code-rate assignment policies, subject to the above-mentioned constraints, particularly with high loss.
Johnson Thie, David S. Taubman
ICIP (3)2
2002 JPEG2000: standard for interactive imaging
abstract
JPEG2000 is the latest image compression standard to emerge from the Joint Photographic Experts Group (JPEG) working under the auspices of the International Standards Organization. Although the new standard does offer superior compression performance to JPEG, JPEG2000 provides a whole new way of interacting with compressed imagery in a scalable and interoperable fashion. This paper provides a tutorial-style review of the new standard, explaining the technology on which it is based and drawing comparisons with JPEG and other compression standards. The paper also describes new work, exploiting the capabilities of JPEG2000 in client-server systems for efficient interactive browsing of images over the Internet.
David S. Taubman, Michael W. Marcellin
Proc. IEEE1
2002 Embedded block coding in JPEG 2000
David S. Taubman, Erik Ordentlich, Marcelo J. Weinberger, Gadiel Seroussi
Signal Process. Image Commun.1
2001 Motion-compensated highly scalable video compression using an adaptive 3D wavelet transform based on lifting
abstract
This paper proposes a new framework for the construction of motion compensated wavelet transforms, with application to efficient highly scalable video compression. Motion compensated transform techniques, as distinct from motion compensated predictive coding, represent a key tool in the development of highly scalable video compression algorithms. The proposed framework overcomes a variety of limitations exhibited by existing approaches. This new method overcomes the failure of frame warping techniques to preserve perfect reconstruction when tracking complex scene motion. It also overcomes some of the limitations of block displacement methods. Specifically, the lifting framework allows the transform to exploit inter-frame redundancy without any dependence on the model selected for estimating and representing motion. A preliminary implementation of the proposed approach was tested in the context of a scalable video compression system, yielding PSNR performance competitive with other results reported in the literature.
Andrew J. Secker, David S. Taubman
ICIP (2)2
2000 Generalized Wiener Reconstruction of Images from Colour Sensor Data Using a Scale Invariant Prior
abstract
An algorithm is described for reconstructing images from colour sensor samples, which need not be aligned nor conform to a rectangular sampling geometry. The algorithm has applications in de-mosaicing digital camera color filter array (CFA) data, and processing other imaging modalities such as scanned images and captured video. A unique scale invariant WSS prior model is described for the uncorrupted surface spectral reflectance functions and used to form linear least mean squared error (LLMSE) optimal reconstructions with constrained support operators. Some important results are established concerning the existence and tractability of the solutions based on this prior.
David S. Taubman
ICIP1
2000 High performance scalable image compression with EBCOT
abstract
A new image compression algorithm is proposed, based on independent embedded block coding with optimized truncation of the embedded bit-streams (EBCOT). The algorithm exhibits state-of-the-art compression performance while producing a bit-stream with a rich set of features, including resolution and SNR scalability together with a "random access" property. The algorithm has modest complexity and is suitable for applications involving remote browsing of large compressed images. The algorithm lends itself to explicit optimization with respect to MSE as well as more realistic psychovisual metrics, capable of modeling the spatially varying visual masking phenomenon.
David S. Taubman
IEEE Trans. Image Process.1
1999 Memory Efficient Scalable Line-based Image Coding
abstract
We study the problem of memory-efficient scalable image compression and investigate some tradeoffs in the complexity versus coding efficiency space. The focus is on a low-complexity algorithm centered around the use of sub-bit-planes, scan-causal modeling, and a simplified arithmetic coder. This algorithm approaches the lowest possible memory usage for scalable wavelet-based image compression and demonstrates that the generation of a scalable bit-stream is not incompatible with a low-memory architecture.
Erik Ordentlich, David S. Taubman, Marcelo J. Weinberger, Gadiel Seroussi, Michael W. Marcellin
Data Compression Conference2
1999 High Performance Scalable Image Compression with Ebcot
abstract
A new image compression algorithm is proposed, based on independent embedded block coding with optimized truncation of the embedded bit-streams (EBCOT). The algorithm exhibits state-of-the-art compression performance while producing a bit-stream with a rich feature set, including resolution and SNR scalability together with a random access property. The algorithm has modest complexity and is extremely well suited to applications involving remote browsing of large compressed images. The algorithm lends itself to explicit optimization with respect to MSE as well as more realistic psychovisual metrics, capable of modeling the spatially varying visual masking phenomenon.
David S. Taubman
ICIP (3)1
1999 Adaptive, Non-Separable Lifting Transforms for Image Compression
abstract
In the context of high performance image compression algorithms, such as that emerging as the JPEG 2000 standard, the wavelet transform has demonstrated excellent compression performance with natural images. Like all waveform coding techniques, however, performance suffers in the neighbourhood of oriented edges clad with artificial imagery such as text and graphics. In this paper, we explore some of the opportunities offered by the framework of lifting for developing adaptive wavelet transforms to improve performance under these conditions.
David S. Taubman
ICIP (3)1
1996 A common framework for rate and distortion based scaling of highly scalable compressed video
abstract
Scalability refers to the ability to modify the resolution and/or bit rate associated with an already compressed data source in order to satisfy requirements which could not be foreseen at the time of compression. A number of researchers have already demonstrated the feasibility of efficient scalable image and video compression. The principle focus of this paper is to describe data structures for highly scalable compressed video, which are able to support simple, generic scaling approaches for both constant bit rate and constant distortion scaling criteria. Interactive video material presents particular challenges when the data stream is to be scaled to maintain an approximately constant level of distortion, rather than just a constant bit rate. Special attention is paid, therefore, to the development of generic, robust scaling algorithms for such applications. The data structures and scaling methodologies developed are particularly appealing for the distribution of highly scalable compressed video over heterogeneous media, because they simultaneously support both variable bit rate (VBR) and constant bit rate (CBR) services with a wide range of available service qualities, using only simple, generic mechanisms for scaling. The performance of the proposed scaling methodologies is experimentally investigated using a highly scalable video compression algorithm, which is able to achieve comparable compression performance to that of the inherently nonscalable MPEG-1 compression standard.
David S. Taubman, Avideh Zakhor
IEEE Trans. Circuits Syst. Video Technol.1
1994 Rate and resolution scalable subband coding of video
abstract
We propose a full colour video compression strategy, based on 3-D subband coding with camera pan compensation, to generate a single, embedded, compressed bit stream supporting multiple decoder display formats and a wide, finely gradated range of bit rates. An experimental implementation of our algorithm produces a single bit stream, from which suitable subsets are extracted to be compatible with many useful decoder frame sizes and frame rates and to satisfy transmission bandwidth constraints ranging from several tens of kilo-bits per second to several mega-bits per second. The reconstructed video quality from any of these bit stream subsets is often found to exceed that obtained from an MPEG-1 implementation, operated with equivalent bit rate constraints, in both perceptual quality and mean squared error.>
David S. Taubman, Avideh Zakhor
ICASSP (5)1
1994 Highly Scalable, Low-Delay Video Compression
abstract
We propose a class of scalable video compression algorithms, within which compression performance may be exchanged for end-to-end delay. Each of the algorithms in this class produces a highly scalable bit stream, from which subsets may be extracted for compatibility with a wide range of display frame sizes, frame rates and bit rate constraints. Moreover, with modest end-to-end delay, the reconstructed video quality associated with any of these subsets is often superior to that obtained using an implementation of the inherently non-scalable MPEG-1 compression standard, operated with equivalent resolution and bit rate constraints. We describe scalable compressed data structures based on a layered substream abstraction, with simple, generic scaling operations, for both constant bit rate and constant distortion scaling criteria.>
David S. Taubman, Avideh Zakhor
ICIP (1)1
1994 Orientation adaptive subband coding of images
abstract
In the subband coding of images, directionality of image features has thus far been exploited very little. The proposed subband coding scheme utilizes orientation of local image features to avoid the highly objectionable Gibbs-like phenomena observed at reconstructed image edges with conventional subband schemes at low bit rates, At comparable bit rates, the subjective image quality obtained by our orientation adaptive scheme is considerably enhanced over a conventional separable subband coding scheme, as well as other separable approaches such as the JPEG compression standard.
David S. Taubman, Avideh Zakhor
IEEE Trans. Image Process.1
1994 Multirate 3-D subband coding of video
abstract
We propose a full color video compression strategy, based on 3-D subband coding with camera pan compensation, to generate a single embedded bit stream supporting multiple decoder display formats and a wide, finely gradated range of bit rates. An experimental implementation of our algorithm produces a single bit stream, from which suitable subsets are extracted to be compatible with many decoder frame sizes and frame rates and to satisfy transmission bandwidth constraints ranging from several tens of kilobits per second to several megabits per second. Reconstructed video quality from any of these bit stream subsets is often found to exceed that obtained from an MPEG-1 implementation, operated with equivalent bit rate constraints, in both perceptual quality and mean squared error. In addition, when restricted to 2-D, the algorithm produces some of the best results available in still image compression.
David S. Taubman, Avideh Zakhor
IEEE Trans. Image Process.1
1993 Orientation adaptive subband coding of images
David S. Taubman, Avideh Zakhor
ISCAS1
1992 A multi-start algorithm for signal adaptive subband systems (image coding)
abstract
A method for optimizing the filter coefficients of signal adaptive subband systems, in which the polyphase transfer matrix is paraunitary, using a mean-squared-error criterion is proposed. An efficient algorithm was developed for locating numerous local optima in the coefficient space, permitting a degree of confidence in the location of globally optimal or near-optimal solutions. Both separable systems and a particular class of nonseparable filter systems are studied. The application of the algorithm to a number of images is described.>
David S. Taubman, Avideh Zakhor
ICASSP1