VLDB 2026 Research / reviewers in the wild / expert
Reji Mathew
dblp:74/1555 · also Reji K. Mathew
· DBLP profile ↗
45ranked-venue papers
19as first author
5since 2021 · last 2023
0000-0003-2940-7325ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 19 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | JPEG Pleno Light Field Encoder with Mesh based View WarpingabstractWe introduce mesh-based view warping to the JPEG Pleno light field coding framework and replace the standardized sample-based forward warping and splatting of reference texture with mesh-based backward warping, which allows for a more disciplined interpolation of the reference texture for predicting the target view. Instead of coding depth maps with JPEG 2000, which is the default option of the JPEG Pleno framework, we employ a recent extension referred to as JPEG 2000 Part 17. This extension utilises breakpoints to describe discontinuity boundary geometry for the purpose of modifying the predict and update lifting steps in the vicinity of detected discontinuities. We directly decode breakpoints and corresponding DWT coefficients onto a mesh and describe a scheme to construct a single, consolidated mesh for a large group of views, borrowing information from multiple coded depth maps. Results show that the cumulative impact of all these modifications enable improved rate-distortion performance in comparison with the default operation of the JPEG Pleno encoder. Yue Li 0034, Reji Mathew, David S. Taubman |
ICIP | 2 |
| 2023 | JPEG 2000 Extensions for Scalable Coding of Discontinuous MediaabstractIn this paper we propose novel extensions to JPEG 2000 for the coding of discontinuous media which includes piecewise smooth imagery such as depth maps and optical flows. These extensions use breakpoints to model discontinuity boundary geometry and apply a breakpoint dependent Discrete Wavelet Transform (BP-DWT) to the input imagery. The highly scalable and accessible coding features provided by the JPEG 2000 compression framework are preserved by our proposed extensions, with the breakpoint and transform components encoded as independent bit streams that can be progressively decoded. Comparative rate-distortion results are provided along with corresponding visual examples which highlight the advantages of using breakpoint representations with accompanying BD-DWT and embedded bit-plane coding. Recently our proposed extensions have been adopted and are in the process of being published as a new Part 17 to the JPEG 2000 family of coding standards. Reji Mathew, Aous Thabit Naman, Yue Li 0034, David S. Taubman |
IEEE Trans. Image Process. | 1 |
| 2022 | Breakpoint Dependent Scalable Coding of Optical Flow VolumeabstractMotion representations that describe motion flow from one single anchor frame to multiple reference frames can be useful for many tasks such as motion compensation (MC) and frame-rate up-sampling. We construct an anchored, multi-flow representation which we refer to as an Optical Flow Volume (OFV). We explore the use of recent breakpoint dependent DWT (BD-DWT) being considered as part of extensions to JPEG 2000 for coding discontinuous media. Breakpoints describe discontinuity boundary geometry and we estimate a single set of breakpoints that can be shared by all individual flows of the OFV. Additionally, as flows are anchored at a common frame, we are able to readily explore inter-flow transforms. Rate scalable results show significant rate-distortion gains for BD-DWT over the 5/3 DWT commonly used with JPEG 2000. The validity and utility of OFV coding are confirmed with accompanying MC results. Reji Mathew, David S. Taubman |
ICIP | 1 |
| 2022 | JPEG Pleno Light Field Encoder with Breakpoint Dependent Affine Wavelet Transform for Disparity MapsabstractThe JPEG Pleno light field encoder can perform disparity compensated view prediction for coding HDCA views. For operating in this prediction mode, disparity maps need to be communicated along with the 2D array of views. In this work we explore a new adaptive wavelet transform for coding disparity maps, known as tri-breakpoint (TriBRK) dependent DWT, that is currently being considered as part of JPEG 2000 Part-17 extensions. We show that the scalable TriBRK dependent DWT, defined on a hierarchical triangular grid, provides rate-distortion gains for coding disparity maps and improves the compression of HDCA views that rely upon the disparity information. Visual quality improvements are also observed, specifically at object boundaries of decoded views. Reji Mathew, David S. Taubman |
ICIP | 1 |
| 2021 | Scalable Coding Of Motion And Depth Fields With Shared BreakpointsabstractA new breakpoint adaptive DWT, referred to as tri-break, is currently being considered by standardization efforts in relation to JPEG 2000 Part 17 extensions. We first provide a summary of the tri-break transform and then explore its performance for coding motion fields. Experimental results show that significant gains can be achieved for coding piecewise smooth motion flows by employing the tri-break transform. We demonstrate the feasibility of utilising a common set of breakpoints for compressing depth maps and motion fields anchored at the same frame. We also extend prior work to decode motion vector fields directly onto a triangular mesh, enabling operations such as view warping to be defined on triangular cells which can be efficiently processed by GPU based architectures. Reji Mathew, Yue Li 0034, David S. Taubman |
ICIP | 1 |
| 2020 | Scalable Mesh Representation for Depth from Breakpoint-Adaptive Wavelet CodingabstractA highly scalable and compact representation of depth data is required in many applications, and it is especially critical for plenoptic multiview image compression frameworks that use depth information for novel view synthesis and interview prediction. Efficiently coding depth data can be difficult as it contains sharp discontinuities. Breakpoint-adaptive discrete wavelet transforms (BPA-DWT) currently being standardized as part of JPEG 2000 Part-17 extensions have been found suitable for coding spatial media with hard discontinuities. In this paper, we explore a modification to the original BPA-DWT by replacing the traditional constant extrapolation strategy with the newly proposed affine extrapolation for reconstructing depth data in the vicinity of discontinuities. We also present a depth reconstruction scheme that can directly decode the BPA-DWT coefficients and breakpoints onto a compact and scalable mesh-based representation which has many potential benefits over the sample-based description. For performing depth compensated view prediction, our proposed triangular mesh representation of the depth data is a natural fit for modern graphics architectures. Yue Li 0034, Reji Mathew, David S. Taubman |
MMSP | 2 |
| 2020 | Rate-Distortion Driven Decomposition of Multiview Imagery to Diffuse and Specular ComponentsabstractIn this work, we propose an overcomplete representation of multiview imagery for the purpose of compression. We present a rate-distortion (R-D) driven approach to decompose multiview datasets into two additive parts which can be interpreted as diffuse and specular content. We choose distinct and different sparsifying transforms for the diffuse and specular components and employ an R-D inspired measure as our optimization cost function to drive the decomposition based solely on compressibility. We first describe a framework which performs data separation in a registered domain to avoid the complexity of warping between views. Then a more comprehensive approach is proposed to separate specular data progressively from coordinates of multiple reference views. Experimental results show a coding gain of up to 0.6 dB for synthetic datasets and up to 0.9 dB for real datasets. Maryam Haghighat, Reji Mathew, David S. Taubman |
IEEE Trans. Image Process. | 2 |
| 2019 | Rate-Distortion Driven Separation of Diffuse and Specular Components in Multiview ImageryabstractIn this work we explore an overcomplete representation of multiview imagery for the purpose of compression. We present a rate-distortion (R-D) driven approach to decompose multiview datasets into two additive parts which can be interpreted as being the diffuse and specular components. We apply different transforms to each component such that the compressibility of input data is improved. We describe a framework which performs the R-D optimized separation in a registered domain to avoid the complexity of warping between views. Experimental results highlight the benefits of the proposed source separation approach in the context of compression. Maryam Haghighat, Reji Mathew, David S. Taubman |
ICIP | 2 |
| 2019 | WaSP Encoder with Breakpoint Adaptive DWT Coding of Disparity MapsabstractAn encoder architecture that employs view warping and sparse prediction (WaSP) has recently been proposed by the JPEG Pleno standards activity for the compression of light field imagery. The proposed WaSP encoder utilises disparity information to exploit the correlation that exists between multiple views and therefore requires disparity data to be communicated. The WaSP framework currently employs JPEG2000 for the coding of disparity maps. In this work we explore the impact of introducing breakpoint adaptive DWT (BPA-DWT) coding of disparity maps. While prior work has shown promising results for the coding of depth maps with breakpoints, its impact on multi-view compression where the depth or disparity data forms part of the communicated side information has not been previously explored. Our investigations show important gains in RD performance coupled with improvements in the visual quality of decoded views by the introduction of BPA-DWT coding of disparity maps. Reji Mathew, David S. Taubman |
ICIP | 1 |
| 2019 | High Throughput Block Coding in the HTJ2K Compression StandardabstractThis paper describes the block coding algorithm that underpins the new High Throughput JPEG 2000 (HTJ2K) standard. The objective of HTJ2K is to overcome the computational complexity of the original block coding algorithm, by providing a drop-in replacement that preserves as much of the JPEG 2000 feature set as possible, while allowing reversible transcoding to/from the original format. We show how the new standard achieves these goals, with high coding efficiency, and extremely high throughput in software. David S. Taubman, Aous Thabit Naman, Reji Mathew |
ICIP | 3 |
| 2019 | Consistent Disparity Synthesis for Inter-View Prediction in Lightfield CompressionabstractFor efficient compression of lightfields that involve many views, it has been found preferable to explicitly communicate disparity/depth information at only a small subset of the view locations. In this study, we focus solely on inter-view prediction, which is fundamental to multi-view imagery compression, and itself depends upon the synthesis of disparity at new view locations. Current HDCA standardization activities consider a framework known as WaSP, that hierarchically predicts views, independently synthesizing the required disparity maps at the reference views for each prediction step. A potentially better approach is to progressively construct a unified multi-layered base-model for consistent disparity synthesis across many views. This paper improves significantly upon an existing base-model approach, demonstrating superior performance to WaSP. More generally, the paper investigates the implications of texture warping and disparity synthesis methods. Yue Li 0034, Reji Mathew, Dominic Rüfenacht, Aous Thabit Naman, David S. Taubman |
PCS | 2 |
| 2019 | Temporal Frame Interpolation With Motion-Divergence-Guided Occlusion HandlingabstractWe present a high-quality temporal frame interpolation (TFI) method that employs piecewise-smooth motion and handles disoccluded regions using the observation that motion discontinuities travel with the foreground object. We derive a “motion discontinuity” likelihood map from the divergence of a motion field between the input frames. Motion which is modeled at the reference frame is mapped to the target frame using a cellular-affine mapping strategy - a process during which regions of disocclusion are readily observed. This information is then used to guide the occlusion-aware, bidirectional FI process. Furthermore, we propose two computationally inexpensive texture optimizations that selectively improve the quality of the interpolated frames in regions around moving objects. The scheme produces very high-quality interpolated frames and outperforms current high-quality state-of-the-art TFI schemes by 2-2.5 dB; the method works with a very low-complexity motion estimation scheme and runs orders of magnitudes faster than its competitors. Dominic Rüfenacht, Reji Mathew, David S. Taubman |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Illumination Estimation and Compensation of Low Frame Rate Video Sequences for Wavelet-Based Video CompressionabstractIn this paper, we are interested in the compression of image sets or video with considerable changes in illumination. We develop a framework to decompose frames into illumination fields and texture in order to achieve sparser representations of frames which is beneficial for compression. Illumination variations or contrast ratio factors among frames are described by a full resolution multiplicative field. First, we propose a Lifting-based Illumination Adaptive Transform (LIAT) framework which incorporates illumination compensation to temporal wavelet transforms. We estimate a full resolution illumination field, taking heed of its spatial sparsity by a rate-distortion (R-D) driven framework. An affine mesh model is also developed as a point of comparison. We find the operational coding cost of the subband frames by modeling a typical t + 2D wavelet video coding system. While our general findings on R-D optimization are applicable to a range of coding frameworks, in this paper, we report results based on employing JPEG 2000 coding tools. The experimental results highlight the benefits of the proposed R-D driven illumination estimation and compensation in comparison with alternative scalable coding methods and non-scalable coding schemes of AVC and HEVC employing weighted prediction. Maryam Haghighat, Reji Mathew, Aous Thabit Naman, David S. Taubman |
IEEE Trans. Image Process. | 2 |
| 2019 | Base-Anchored Model for Highly Scalable and Accessible Compression of Multiview ImageryabstractWe present a compression scheme for multiview imagery that facilitates high scalability and accessibility of the compressed content. Our scheme relies upon constructing at a single base view, a disparity model for a group of views, and then utilizing this base-anchored model to infer disparity at all views belonging to the group. We employ a hierarchical disparity-compensated inter-view transform where the corresponding analysis and synthesis filters are applied along the geometric flows defined by the base-anchored disparity model. The output of this inter-view transform along with the disparity information is subjected to spatial wavelet transforms and embedded block-based coding. Rate-distortion results reveal superior performance to the x.265 anchor chosen by the JPEG Pleno standards activity for the coding of multiview imagery captured by high-density camera arrays. Dominic Rüfenacht, Aous Thabit Naman, Reji Mathew, David S. Taubman |
IEEE Trans. Image Process. | 3 |
| 2018 | Rate-Distortion Optimized Illumination Estimation for Wavelet-Based Video CodingabstractWe propose a rate-distortion optimized framework for estimating illumination changes (lighting variations, fade in/out effects) in a highly scalable coding system. Illumination variations are realized using multiplicative factors in the image domain and are estimated considering the coding cost of the illumination field and input frames which are first subject to a temporal Lifting-based Illumination Adaptive Transform (LIAT). The coding cost is modelled by an ℓ1-norm optimization problem which is derived to approximate a quadratic-log function which emerges from rate-distortion considerations. The optimization problem is solved using ADMM. The proposed solution works the same or better than a mesh-based approach proposed in prior work, where sparsity was controlled by explicitly choosing mesh parameters. In the compression-inspired formulation presented here, sparsity is discovered automatically through the solution of a convex program that depends only on a target rate-distortion operating point. Maryam Haghighat, Reji Mathew, Aous Thabit Naman, Sean I. Young, David S. Taubman |
ICASSP | 2 |
| 2017 | Lifting-based Illumination Adaptive Transform (LIAT) using mesh-based illumination modellingabstractState-of-the-art video coding techniques employ block-based illumination compensation to improve coding efficiency. In this work, we propose a Lifting-based Illumination Adaptive Transform (LIAT) to exploit temporal redundancy among frames that have illumination variations, such as the frames of low frame rate video or multi-view video. LIAT employs a mesh-based spatially affine model to represent illumination variations between two frames. In LIAT, transformed frames are jointly compressed, together with illumination information, into a layered rate-distortion optimal codestream, using the JPEG2000 format. We show that the LIAT framework significantly improves compression efficiency of temporal subband transforms for both predictive and more general transforms with predict and update steps. Maryam Haghighat, Reji Mathew, Aous Thabit Naman, David S. Taubman |
ICIP | 2 |
| 2016 | Efficient action recognition from compressed depth mapsabstractWe propose an efficient action recognition scheme based solely on compressed depth maps. Each depth map is coded by a recently proposed scalable encoder that employs multi-scale breakpoints and an adaptive discrete wavelet transform (DWT). DWT coefficients describe smooth variations in depth while breakpoints communicate sharp boundaries. Both of these attributes are extracted from the bit-stream and utilized to construct features which are subject to a classification scheme for human action recognition. By extracting features from the compressed bit-stream computational complexity is significantly reduced thereby making the proposed scheme suitable for real-time applications. A L2-regularized collaborative representation classifier is employed for classification. The proposed scheme is computationally more efficient when compared with conventional approaches. Experimental results on the MSR 3D action dataset validate the effectiveness and efficiency of our proposed scheme. Jie Miao, Xiaoyi Jia, Reji Mathew, Xiangmin Xu 0001, David S. Taubman, Chunmei Qing |
ICIP | 3 |
| 2016 | Optimization and compression of geometry discontinuities for graph-based representation of piecewise smooth mediaabstractEarlier research has shown the efficacy of using geometry discontinuities, encoded using “breakpoints,” in improving the coding efficiency of piecewise smooth media, such as depth maps and motion flows. This work proposes a new structure for encoding these breakpoints that is more suited for piecewise affine media, such as affine motion flows. Here, we choose to employ belief propagation over a graphical model to discover a set of breakpoints that is rate-distortion optimal. Similar to earlier works, the discovered breakpoints are employed in a breakpoint adaptive (BPA) wavelet decomposition of the media under consideration; the resulting BPA wavelet coefficients and the proposed “affine” breakpoints are then compressed using the JPEG2000 format, in a way that provides spatial and quality scalability as well as accessibility. We show that the coding cost of affine breakpoints is moderate, and that, for coding motion flows, the coding efficiency of the proposed approach is comparable to H.264. Aous Thabit Naman, David S. Taubman, Reji Mathew |
ICIP | 3 |
| 2016 | Optimizing block-coded motion parameters with block-partition graphsabstractWe address the problem of optimizing block-coded motion parameters for use inside typical motion-compensating video encoders. We cast the given discrete problem as a nonsmooth nonconvex optimization problem which is defined over some graph, and solve it using the split primal-dual hybrid gradient algorithm. Although computational efficiency is not the main focus of this paper, an efficient, parallelized implementation of our proposed approach can be used as a way of performing rate-distortion optimal motion estimation in video encoders such as those following the H.264 or HEVC standard. Results from our experiments highlight the degree of sub-optimality demonstrated by motion parameters that have been computed by H.264 block matching algorithms. Sean I. Young, Reji Mathew, David S. Taubman |
ICIP | 2 |
| 2016 | Higher-order motion models for temporal frame interpolation with applications to video codingabstractWe have recently proposed a motion-centric temporal frame interpolation (TFI) method, called BAM-TFI, which is able to produce high quality interpolated frames under a constant velocity assumption. However, for objects that do not follow constant velocity motion, the predictions, although credible, will differ from the “true” target frames, leading to high prediction residuals. In this paper, we show how higher-order motion models can be incorporated into the BAM-TFI scheme to interpolate frames that better predict the target frames. This opens up the door to a seamless integration of TFI with a video coding scheme. Comparisons on a variety of both synthetic and natural video sequences highlight the benefits of a second-order motion model. We further integrate the proposed TFI scheme into HEVC; preliminary comparisons with HEVC show promising results. Dominic Rüfenacht, Reji Mathew, David S. Taubman |
PCS | 2 |
| 2016 | Optimized decoding of JPEG images based on generalized graph LaplaciansabstractWe address the problem of optimizing the decoding of JPEG-compressed images, employing an approach based on the “generalized” graph Laplacian, a higher-order generalization of the usual graph Laplacian. The optimal decoding problem is formulated as a non-smooth but convex problem over a graph, and solved via the alternating directions method of multipliers. While similar, graph-based optimized decoding techniques exist, the use of our generalized graph Laplacian enables better recovery of the original smooth image from the transform coefficients, especially those coded at low bit-rates. Experimental results highlight the performance of our proposed approach in terms of reconstructed distortion and visual quality. Comparisons with the method based on the usual graph Laplacian and other state-of-the-art methods are also given. Sean I. Young, Aous Thabit Naman, Reji Mathew, David S. Taubman |
PCS | 3 |
| 2016 | A Novel Motion Field Anchoring Paradigm for Highly Scalable Wavelet-Based Video CodingabstractExisting video coders anchor motion fields at frames that are to be predicted. In this paper, we demonstrate how changing the anchoring of motion fields to reference frames has some important advantages over conventional anchoring. We work with piecewise-smooth motion fields, and use breakpoints to signal discontinuities at moving object boundaries. We show how discontinuity information can be used to resolve double mappings arising when motion is warped from reference to target frames. We present an analytical model that allows to determine weights for texture, motion, and breakpoints to guide the rate-allocation for scalable encoding. Compared with the conventional way of anchoring motion fields, the proposed scheme requires fewer bits for the coding of motion; furthermore, the reconstructed video frames contain fewer ghosting artefacts. The experimental results show the superior performance compared with the traditional anchoring, and demonstrate the high scalability attributes of the proposed method. Dominic Rüfenacht, Reji Mathew, David S. Taubman |
IEEE Trans. Image Process. | 2 |
| 2015 | Residue boundary histograms for action recognition in the compressed domainabstractTraditional action recognition approaches are too slow for real-time or large-scale applications. This problem has been tackled by replacing optical flow with motion vectors from the compressed domain. Yet further usage of compressed domain information for action recognition is possible. Discrete cosine transform (DCT) coefficients, which correspond to residue data, represent information which the block based motion vectors fail to capture. We propose a set of residue boundary histograms (RBH) features for action recognition, separating each DCT block into four parts to obtain four small residue maps and then encoding each residue map by histogram-based descriptors to obtain local features. Experimental results on three challenging datasets show that proposed RBH features improve upon motion vector based features significantly. While more than 100× faster, the results are highly competitive compared with traditional action recognition approaches. Jie Miao, Xiangmin Xu 0001, Reji Mathew |
ICIP | 3 |
| 2015 | Motion blur modelling for hierarchically anchored motion with discontinuitiesabstractWe have previously proposed a scheme for representing motion with motion discontinuities which has beneficial properties in terms of compactness (efficiency) and scalability. This so-called BIHA scheme has applications in video coding as well as temporal frame interpolation. In both cases, modelling of motion discontinuities has proven to be valuable. In these earlier works, we have ignored effects of motion blur, which can result in artificial sharp transitions of texture information at moving object boundaries. In this paper, we extend the BIHA framework to account for motion blur. Experimental results show significant improvements over the original BIHA scheme in texture blending regions, resulting in more visually pleasing predictions, as well as better rate-distortion performance. Dominic Rüfenacht, Reji Mathew, David S. Taubman |
MMSP | 2 |
| 2015 | Optimization of optical flow for scalable codingabstractOptical flow or dense motion field representations provide an alternative to block based schemes and are capable of describing smooth motion flows without introducing any artificial block boundaries. However dense motion fields have not been widely adopted for video coding applications principally due to the prohibitive cost of communicating such detailed motion representations. In this paper our focus is on estimating dense motion fields subject to R-D constraints derived from wavelet based coding of the motion data. We develop a probabilistic framework for estimating dense motion fields that is capable of evaluating the tradeoff between sparsity in the transformed domain and motion compensated distortion. R-D results show that we are able to improve upon traditional optical flow descriptions. Reji Mathew, Sean I. Young, David S. Taubman |
PCS | 1 |
| 2015 | Bidirectional, occlusion-aware temporal frame interpolation in a highly scalable video settingabstractWe present a bidirectional, occlusion-aware temporal frame interpolation (BOA-TFI) scheme that builds upon our recently proposed highly scalable video coding scheme. Unlike previous TFI methods, our scheme attempts to put “correct” information in problematic regions around moving objects. From a “parent” motion field between two existing reference frames, we compose motion from both reference frames to the target frame. These motion fields, together with motion discontinuity information, are then warped to the target frame - a process during which we discover valuable information about disocclusions, which we then use to guide the bidirectional prediction of the interpolated frame. The scheme can be used in any state-of-the-art codec, but is most beneficial if used in conjunction with a highly scalable video coder. Evaluation of the method on synthetic data allows us to shine a light on problematic regions around moving object boundaries, which has not been the focus of previous frame interpolation methods. The proposed frame interpolation method yields credible results, and compares favourably to current state-of-the-art frame interpolation methods. Dominic Rüfenacht, Reji Mathew, David S. Taubman |
PCS | 2 |
| 2015 | Spatial induction policies for scalable depth codingabstractAn edge-adaptive depth coding proposed in [1] uses a multi-resolution field of breakpoints to encode edges in a scene's geometry. In order to remain efficient, this scheme uses spatial induction to infer the location of breakpoints from the surrounding geometry where possible. This intelligence means only a subset of the breakpoints need to be encoded. The original proposal employs a simple policy, using linear interpolation between known points to induce breakpoints. Whilst this is effective, a more sophisticated policy could better exploit the available breakpoint data in synthesising edges. This paper presents our findings regarding the potential benefits of higher order spatial induction policies (SIPs). Mitchell S. Ward, David S. Taubman, Reji Mathew |
PCS | 3 |
| 2015 | Motion estimation with accurate boundariesabstractThis paper investigates several techniques that increase the accuracy of motion boundaries in estimated motion fields of a local dense estimation scheme. In particular, we examine two matching metrics, one is MSE in the image domain and the other one is a recently proposed multiresolution metric that has been shown to produce more accurate motion boundaries. We also examine several different edge-preserving filters. The edge-aware moving average filter, proposed in this paper, takes an input image and the result of an edge detection algorithm, and outputs an image that is smooth except at the detected edges. Compared to the adoption of edge-preserving filters, we find that matching metrics play a more important role in estimating accurate and compressible motion fields. Nevertheless, the proposed filter may provide further improvements in the accuracy of the motion boundaries. These findings can be very useful for a number of recently proposed scalable interactive video coding schemes. Rui Xu 0030, Aous Thabit Naman, Reji Mathew, Dominic Rüfenacht, David S. Taubman |
PCS | 3 |
| 2014 | Hierarchical anchoring of motion fields for fully scalable video codingabstractTraditional video codecs anchor motion fields in the frame that is to be predicted, which is natural in a non-scalable context. In this paper, we propose a hierarchical anchoring of motion fields at reference frames, which allows to “reuse” them at finer temporal levels - a very desirable property for temporal scalability. The main challenge using this approach is that the motion fields need to be warped to the target frames, leading to disocclusions and motion folding in the warped motion fields. We show how to resolve motion folding ambiguities that occur in the vicinity of moving object boundaries by using breakpoint fields that have recently been proposed for the scalable coding of motion. During the motion field warping process, we obtain disocclusion and folding maps on-the-fly, which are used to control the temporal update step of the Haar wavelet. Results on synthetic data show that the proposed hierarchical anchoring scheme outperforms the traditional way of anchoring motion fields. Dominic Rüfenacht, Reji Mathew, David S. Taubman |
ICIP | 2 |
| 2014 | Bidirectional hierarchical anchoring of motion fields for scalable video codingabstractThe ability to predict motion fields at finer temporal scales from coarser ones is a very desirable property for temporal scalability. This is at best very difficult in current state-of-the-art video codecs (i.e., H.264, HEVC), where motion fields are anchored in the frame that is to be predicted (target frame). In this paper, we propose to anchor motion fields in the reference frames. We show how from only one fully coded motion field at the coarsest temporal level as well as breakpoints which signal discontinuities in the motion field, we are able to reliably predict motion fields used at finer temporal levels. This significantly reduces the cost for coding the motion fields. Results on synthetic data show improved rate-distortion (R-D) performance and superior scalability, when compared to the traditional way of anchoring motion fields. Dominic Rüfenacht, Reji Mathew, David S. Taubman |
MMSP | 2 |
| 2014 | Embedded coding of optical flow fields for scalable video compressionabstractAn embedded coding scheme for dense motion (optical flow) fields is proposed. Such a scheme is particularly useful in scalable video compression where one must compensate for inter-frame motion at various visual qualities and resolutions. However, the high cost of coding such fields has often made this option prohibitive. Using our previously developed `breakpoint'-adaptive wavelet transform, we show that it is possible to code dense motion fields efficiently while simultaneously endowing the coded motion representation with embedded resolution and quality scalability attributes. Performance comparisons with the traditional non-scalable block-based model are also made and presented with the aid of a modified H.264/AVC JM reference encoder. Sean I. Young, Reji Mathew, David S. Taubman |
MMSP | 2 |
| 2013 | Robust sum of Linear-Log Squared Differences distortion measure and its applicationsabstractA robust distortion measure known as the Sum of Linear-Log Squared Differences (SLLSD) derived analytically from the R-D optimality conditions for compression is proposed with particular applications in motion estimation (ME) and coding. When used in the context of ME, this new measure is shown both to improve motion robustness (accuracy) and lower the Lagrangian cost (coding) unlike the typically employed Sum of Squared Differences (SSD) measure. Our measure is defined in the image domain and does not necessitate or assume the use of a particular transform. Its relationship to other M-estimators is briefly discussed and coding related performance results presented with the aid of a suitably modified H.264/AVC JM reference encoder. Sean I. Young, Reji Mathew, David S. Taubman |
PCS | 2 |
| 2013 | Scalable Coding of Depth Maps With R-D Optimized EmbeddingabstractRecent work on depth map compression has revealed the importance of incorporating a description of discontinuity boundary geometry into the compression scheme. We propose a novel compression strategy for depth maps that incorporates geometry information while achieving the goals of scalability and embedded representation. Our scheme involves two separate image pyramid structures, one for breakpoints and the other for sub-band samples produced by a breakpoint-adaptive transform. Breakpoints capture geometric attributes, and are amenable to scalable coding. We develop a rate-distortion optimization framework for determining the presence and precision of breakpoints in the pyramid representation. We employ a variation of the EBCOT scheme to produce embedded bit-streams for both the breakpoint and sub-band data. Compared to JPEG 2000, our proposed scheme enables the same the scalability features while achieving substantially improved rate-distortion performance at the higher bit-rate range and comparable performance at the lower rates. Reji Mathew, David S. Taubman, Pietro Zanuttigh |
IEEE Trans. Image Process. | 1 |
| 2012 | Highly Scalable Coding of Depth Maps with Arc BreakpointsabstractRecent work highlights the importance of incorporating geometry information into the compression of depth maps. For many applications, features such as resolution scalability and embedded coding are also highly desirable. JPEG 2000 offers these scalability features but suffers from poor compression performance in the vicinity of strong discontinuities. We propose a novel compression strategy for depth maps that incorporates geometry information while retaining the highly scalable coding properties of JPEG 2000. Our scheme involves two separate image pyramid structures, one for arc breakpoints and other for sub-band samples produced by a breakpoint-adaptive transform. Breakpoints capture geometric attributes and are also amenable to scalable coding. We develop an R-D optimization framework for the breakpoint data. We also use a variation of the EBCOT scheme to produce embedded bit-streams for both the breakpoint and sub-band data, allowing them to be independently and incrementally sequenced based on R-D considerations. Reji Mathew, Pietro Zanuttigh, David S. Taubman |
DCC | 1 |
| 2012 | Scalable depth maps with R-D optimized embeddingabstractRecent work has highlighted the importance of incorporating geometry information into the compression of depth maps. In prior approaches however the geometry information is not resolution scalable nor amenable to embedded coding. In this paper we propose a novel compression strategy for depth maps that incorporates geometry information while achieving the goals of scalability and embedded representation. Our scheme involves two separate image pyramid structures, one for breakpoints and other for sub-band samples produced by a breakpoint-adaptive transform. Breakpoints capture geometric attributes and are amenable to scalable coding. We develop an R-D optimization framework for the breakpoint data. We also use a variation of the EBCOT scheme to produce embedded bit-streams for both the breakpoint and sub-band data, allowing them to be independently and incrementally sequenced based on R-D considerations. Reji Mathew, David S. Taubman, Pietro Zanuttigh |
MMSP | 1 |
| 2011 | Scalable Modeling of Motion and Boundary Geometry With Quad-Tree Node MergingabstractQuad-tree structures are often used to model motion between frames of a video sequence. However, a fundamental limitation of the quad-tree structure is that it can only capture horizontal and vertical edge discontinuities at dyadically related locations. To address this limitation, recent work has focused on the introduction of geometry information to nodes of tree structured motion representations. In this paper we explore modeling boundary geometry and motion with separate quad-tree structures; thereby enabling each attribute to be refined separately. Recent work into quad-tree representations has also highlighted the benefits of leaf merging. We extend the leaf merging paradigm to incorporate both geometry and motion attributes. We also explore resolution scalability of the merged dual tree representation and present experimental results which demonstrate both rate-distortion and scabalility performance. Theoretical investigations conducted in this paper reveal that to achieve optimal rate-distortion behavior, quad-tree motion models need to incorporate both geometry information and node merging. Reji Mathew, David S. Taubman |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2010 | Quad-Tree Motion Modeling With Leaf MergingabstractIn this paper, we are concerned with the modeling of motion between frames of a video sequence. Typically, it is not possible to represent the motion between frames by a single model and therefore a quad-tree structure is often employed where smaller, variable size regions or blocks are allowed to take on separate motion models. Previous work into quad-tree representations has demonstrated the sub-optimal performance of quad-trees where the dependency between neighboring leaf nodes with different parents is not exploited. Leaf merging has been proposed to rectify this performance loss as it allows joint coding and optimization of related nodes. In this paper, we describe how the merging step can be incorporated into quad-tree motion representations for a range of motion modeling contexts. In particular, we study the impact of rate-distortion optimized merging for two motion coding schemes, these being spatially predictive coding, as used in H.264, and hierarchical coding. We present experimental results which demonstrate that node merging can provide significant gains for both the hierarchical and spatial prediction schemes. Interestingly, experimental results also show that in the presence of merging, the rate-distortion performance of hierarchical coding is comparable to that of spatial prediction. We pursue the case of hierarchical coding further in this paper, introducing polynomial motion models to the quad-tree representation and exploring resolution scalability of the merged quad-tree structure. We also present a theoretical study of the impact of leaf merging in modeling motion, identifying the inherent advantages of merging which give rise to a more efficient description of frame motion. Reji Mathew, David S. Taubman |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Joint scalable modeling of motion and boundary geometry with quad-tree node mergingabstractQuad-tree structures are often used to model motion between frames of a video sequence. In this study we are interested in the bit-rate efficiency and resolution scalability of quad-tree motion models. Publications have reported improvements to motion modeling efficiency achieved by introducing geometry information to nodes of quad-tree structures. The benefits of leaf merging to tree based representations have also been highlighted. In keeping with these findings we explore scalability in the context of merged quad-tree representations, modeling the underlying motion flow with joint geometry and motion models. We employ hierarchical coding of the merged quad-tree and ensure that the merging process retains the property of resolution scalability. We show that the performance of scalable coding can be significantly improved by incorporating a new cost objective which takes into account the possibility of scalable decoding. Experimental results reveal that these scalability improvements are achieved without significant loss in overall efficiency and with competitive performance at all resolutions. Reji Mathew, David S. Taubman |
ICIP | 1 |
| 2008 | Joint motion and geometry modeling with quad-tree leaf mergingabstractQuad-tree structures are often used to model motion between frames of a video sequence. However, a fundamental limitation of the quad-tree structure is that it can only capture horizontal and vertical edge discontinuities at dyadically related locations. To address this limitation recent work has focused on the introduction of geometry information to nodes of tree structured motion representations. Recent research into quad-tree structures have also demonstrated the importance of leaf merging. In this paper we create a general quadtree structure, well suited to jointly representing geometry and motion information. We then explore tree pruning and leaf merging, where the geometry information is treated as an equal partner with motion. To achieve an efficient joint representation we improve on the estimation algorithm that detects boundary geometry and introduce polynomial motion models. Experimental results show that the approach taken in this paper provides significant improvement over previous quad-tree based motion representation schemes. Reji Mathew, David S. Taubman |
ICIP | 1 |
| 2008 | Motion modeling with separate quad-tree structures for geometry and motionabstractQuad-tree structures are often used to model motion between frames of a video sequence. However, a fundamental limitation of the quad-tree structure is that it can only capture horizontal and vertical edge discontinuities at dyadically related locations. To address this limitation recent work has focused on the introduction of geometry information to nodes of tree structured motion representations. In this paper we explore modeling boundary geometry and motion with separate quadtree structures. Recent work into quad-tree representations have also highlighted the benefits of leaf merging. We extend the leaf merging paradigm to incorporate separate tree structures for boundary geometry and motion. To achieve an efficient joint representation we introduce polynomial motion models and piecewise linear boundary geometry to our quad-tree structures. Experimental results show that the approach taken in this paper provides significant improvement over previous quad-tree based motion representation schemes. Reji Mathew, David S. Taubman |
MMSP | 1 |
| 2007 | Motion Modeling with Geometry and Quad-tree Leaf MergingabstractQuad-tree structures are often used to model motion between frames of a video sequence. However, a fundamental limitation of the quadtree structure is that it can only capture horizontal and vertical edge discontinuities at dyadically related locations. In this paper we seek to address this limitation by introducing geometry information into the nodes of a pruned quad-tree. We start with a typical optimally pruned quad-tree where each node is allowed to model motion. Then for each node in the tree, we consider augmenting the node's motion model with a linear geometry model. Experimental results show that the introduction of geometry information improves the performance of quad-trees in representing motion. Recent work into quad-tree representations have highlighted the benefits of leaf merging. In this paper we extend the leaf merging paradigm to incorporate both geometry and motion information, allowing the creation of regions that have both motion and geometry attributes. Reji Mathew, David S. Taubman |
ICIP (2) | 1 |
| 2006 | Hierarchical and Polynomial Motion Modeling with Quad-Tree Leaf MergingabstractRecent work into quad-tree representations has commented on the sub-optimal performance of quad-trees due to not exploiting the dependency between neighboring leaf nodes with different parents. Leaf merging therefore has been proposed to rectify this performance loss. In the context of quad-tree motion models, the performance of leaf merging, was recently demonstrated in which an optimally pruned H.264 motion model was taken as the starting point for subsequent merging steps. In this paper we explore the performance of leaf merging over a wider range of conditions by starting with a general tree structure where each node is allowed to represent motion using either polynomial models or a single translational vector. We also explore two cases of motion prediction methods, these being spatial prediction and hierarchical prediction. Experimental results show the benefit of merging across these broad conditions. In comparison to the merged H.264 model reported in the previous work, substantial gains are evident. Furthermore we explore applying the merging principles to branch nodes of a quad-tree to achieve efficient hierarchical motion representation. Reji Mathew, David S. Taubman |
ICIP | 1 |
| 2005 | Detecting New Stable Objects In Surveillance VideoabstractWe describe a novel method to detect new stable objects in video. This includes detecting new objects that appear in a scene and remain stationary for a period of time. Examples include detecting a dropped bag or a parked car. Our method utilizes the state transition history (or a record of the "life cycle") of individual Gaussian distributions in a Gaussian Mixture Model (GMM) used to model the background. In typical implementations of the GMM, this state transition information is ignored however we show that by observing and retaining the history of state transitions of individual distributions, it is possible to detect long term changes in a scene. In particular we identify changes to the most probable background distribution and impose certain conditions on the characteristics and temporal behavior of this distribution. Results presented in this paper illustrate the success of the proposed method and its relevance to surveillance applications. Reji Mathew, Zhenghua Yu, Jian Zhang 0002 |
MMSP | 1 |
| 1999 | Efficient layered video coding using data partitioning
Reji Mathew, John F. Arnold |
Signal Process. Image Commun. | 1 |
| 1997 | Layered coding using bitstream decomposition with drift correctionabstractIt is well known that layered video coding is useful both for service interworking and as an aid to error resilience. A major drawback of layered coding is that it invariably results in a reduction in overall coding efficiency for the high quality service. An attractive approach would be one in which the high quality service is coded, and then the resulting bitstream is decomposed in such a way that the lower resolution services can be reconstructed using only a subset of the total generated data. This would mean that there would be no impact on the coding efficiency of the highest quality service. In this paper we demonstrate that the quality of the lower layer when this approach is used is fundamentally limited by drift. It is shown that even a coarsely quantized, low rate correction signal provides a major benefit to the quality of this lower layer service while still not impacting significantly on the coding efficiency of the high quality service. Of course, some overhead is introduced in the coding of the lower quality service but the total overhead is still significantly better than simulcast. Reji Mathew, John F. Arnold |
IEEE Trans. Circuits Syst. Video Technol. | 1 |