EDBT 2026 Demo / reviewers in the wild / expert
Aous Thabit Naman
dblp:29/5419
· DBLP profile ↗
45ranked-venue papers
16as first author
8since 2021 · last 2025
0000-0002-5774-7143ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 15 first-author · 8 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-View Disparity Estimation Using the Gradient Consistency ModelabstractVariational approaches to disparity estimation typically use a linearised brightness constancy constraint, which only applies in smooth regions and over small distances. Accordingly, current variational approaches rely on a schedule to progressively include image data. This paper proposes the use of Gradient Consistency information to assess the validity of the linearisation; this information is used to determine the weights applied to the data term as part of an analytically inspired Gradient Consistency Model. The Gradient Consistency Model penalises the data term for view pairs that have a mismatch between the spatial gradients in the source view and the spatial gradients in the target view. Instead of relying on a tuned or learned schedule, the Gradient Consistency Model is self-scheduling, since the weights evolve as the algorithm progresses. We show that the Gradient Consistency Model outperforms standard coarse-to-fine schemes and the recently proposed progressive inclusion of views approach in both rate of convergence and accuracy. James L. Gray, Aous Thabit Naman, David S. Taubman |
IEEE Trans. Image Process. | 2 |
| 2024 | Exploration of Learned Lifting-Based Transform Structures for Fully Scalable and Accessible Wavelet-Like Image CompressionabstractThis paper provides a comprehensive study on features and performance of different ways to incorporate neural networks into lifting-based wavelet-like transforms, within the context of fully scalable and accessible image compression. Specifically, we explore different arrangements of lifting steps, as well as various network architectures for learned lifting operators. Moreover, we examine the impact of the number of learned lifting steps, the number of channels, the number of layers and the support of kernels in each learned lifting operator. To facilitate the study, we investigate two generic training methodologies that are simultaneously appropriate to a wide variety of lifting structures considered. Experimental results ultimately suggest that retaining fixed lifting steps from the base wavelet transform is highly beneficial. Moreover, we demonstrate that employing more learned lifting steps and more layers in each learned lifting operator do not contribute strongly to the compression performance. However, benefits can be obtained by utilizing more channels in each learned lifting operator. Ultimately, the learned wavelet-like transform proposed in this paper achieves over 25% bit-rate savings compared to JPEG 2000 with compact spatial support. Xinyue Li 0002, Aous Thabit Naman, David S. Taubman |
IEEE Trans. Image Process. | 2 |
| 2023 | Improved Transform Structures for Learned Wavelet-Like Fully Scalable Image CompressionabstractThis paper studies features and performance of different structures for learned wavelet-like transforms in fully scalable image compression. Specifically, we explore different neural network topologies and various arrangements of lifting steps to improve the existing wavelet transform. Experimental results strongly suggest that the proposed proposal-opacity network topology, comprising a collection of linear predictions modulated by non-linear opacities, performs better than the other considered designs. Results also show that augmenting a good base wavelet transform with two learned lifting steps performs significantly better than other learned lifting structures, achieving 26.8% bit-rate savings over JPEG 2000. Xinyue Li 0002, Aous Thabit Naman, David S. Taubman |
MMSP | 2 |
| 2023 | JPEG 2000 Extensions for Scalable Coding of Discontinuous MediaabstractIn this paper we propose novel extensions to JPEG 2000 for the coding of discontinuous media which includes piecewise smooth imagery such as depth maps and optical flows. These extensions use breakpoints to model discontinuity boundary geometry and apply a breakpoint dependent Discrete Wavelet Transform (BP-DWT) to the input imagery. The highly scalable and accessible coding features provided by the JPEG 2000 compression framework are preserved by our proposed extensions, with the breakpoint and transform components encoded as independent bit streams that can be progressively decoded. Comparative rate-distortion results are provided along with corresponding visual examples which highlight the advantages of using breakpoint representations with accompanying BD-DWT and embedded bit-plane coding. Recently our proposed extensions have been adopted and are in the process of being published as a new Part 17 to the JPEG 2000 family of coding standards. Reji Mathew, Aous Thabit Naman, Yue Li 0034, David S. Taubman |
IEEE Trans. Image Process. | 2 |
| 2022 | A Neural Network Lifting Based Secondary Transform for Improved Fully Scalable Image Compression in Jpeg 2000abstractThis paper proposes a secondary transform to improve wavelet-based image compression schemes such as JPEG 2000. The proposed approach includes two neural network steps, a high-to-low step followed by a low-to-high step. The high-to-low step suppresses aliasing in the low-pass band by using the detail bands at the same resolution, while the low-to-high step aims to further remove redundant information from the detail bands so as to achieve higher energy compaction. We employ the same neural networks for these two steps at every level of the hierarchical discrete wavelet transform (DWT) decomposition. We demonstrate that the visual quality of the LL bands at different resolutions can be dramatically enhanced, along with greatly improved coding efficiency up to 2dB compared with JPEG 2000, while preserving the full quality and resolution scalability and spatial random access features of JPEG 2000. Xinyue Li 0002, Aous Thabit Naman, David S. Taubman |
ICIP | 2 |
| 2022 | Gradient Consistency Based Multi-Scale Optical FlowabstractEstimating optical flow and/or depth from multiple views or frames, all with multiple scales can be considered a data selection problem. We propose a gradient consistency based multi-scale optical flow technique which aims to address this data selection problem. We propose a method to assess gradient consistency and show how to use it to appropriately downweight data with poor consistency. We also introduce a spatial regularisation term that exploits optical flow acceleration between consecutive frame pairs as part of a global variational framework, coupling data between consecutive frame pairs. The proposed gradient consistency based multi-scale approach produces improved results on both real and synthetic scenes compared to a coarse to fine framework especially with limited numbers of warps. James L. Gray, Aous Thabit Naman, David S. Taubman |
MMSP | 2 |
| 2021 | Welsch Based Multiview Disparity EstimationabstractIn this work, we explore disparity estimation from a high number of views. We experimentally identify occlusions as a key challenge for disparity estimation for applications with high numbers of views. In particular, occlusions can actually result in a degradation in accuracy as more views are added to a dataset. We propose the use of a Welsch loss function for the data term in a global variational framework for disparity estimation. We also propose a disciplined warping strategy and a progressive inclusion of views strategy that can reduce the need for coarse to fine strategies that discard high spatial frequency components from the early iterations. Experimental results demonstrate that the proposed approach produces superior and/or more robust estimates than other conventional variational approaches. James L. Gray, Aous Thabit Naman, David S. Taubman |
ICIP | 2 |
| 2021 | Machine-Learning Based Secondary Transform for Improved Image Compression in JPEG2000abstractThis paper proposes a convolutional neural network (CNN) based secondary transform for not only improved coding efficiency in the JPEG2000 image compression format, but also to produce more appealing approximation sub-bands at different resolutions. The CNN in this work exploits information in detail sub-bands to predict some of the aliasing information in the corresponding approximation or low-pass sub-band; this reduction in aliasing, although not perfect, improves the compressibility of the “cleaned” approximation subband. This process is repeated in subsequent wavelet decomposition levels to further improve coding efficiency. Experimental results show that, at high bit rates, the proposed network outperforms conventional JPEG2000 compression framework by up to 1.2 dB, especially for images with strong geometric flow. Xinyue Li 0002, Aous Thabit Naman, David S. Taubman |
ICIP | 2 |
| 2020 | Adaptive Secondary Transform For Improved Image Coding Efficiency In JPEG2000abstractThis paper proposes a secondary transform for wavelet based image compression, whose aim is to predict detail sub-bands from the corresponding low-pass sub-band so as to improve energy compaction. The proposed prediction strategy exploits the orientation of local image features to untangle aliasing components in the low-pass sub-band, and achieves prediction by projecting the cleaned synthesized low-pass sub-band back into the wavelet domain. At low bit rates, the proposed scheme can considerably reduce distortion in a JPEG 2000 based compression framework, especially for images with strong geometric flow. Xinyue Li 0002, Aous Thabit Naman, David S. Taubman |
ICIP | 2 |
| 2020 | Encoding High-Throughput Jpeg2000 (Htj2k) Images On A GpuabstractHigh-Throughput JPEG2000 (HTJ2K) is a new addition to the JPEG2000 suite of coding tools; it has been recently approved as Part-15 of the JPEG2000 standard, and the JPH file extension has been designated for it. The HTJ2K employs a new “fast” block coder that can achieve higher encoding and decoding throughput than a conventional JPEG2000 (C-J2K) encoder. The higher throughput is achieved because the HTJ2K codec processes wavelet coefficients in a smaller number of steps than C-J2K. Moreover, the HTJ2K block coder is also more amenable to parallelizable high-speed software and hardware implementations. The HTJ2K retains most of the features and capabilities of JPEG2000, and it also supports lossless transcoding between HTJ2K and already compressed C-J2K images. Quality scalability however is more limited than C-J2K. In a recent work, we presented preliminary performance results for decoding HTJ2K images on a GPU. In this work, we present a GPU-based HTJ2K encoder; we also present early encoding results for such an encoder, showing that it is possible to encode 4K4:4:4 HDR videos at more than 70 frames per second (fps) on a low-end card, while a high-end GPU can encode such videos at more than 400 fps. Aous Thabit Naman, David S. Taubman |
ICIP | 1 |
| 2020 | Graph Laplacian Regularization for Robust Optical Flow EstimationabstractThis paper proposes graph Laplacian regularization for robust estimation of optical flow. First, we analyze the spectral properties of dense graph Laplacians and show that dense graphs achieve a better trade-off between preserving flow discontinuities and filtering noise, compared with the usual Laplacian. Using this analysis, we then propose a robust optical flow estimation method based on Gaussian graph Laplacians. We revisit the framework of iteratively reweighted least-squares from the perspective of graph edge reweighting, and employ the Welsch loss function to preserve flow discontinuities and handle occlusions. Our experiments using the Middlebury and MPI-Sintel optical flow datasets demonstrate the robustness and the efficiency of our proposed approach. Sean I. Young, Aous Thabit Naman, David S. Taubman |
IEEE Trans. Image Process. | 2 |
| 2019 | Solving Vision Problems via FilteringabstractWe propose a new, filtering approach for solving a large number of regularized inverse problems commonly found in computer vision. Traditionally, such problems are solved by finding the solution to the system of equations that expresses the first-order optimality conditions of the problem. This can be slow if the system of equations is dense due to the use of nonlocal regularization, necessitating iterative solvers such as successive over-relaxation or conjugate gradients. In this paper, we show that similar solutions can be obtained more easily via filtering, obviating the need to solve a potentially dense system of equations using slow iterative methods. Our filtered solutions are very similar to the true ones, but often up to 10 times faster to compute. Sean I. Young, Aous Thabit Naman, Bernd Girod, David S. Taubman |
ICCV | 2 |
| 2019 | Leveraging the Discrete Cosine Basis for Better Motion Modelling in Highly Textured Video SequencesabstractMotion modelling plays a central role in video compression. This role is even more critical in highly textured video sequences, whereby a small error can produce large residuals that are costly to compress. While the translational motion model employed by existing coding standards, such as HEVC, is sufficient in most cases, using higher order models is beneficial; for this reason, the upcoming video coding standard, VVC, employs a 4-parameter affine model. In this work, we explore the use of the discrete cosine basis for motion modelling in highly textured video sequences, and show that this is beneficial. In particular, we use a single high-order model to describe a frame's motion; we employ this motion to produce an extra prediction reference, which is added to the HEVC list of references. Experimental results show that a median delta bit rate of 4.44% is achievable over conventional HEVC if this extra reference frame is used in addition to the temporal references offered by HEVC. Ashek Ahmmed, Aous Thabit Naman, Mark R. Pickering |
ICIP | 2 |
| 2019 | Decoding High-Throughput Jpeg2000 (HTJ2K) On A GabstractHigh-throughput JPEG2000 (HTJ2K), also known as JPEG 2000 Part 15, is the most recent addition to the JPEG2000 suite of coding tools. Ee file extension JPH has been designated for compressed images employing this new part of the standard. Eis new part describes a "fast" block coder for the JPEG 2000 format, while retaining most other JPEG2000 features and capabilities intact. Ee HTJ2K block coder is amenable to parallelizable high-speed encoding and decoding implementations; moreover, it is designed to allow lossless transcoding of already compressed JPEG2000 images that employ the regular block coder. HTJ2K supports the scalability options available in the JPEG2000 format except for quality scalability, which is available only to a limited extent. Eis work gives a high-level overview of this new block coder; we also present preliminary performance results for a GPU implementation. We show that a low-end GPU can decode 4K 4:4:4 12-bit videos at more than 60 frames per second (fps) while a high-end GPU can decode 8K HDR videos at more than 120 fps. Aous Thabit Naman, David S. Taubman |
ICIP | 1 |
| 2019 | High Throughput Block Coding in the HTJ2K Compression StandardabstractThis paper describes the block coding algorithm that underpins the new High Throughput JPEG 2000 (HTJ2K) standard. The objective of HTJ2K is to overcome the computational complexity of the original block coding algorithm, by providing a drop-in replacement that preserves as much of the JPEG 2000 feature set as possible, while allowing reversible transcoding to/from the original format. We show how the new standard achieves these goals, with high coding efficiency, and extremely high throughput in software. David S. Taubman, Aous Thabit Naman, Reji Mathew |
ICIP | 2 |
| 2019 | Consistent Disparity Synthesis for Inter-View Prediction in Lightfield CompressionabstractFor efficient compression of lightfields that involve many views, it has been found preferable to explicitly communicate disparity/depth information at only a small subset of the view locations. In this study, we focus solely on inter-view prediction, which is fundamental to multi-view imagery compression, and itself depends upon the synthesis of disparity at new view locations. Current HDCA standardization activities consider a framework known as WaSP, that hierarchically predicts views, independently synthesizing the required disparity maps at the reference views for each prediction step. A potentially better approach is to progressively construct a unified multi-layered base-model for consistent disparity synthesis across many views. This paper improves significantly upon an existing base-model approach, demonstrating superior performance to WaSP. More generally, the paper investigates the implications of texture warping and disparity synthesis methods. Yue Li 0034, Reji Mathew, Dominic Rüfenacht, Aous Thabit Naman, David S. Taubman |
PCS | 4 |
| 2019 | Illumination Estimation and Compensation of Low Frame Rate Video Sequences for Wavelet-Based Video CompressionabstractIn this paper, we are interested in the compression of image sets or video with considerable changes in illumination. We develop a framework to decompose frames into illumination fields and texture in order to achieve sparser representations of frames which is beneficial for compression. Illumination variations or contrast ratio factors among frames are described by a full resolution multiplicative field. First, we propose a Lifting-based Illumination Adaptive Transform (LIAT) framework which incorporates illumination compensation to temporal wavelet transforms. We estimate a full resolution illumination field, taking heed of its spatial sparsity by a rate-distortion (R-D) driven framework. An affine mesh model is also developed as a point of comparison. We find the operational coding cost of the subband frames by modeling a typical t + 2D wavelet video coding system. While our general findings on R-D optimization are applicable to a range of coding frameworks, in this paper, we report results based on employing JPEG 2000 coding tools. The experimental results highlight the benefits of the proposed R-D driven illumination estimation and compensation in comparison with alternative scalable coding methods and non-scalable coding schemes of AVC and HEVC employing weighted prediction. Maryam Haghighat, Reji Mathew, Aous Thabit Naman, David S. Taubman |
IEEE Trans. Image Process. | 3 |
| 2019 | Base-Anchored Model for Highly Scalable and Accessible Compression of Multiview ImageryabstractWe present a compression scheme for multiview imagery that facilitates high scalability and accessibility of the compressed content. Our scheme relies upon constructing at a single base view, a disparity model for a group of views, and then utilizing this base-anchored model to infer disparity at all views belonging to the group. We employ a hierarchical disparity-compensated inter-view transform where the corresponding analysis and synthesis filters are applied along the geometric flows defined by the base-anchored disparity model. The output of this inter-view transform along with the disparity information is subjected to spatial wavelet transforms and embedded block-based coding. Rate-distortion results reveal superior performance to the x.265 anchor chosen by the JPEG Pleno standards activity for the coding of multiview imagery captured by high-density camera arrays. Dominic Rüfenacht, Aous Thabit Naman, Reji Mathew, David S. Taubman |
IEEE Trans. Image Process. | 2 |
| 2019 | COGL: Coefficient Graph Laplacians for Optimized JPEG Image DecodingabstractWe address the problem of decoding joint photographic experts group (JPEG)-encoded images with less visual artifacts. We view the decoding task as an ill-posed inverse problem and find a regularized solution using a convex, graph Laplacian-regularized model. Since the resulting problem is non-smooth and entails non-local regularization, we use fast high-dimensional Gaussian filtering techniques with the proximal gradient descent method to solve our convex problem efficiently. Our patch-based "coefficient graph" is better suited than the traditional pixel-based ones for regularizing smooth non-stationary signals such as natural images and relates directly to classic non-local means de-noising of images. We also extend our graph along the temporal dimension to handle the decoding of M-JPEG-encoded video. Despite the minimalistic nature of our convex problem, it produces decoded images with similar quality to other more complex, state-of-the-art methods while being up to five times faster. We also expound on the relationship between our method and the classic ANCE method, reinterpreting ANCE from a graph-based regularization perspective. Sean I. Young, Aous Thabit Naman, David S. Taubman |
IEEE Trans. Image Process. | 2 |
| 2018 | Rate-Distortion Optimized Illumination Estimation for Wavelet-Based Video CodingabstractWe propose a rate-distortion optimized framework for estimating illumination changes (lighting variations, fade in/out effects) in a highly scalable coding system. Illumination variations are realized using multiplicative factors in the image domain and are estimated considering the coding cost of the illumination field and input frames which are first subject to a temporal Lifting-based Illumination Adaptive Transform (LIAT). The coding cost is modelled by an ℓ1-norm optimization problem which is derived to approximate a quadratic-log function which emerges from rate-distortion considerations. The optimization problem is solved using ADMM. The proposed solution works the same or better than a mesh-based approach proposed in prior work, where sparsity was controlled by explicitly choosing mesh parameters. In the compression-inspired formulation presented here, sparsity is discovered automatically through the solution of a convex program that depends only on a target rate-distortion operating point. Maryam Haghighat, Reji Mathew, Aous Thabit Naman, Sean I. Young, David S. Taubman |
ICASSP | 3 |
| 2018 | Enhanced Homogeneous Motion Discovery Oriented Prediction for Key Intermediate FramesabstractConventional video compression systems use motion model to approximate the geometry of moving object boundaries. Motion model can be relieved from describing discontinuities in the underlying motion field, by employing motion hint that exploits the spatial structure of reference frames to infer appropriate boundaries for the future ones. However, estimation of highly accurate motion hint is computationally demanding, in particular for high resolution video sequences. Leveraging on the advantages of homogeneous motion discovery oriented prediction, in this paper, we propose to tune the intra-domain motion uniformity for B-frames as per the frame's reference utility. Experimental results show an improved bit rate savings compared to the approach where no such selective tuning is enforced. Ashek Ahmmed, Aous Thabit Naman, David S. Taubman |
PCS | 2 |
| 2018 | Responsive high throughput congestion control for interactive applications over SDN-enabled networks
Aous Thabit Naman, Yu Wang 0131, Hassan Habibi Gharakheili, Vijay Sivaraman, David S. Taubman |
Comput. Networks | 1 |
| 2017 | Lifting-based Illumination Adaptive Transform (LIAT) using mesh-based illumination modellingabstractState-of-the-art video coding techniques employ block-based illumination compensation to improve coding efficiency. In this work, we propose a Lifting-based Illumination Adaptive Transform (LIAT) to exploit temporal redundancy among frames that have illumination variations, such as the frames of low frame rate video or multi-view video. LIAT employs a mesh-based spatially affine model to represent illumination variations between two frames. In LIAT, transformed frames are jointly compressed, together with illumination information, into a layered rate-distortion optimal codestream, using the JPEG2000 format. We show that the LIAT framework significantly improves compression efficiency of temporal subband transforms for both predictive and more general transforms with predict and update steps. Maryam Haghighat, Reji Mathew, Aous Thabit Naman, David S. Taubman |
ICIP | 3 |
| 2016 | Optimization and compression of geometry discontinuities for graph-based representation of piecewise smooth mediaabstractEarlier research has shown the efficacy of using geometry discontinuities, encoded using “breakpoints,” in improving the coding efficiency of piecewise smooth media, such as depth maps and motion flows. This work proposes a new structure for encoding these breakpoints that is more suited for piecewise affine media, such as affine motion flows. Here, we choose to employ belief propagation over a graphical model to discover a set of breakpoints that is rate-distortion optimal. Similar to earlier works, the discovered breakpoints are employed in a breakpoint adaptive (BPA) wavelet decomposition of the media under consideration; the resulting BPA wavelet coefficients and the proposed “affine” breakpoints are then compressed using the JPEG2000 format, in a way that provides spatial and quality scalability as well as accessibility. We show that the coding cost of affine breakpoints is moderate, and that, for coding motion flows, the coding efficiency of the proposed approach is comparable to H.264. Aous Thabit Naman, David S. Taubman, Reji Mathew |
ICIP | 1 |
| 2016 | Homogeneous motion discovery oriented reference frame for high efficiency video codingabstractTraditional video coding uses the motion model to approximate geometric boundaries of moving objects where motion discontinuities occur. Motion hints based inter-frame prediction paradigm moves away from this redundant approach and employs an innovative framework consisting of motion hint fields that are continuous and invertible, at least, over their respective domains. However, estimation of motion hint is computationally demanding, in particular for high resolution video sequences. In this paper, we propose to discover motion models and their associated masks over the current frame and then use these models and masks to form a prediction of the current frame. The prediction process is computationally simpler and experimental results show that a savings in bit rate of 2.3% is achievable over standalone HEVC if this predicted frame is used as an additional reference frame. Ashek Ahmmed, David S. Taubman, Aous Thabit Naman, Mark R. Pickering |
PCS | 3 |
| 2016 | Optimized decoding of JPEG images based on generalized graph LaplaciansabstractWe address the problem of optimizing the decoding of JPEG-compressed images, employing an approach based on the “generalized” graph Laplacian, a higher-order generalization of the usual graph Laplacian. The optimal decoding problem is formulated as a non-smooth but convex problem over a graph, and solved via the alternating directions method of multipliers. While similar, graph-based optimized decoding techniques exist, the use of our generalized graph Laplacian enables better recovery of the original smooth image from the transform coefficients, especially those coded at low bit-rates. Experimental results highlight the performance of our proposed approach in terms of reconstructed distortion and visual quality. Comparisons with the method based on the usual graph Laplacian and other state-of-the-art methods are also given. Sean I. Young, Aous Thabit Naman, Reji Mathew, David S. Taubman |
PCS | 2 |
| 2016 | Motion Estimation Based on Mutual Information and Adaptive Multi-Scale ThresholdingabstractThis paper proposes a new method of calculating a matching metric for motion estimation. The proposed method splits the information in the source images into multiple scale and orientation subbands, reduces the subband values to a binary representation via an adaptive thresholding algorithm, and uses mutual information to model the similarity of corresponding square windows in each image. A moving window strategy is applied to recover a dense estimated motion field whose properties are explored. The proposed matching metric is a sum of mutual information scores across space, scale, and orientation. This facilitates the exploitation of information diversity in the source images. Experimental comparisons are performed amongst several related approaches, revealing that the proposed matching metric is better able to exploit information diversity, generating more accurate motion fields. Rui Xu 0030, David S. Taubman, Aous Thabit Naman |
IEEE Trans. Image Process. | 3 |
| 2015 | Motion hints mode for macroblock coding in bi-predictive slicesabstractRecent advances in motion modelling have largely focused on careful partitioning of motion blocks in the vicinity of object boundaries. The need for such fine partitioning can be avoided by using motion hints which provide a global description of motion over specific domains. Experimental results show that, with a hybrid setting, more than 50% of the motion discontinuity macroblocks are coded using the motion hints mode in low bit rate cases. The use of this mode leads to a gain of prediction PSNR of 1.11 dB, or equivalently 17.05% savings in bit rate, when compared to the H.264/AVC reference and considering both low and high bit rate applications. Ashek Ahmmed, Md. Jahangir Alam 0005, Aous Thabit Naman, Mark R. Pickering, David S. Taubman |
PCS | 3 |
| 2015 | Motion estimation with accurate boundariesabstractThis paper investigates several techniques that increase the accuracy of motion boundaries in estimated motion fields of a local dense estimation scheme. In particular, we examine two matching metrics, one is MSE in the image domain and the other one is a recently proposed multiresolution metric that has been shown to produce more accurate motion boundaries. We also examine several different edge-preserving filters. The edge-aware moving average filter, proposed in this paper, takes an input image and the result of an edge detection algorithm, and outputs an image that is smooth except at the detected edges. Compared to the adoption of edge-preserving filters, we find that matching metrics play a more important role in estimating accurate and compressible motion fields. Nevertheless, the proposed filter may provide further improvements in the accuracy of the motion boundaries. These findings can be very useful for a number of recently proposed scalable interactive video coding schemes. Rui Xu 0030, Aous Thabit Naman, Reji Mathew, Dominic Rüfenacht, David S. Taubman |
PCS | 2 |
| 2014 | Overlapping motion hints with polynomial motion for video communicationabstractIn a recent work, we propose the use of motion hints for communicating motion. A motion hint describes motion that is accurate (describes the actual motion) for only a region inside a domain associated with that motion hint; it is the job of the client or decoder to decide the exact region of applicability (ROA). The motion described by a motion hint is invertible and global; that is, it allows the prediction of the ROA associated with a motion hint from any frame that has that hint. Motion hints are applicable to closed-loop prediction, but they are more useful in open-loop prediction scenarios, such as remote browsing of surveillance footage, communicated by a JPIP server, which is the focus of this work. This work proposes a probabilistic multi-scale framework to identifying the ROA; the framework is applicable to multiple overlapping motion hints. The proposed approach is localized (and therefore amenable to parallel processing) and robust to noise, quantization, and changes in contrast. We show that motion hints can be used for real video sequences, and we also present results for the case of three overlapping motion hints (one background and two overlapping foregrounds). Aous Thabit Naman, David S. Taubman, Rui Xu 0030 |
ICIP | 1 |
| 2014 | Block motion matching on directional subbands with interband suppressionabstractWe propose a novel matching metric for dense block-based true-motion estimation. Block-based matching scores are calculated on directional detail bands of a steerable pyramid after conversion to a 2-bit representation. The proposed non-linear transform involves a novel inter-band suppression mechanism so that matching scores can be accumulated across resolutions and directions. We show that edge features of different strength that occupy similar locations can be isolated in different directional bands, after which the 2-bit transformation effectively equalises their contrast. This considerably improves the robustness of motion estimation procedure when compared to other matching metrics, including our prior work and the commonly used MSE. Rui Xu 0030, David S. Taubman, Aous Thabit Naman |
ICIP | 3 |
| 2014 | Flexible Synthesis of Video Frames Based on Motion HintsabstractIn this paper, we propose the use of "motion hints" to produce interframe predictions. A motion hint is a loose and global description of motion that can be communicated using metadata; it describes a continuous and invertible motion model over multiple frames, spatially overlapping other motion hints. A motion hint provides a reasonably accurate description of motion but only a loose description of where it is applicable; it is the task of the client to identify the exact locations where this motion model is applicable. The focus of this paper is a probabilistic multiscale approach to identifying these locations of applicability; the method is robust to noise, quantization, and contrast changes. The proposed approach employs the Laplacian pyramid; it generates motion hint probabilities from observations at each scale of the pyramid. These probabilities are then combined across the scales of the pyramid starting from the coarsest scale. The computational cost of the approach is reasonable, and only the neighborhood of a pixel is employed to determine a motion hint probability, which makes parallel implementation feasible. This paper also elaborates on how motion hint probabilities are exploited in generating interframe predictions. The scheme of this paper is applicable to closed-loop prediction, but it is more useful in open-loop prediction scenarios, such as using prediction in conjunction with remote browsing of surveillance footage, communicated by a JPEG2000 Interactive Protocol (JPIP) server. We show that the interframe predictions obtained using the proposed approach are good both visually and in terms of PSNR. Aous Thabit Naman, David S. Taubman |
IEEE Trans. Image Process. | 1 |
| 2014 | Nonlinear Transform for Robust Dense Block-Based Motion EstimationabstractWe present a noniterative multiresolution motion estimation strategy, involving block-based comparisons in each detail band of a Laplacian pyramid. A novel matching score is developed and analyzed. The proposed matching score is based on a class of nonlinear transformations of Laplacian detail bands, yielding 1-bit or 2-bit representations. The matching score is evaluated in a dense full-search motion estimation setting, with synthetic video frames and an optical flow data set. Together with a strategy for combining the matching scores across resolutions, the proposed method is shown to produce smoother and more robust estimates than mean square error (MSE) in each detail band and combined. It tolerates more of nontranslational motion, such as rotation, validating the analysis, while providing much better localization of the motion discontinuities. We also provide an efficient implementation of the motion estimation strategy and show that the computational complexity of the approach is closely related to the traditional MSE block-based full-search motion estimation procedure. Rui Xu 0030, David S. Taubman, Aous Thabit Naman |
IEEE Trans. Image Process. | 3 |
| 2013 | A soft measure for identifying structure from randomness in imagesabstractThis paper presents a novel measure for identifying strong structure features, such as edges, from randomness, such as regions predominated by noise, within an image. The proposed structural measure is localized in space and scale; for a given scale, it gives values close to one in the vicinity of strong structures and close to zero in regions predominated by noise. The proposed structural measure is a primitive operation that can be used in a wide variety of image analysis techniques to identify regions which has structure; for example, motion estimation is more meaningful in structured regions than in regions filled with noise. The first innovation in this work is in converting an image into a ternary feature map that are rather resistant to noise and changes in illumination. The second is the structural measure, which is derived from the degree of non-uniformity amongst the magnitudes of the DFT coefficients obtained over a small window within the ternary maps. In this work, we show that the proposed structural measure is robust and gives a good indication of the strength of structure when compared to alternate strategies; moreover, we show that the computational cost of the proposed structural measure is reasonable. Aous Thabit Naman, David S. Taubman |
ICIP | 1 |
| 2013 | Inter-frame prediction using motion hintsabstractWe recently proposed a novel approach that employs motion hints for inter-frame prediction. Motion hints are a loose and global description of motion communicated as metadata; they specify motion but they leave it to the client/decoder to find the exact locations where motion is applicable. This work proposes a multi-scale approach for identifying these exact locations, which are then used with the available reference frames to generate an inter-frame prediction. The proposed approach is localized and robust to noise and illumination changes. The scheme of this work is applicable to close-loop prediction, but it is more useful in open-loop prediction scenarios, such as using prediction in conjunction with remote browsing of surveillance footage, communicated by a JPIP server. We show that, with reasonably accurate motion, it is possible to produce good inter-frame predictions visually and in terms of PSNR. Aous Thabit Naman, Rui Xu 0030, David S. Taubman |
ICIP | 1 |
| 2013 | Motion segmentation initialization strategies for bi-directional inter-frame predictionabstractExperimental results and the latest standards have proved that segmentation based video coding systems can outperform the traditional block-based video coding systems. However, this approach requires the simultaneous estimation of both the shape and motion of moving objects in a video scene. In most of the cases neither the shape nor the motion are known initially. Another critical aspect of this tightly-coupled relationship is that inaccurate motion estimation may cause poor segmentation and erroneous segmentation may negatively impact motion estimation. While some of the existing approaches require user intervention and some use clues such as depth, colour or occlusion to separate the foreground from the background, we propose to use motion reliability information for this purpose. This is because the ingredients necessary for the calculation of motion reliability are the by-product of block-based motion estimation and compensation between the reference frames. Therefore, they require very little or no increase in the computational overhead. In this paper, we explore several motion segmentation initialization strategies based on motion reliability. The performances of these initialization approaches are investigated, in terms of the PSNR, for the predicted inter-frames. Ashek Ahmmed, Rui Xu 0030, Aous Thabit Naman, Md. Jahangir Alam 0005, Mark R. Pickering, David S. Taubman |
MMSP | 3 |
| 2013 | Motion hints based inter-frame prediction for hybrid video codingabstractExperimental results and the latest standards have proved video coding systems with the ability to adapt the size and shape of the motion estimation area to the objects in the scene can outperform the traditional block-based video coding systems. In this paper, a segmentation-based coding strategy that employs bi-directional motion hints for interframe prediction is proposed. The appealing thing about motion hints is that they are continuous and invertible, even though the observed motion field for a frame will be discontinuous and non-invertible. The proposed scheme outperforms the rate-distortion performance of H.264/AVC reference by 1.1 dB and a bit rebate of 26.6% is achieved. Ashek Ahmmed, Md. Jahangir Alam 0005, Mark R. Pickering, Rui Xu 0030, Aous Thabit Naman, David S. Taubman |
PCS | 5 |
| 2011 | Efficient communication of video using metadataabstractIn traditional video coding schemes, motion information is tightly coupled to the prediction strategy. In this preliminary work, we depart from this model by utilizing metadata to convey motion information to the client; in particular, metadata conveys crude boundaries of objects together with motion information for these objects. Here, we are interested in applications where metadata itself carries semantics that the client is interested in, such as tracking information in surveillance applications. To keep things simple, we focus on the case where we have a single object to track. Therefore, we model each frame as a background region with a foreground region/object, enclosing each region by a quadrilateral that identifies it. The foreground quadrilateral does not follow the exact boundaries of the foreground object; it leaves the task of identifying these boundaries to the client. The advantages of metadata is that it provides a global representation of motion, which allows predicting a given object from potentially all the frames that contain that object. The approach is applicable in fully open loop systems such as in the case of the JPEG2000-Based Scalable Interactive Video (JSIV) paradigm. In this work, we present the concepts behind the proposed approach and detail the modifications introduced to the JSIV server and client policies, presenting some promising preliminary results. Aous Thabit Naman, Duncan Edwards, David S. Taubman |
ICIP | 1 |
| 2011 | JPEG2000-Based Scalable Interactive Video (JSIV)abstractWe propose a novel paradigm for interactive video streaming and we coin the term JPEG2000-based scalable interactive video (JSIV) for it. JSIV utilizes JPEG2000 to independently compress the original video sequence frames and provide for quality and spatial resolution scalability. To exploit interframe redundancy, JSIV utilizes prediction and conditional replenishment of code-blocks aided by a server policy that optimally selects the number of quality layer for each code-block transmitted and a client policy that makes most of the received (distorted) frames. It is also possible for JSIV to employ motion compensation; however, we leave this topic to future work. To optimally solve the server transmission problem, a Lagrangian-style rate-distortion optimization procedure is employed. In JSIV, a wide variety of frame prediction arrangements can be employed including hierarchical B-frames of the scalable video coding (SVC) extension of the H.264/AVC standard. JSIV provides considerably better interactivity compared to existing schemes and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. Experimental results for surveillance footage, which does not suffer from the absence of motion compensation, show that JSIV's performance is comparable to that of SVC in some usage scenarios while JSIV performs better in others. Aous Thabit Naman, David S. Taubman |
IEEE Trans. Image Process. | 1 |
| 2011 | JPEG2000-Based Scalable Interactive Video (JSIV) With Motion CompensationabstractIn a recent work, the authors proposed a novel paradigm for interactive video streaming and coined the term JPEG2000-Based Scalable Interactive Video (JSIV) for it. In this work, we investigate JSIV when motion compensation is employed to improve prediction, something that was intentionally left out in our earlier treatment. JSIV relies on three concepts: storing the video sequence as independent JPEG2000 frames to provide quality and spatial resolution scalability, prediction and conditional replenishment of code-blocks to exploit inter-frame redundancy, and loosely coupled server and client policies in which a server optimally selects the number of quality layers for each code-block transmitted and a client makes the most of the received (distorted) frames. In JSIV, the server transmission problem is optimally solved using Lagrangian-style rate-distortion optimization. The flexibility of JSIV enables us to employ a wide variety of frame prediction arrangements, including hierarchical B-frames. JSIV provides considerably better interactivity compared with existing schemes and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. Experimental results show that JSIV's performance is inferior to that of SVC in conventional streaming applications while JSIV performs better in interactive browsing applications. Aous Thabit Naman, David S. Taubman |
IEEE Trans. Image Process. | 1 |
| 2010 | Predictor selection using quantization intervals in JPEG2000-Based Scalable Interactive Video (JSIV)abstractThe authors have recently introduced the JPEG2000-Based Scalable Interactive Video (JSIV) paradigm. JSIV relies on JPEG2000 format for providing scalability and accessibility, and on motion compensation and conditional replenishment to exploit temporal redundancy. JSIV can provide considerably better interactivity compared to existing video streaming practices, and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. This work extends our previous work by providing server and client policies that can exploit the client's knowledge about the quantization intervals of received samples in selecting a favorable predictor in dyadic hierarchical B-frame arrangement that does not employ motion compensation. We also demonstrate the flexibility of the JSIV paradigm by showing an improved client policy working with a non-improved server policy without any negative impact on reconstructed video. Aous Thabit Naman, David S. Taubman |
ICIP | 1 |
| 2009 | Rate-distortion optimized JPEG2000-based scalable interactive video (JSIV) with motion and quantization bin side-informationabstractThe authors have recently proposed a paradigm that can potentially provide for considerably better interactivity compared to existing practices and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. The proposed paradigm relies on JPEG2000 format for providing scalability, flexibility, and accessibility; and on transmitting a server-optimized selection of code-blocks and motion side-information. Motion compensation and conditional replenishment are employed to reduce needed bandwidth. This work extends the previous work by providing server and client policies that allow for a realistic implementation and by introducing the use of coarsely quantized code-blocks in improving prediction. This work introduces the concepts, formulates the policies and optimization problems, proposes solutions, and compares the performance to alternate strategies. Aous Thabit Naman, David S. Taubman |
ICIP | 1 |
| 2008 | Rate-distortion optimized delivery of JPEG2000 compressed video with hierarchical motion side informationabstractStreaming video as a sequence of JPEG2000 images provides the scalability, flexibility, and accessibility at a wide range of bit-rates that is lacking from the current motion-compensated predictive video coding standards; however, streaming this sequence requires considerably more bandwidth. The authors have recently proposed a novel approach that reduces the required bandwidth; this approach uses motion compensation and conditional replenishment of the JPEG2000 code-blocks, aided by server-optimized selection of these code-blocks. This work extends the previous work to the case of hierarchical arrangement of frames, similar to the hierarchical B-frames of the SVC scalable video coding extension of the H.264/AVC standard. We employ a Lagrangian-style rate-distortion optimization procedure to the server transmission problem and compare the performance to that of streaming individual frames and also to that of predictive video coding. The proposed approach can serve a diverse range of client requirements and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. This paper introduces the concepts, formulates the optimization problem, proposes a solution, and compares the performance to alternate strategies. Aous Thabit Naman, David S. Taubman |
ICIP | 1 |
| 2008 | Distortion estimation for optimized delivery of JPEG2000 compressed video with motionabstractA JPEG2000 compressed video sequence can provide better support for scalability, flexibility, and accessibility at a wider range of bit-rates than the current motion-compensated predictive video coding standards; however, it requires considerably more bandwidth to stream. The authors have recently proposed a novel approach that reduces the required bandwidth; this approach uses motion compensation and conditional replenishment of JPEG2000 code-blocks, aided by server-optimized selection of these code-blocks. The proposed approach can serve a diverse range of client requirements and can adapt immediately to interactive changes in client interests, such as forward or backward playback and zooming into individual frames. This work extends the previous work by approximating the distortion associated with the decisions made by the server without the need to recreate the actual video sequence at the server. The proposed distortion estimation algorithm is general and can be applied to various frames arrangements. Here, we choose to employ it in a hierarchical arrangement of frames, similar to the hierarchical B-frames of the SVC scalable video coding extension of the H.264/AVC standard. We employ a Lagrangian-style rate-distortion optimization procedure to the server transmission problem and compare the performance of both distortion estimation and exact distortion calculation cases against streaming individual frames and SVC. Results obtained suggest that the distortion estimation algorithm considerably reduces the amount of calculation needed by the server without enormously degrading the performance compared to the exact distortion calculation case. This work introduces the concepts, formulates the estimation and optimization problems, proposes a solution, and compares the performance to alternate strategies. Aous Thabit Naman, David S. Taubman |
MMSP | 1 |
| 2007 | A Novel Paradigm for Optimized Scalable Video Transmission Based on JPEG2000 with MotionabstractA novel paradigm for optimized scalable video streaming is presented. The paradigm proposes transmission of motion vectors and selected code-blocks of the JPEG 2000 representation of each new frame, instead of frame differences as in existing methods. These code-blocks are selected to achieve the highest MSE. This paradigm overcomes the flexibility and accessibility limitations imposed by predictive motion compensated video by relying on the JPEG 2000 stream features for spatial scalability and on motion compensation and server-optimized conditional replenishment for temporal redundancy reduction. It is expected that real-time and interactive applications, such as teleconferencing and surveillance, would benefit most from this paradigm. This paper introduces the paradigm, formulates an optimization procedure for one simple case where it can be applied and compares its performance with alternate strategies. Aous Thabit Naman, David S. Taubman |
ICIP (5) | 1 |