VLDB 2026 Research / reviewers in the wild / expert
Debargha Mukherjee
dblp:35/4322
· DBLP profile ↗
92ranked-venue papers
29as first author
17since 2021 · last 2026
0000-0002-9380-7377ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 88 · 28 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive AV2 In-loop Filtering via Guided Neural Model with Vectorized Quantization
Kequan Mao, Dandan Ding, Urvang Joshi, Debargha Mukherjee |
ISCAS | 4 |
| 2025 | Super Resolution-Based Video Coding via Lightweight Implicit Neural ModelingabstractThe super-resolution (SR)-based coding tool is widely employed in modern video coding standards. By encoding video frames at a reduced resolution and then restoring them to their original resolution during the in-loop filtering stage, this tool helps to further reduce the bitrate and improve the coding performance. Current video coding standards typically devise rule-based SR methods in their codecs, compromising the coding efficiency to maintain low computational complexity. As deep neural network (DNN)-based SR methods are proving more effective than rule-based approaches, this paper proposes integrating the neural SR into video codecs to enhance coding performance while minimizing the computational cost. To this end, we propose a Lightweight Implicit Neural Model (LIM). Specifically, our LIM, consisting of Lightweight Feature Aggregation Network (LFANet) and Coordinate Upsampling Network Based on B-spline Representation (CURNet), is developed to support SR-based coding at an arbitrary scale. We exemplify the proposed method on the ongoing AVM reference software and conduct extensive experiments to demonstrate its effectiveness. Compared with anchored AVM, our method improves the BD-Rate by 5.52%, which significantly outperforms state-of-the-art works. Meanwhile, its computational complexity is much lower than others, having only 22.5k parameters and 18.7k FLOPs/pixel complexity, which is attractive to real-world applications. Xianlu Bian, Dandan Ding, Urvang Joshi, Debargha Mukherjee |
DCC | 5 |
| 2025 | Joint Optimization of Primary and Secondary Transforms Using Rate-Distortion Optimized Transform DesignabstractData-dependent transforms are increasingly being incorporated into next-generation video coding systems such as AVM, a codec under development by the Alliance for Open Media (AOM), and VVC. To circumvent the computational complexities associated with implementing non-separable data-dependent transforms, combinations of separable primary transforms and non-separable secondary transforms have been studied and integrated into video coding standards. These codecs often utilize rate-distortion optimized transforms (RDOT) to ensure that the new transforms complement existing transforms like the DCT and the ADST. In this work, we propose an optimization framework for jointly designing primary and secondary transforms from data through a rate-distortion optimized clustering. Primary transforms are assumed to follow a path-graph model, while secondary transforms are non-separable. We empirically evaluate our proposed approach using AVM residual data and demonstrate that 1) the joint clustering method achieves lower total RD cost in the RDOT design framework, and 2) jointly optimized separable path-graph transforms (SPGT) provide better coding efficiency compared to separable KLTs obtained from the same data. Darukeesan Pakiyarajah, Eduardo Pavez, Antonio Ortega, Debargha Mukherjee, Onur G. Guleryuz, Keng-Shih Lu, Kruthika Koratti Sivakumar |
ICIP | 4 |
| 2025 | AVM In-loop Filtering using Adaptive Hardware-friendly Neural Networks
Urvang Joshi, Akshaya Purohit, Debargha Mukherjee, Shan Li 0001, Randy Hsin, In Suk Chong |
PCS | 3 |
| 2024 | Standard Compatible Efficient Video Coding with Jointly Optimized Neural WrappersabstractWe present a standard-compatible video coding scheme with end-to-end optimized neural wrapper over standard video codecs that achieves significant rate-distortion (R-D) performance gains and is still efficient in decoding. We train a pair of pre- and post-processor using a differential JPEG proxy. The pre-processor applies a learned transform to the video and downsamples the video by a factor of 2. It generates a bottleneck video to be coded by a standard codec as a YUV sequence. The post-processor takes the decoded bottleneck video, does the inverse transform, and upsamples it to the original resolution. We follow the design in [1] , where we configure downsample using a layer of strided convolution. We optimize the post-processor for efficiency by replacing convolutions with kernel size larger than 1×1 to depth-wise convolutions [2] . Yueyu Hu, Onur G. Guleryuz, Debargha Mukherjee, Yao Wang 0001 |
DCC | 4 |
| 2024 | Standard Compliant Video Coding Using Low Complexity, Switchable Neural WrappersabstractThe proliferation of high resolution videos posts great storage and bandwidth pressure on cloud video services, driving the development of next-generation video codecs. Despite great progress made in neural video coding, existing approaches are still far from economical deployment considering the complexity and rate-distortion performance tradeoff. To clear the roadblocks for neural video coding, in this paper we propose a new framework featuring standard compatibility, high performance, and low decoding complexity. We employ a set of jointly optimized neural pre and post-processors, wrapping a standard video codec, to encode videos at different resolutions. The rate-distorion optimal downsampling ratio is signaled to the decoder at the per-sequence level for each target rate. We design a low complexity neural post-processor architecture that can handle different upsampling ratios. The change of resolution exploits the spatial redundancy in high-resolution videos, while the neural wrapper further achieves rate-distortion performance improvement through end-to-end optimization with a codec proxy. Our light-weight post-processor architecture has a complexity of 516 MACs / pixel, and achieves 9.3% BD-Rate reduction over VVC on the UVG dataset, and $6.4 \%$ on AOM CTC Class A1. Our approach has the potential to further advance the performance of the latest video coding standards using neural processing with minimal added complexity. Yueyu Hu, Onur G. Guleryuz, Debargha Mukherjee, Yao Wang 0001 |
ICIP | 4 |
| 2024 | ELIM: Extremely Low-Complexity Implicit Neural Model for Super Resolution-Based CodingabstractThe super-resolution (SR)-based coding, which encodes a frame at a reduced resolution to achieve a lower bitrate, is a prevalent tool used in modern video coding standards. Accordingly, the low-resolution frame is restored to the full resolution at the reconstruction stage for subsequent reference. Therefore, the resolution restoration algorithm significantly affects the coding performance. This paper devises a highly efficient and extremely low complexity implicit neural model (ELIM) for SR-based encoding to support arbitrary scale factors. Specifically, ELIM consists of two stages: Feature Aggregation and Coordinate Upsampling. In Feature Aggregation, we embed a simplified attention block to the U-Net style framework to collect valuable information while reducing computational complexity through downsampling. In Coordinate Upsampling, in addition to the extracted content features, information including coordinate relative location and pixel cell size is fused to achieve better performance. We exemplify ELIM on the AV2 codec (the next generation of AV1). Extensive experiments demonstrate its superior performance: it achieves 4.17% BD-Rate gains over the anchor AV2 reference software with only 5,185 flops/pixel, significantly surpassing existing methods. The low complexity of ELIM is attractive to real applications. Dandan Ding, Urvang Joshi, Debargha Mukherjee |
PCS | 5 |
| 2023 | Neural Adaptive Loop Filtering for Video Coding: Exploring Multi-Hypothesis Sample RefinementabstractAdaptive loop filtering (ALF) is extensively investigated for lossy video coding to mitigate compression noise. Numerous learning-based ALFs have emerged recently and improved the coding efficiency significantly through the use of complexity-intensive, large-scale models trained on excessive samples, making it impractical for real-life applications. By contrast, lightweight, small-scale ALF models cannot promise convincing performance and model generalization. In principle, the ALF estimates the sample distortion for restoration. Instead of directly approximating the distortion as in existing solutions, we reformulate it as a Multi-hypothesis Sample Refinement (MSR) problem. To this end, we first generate multiple distortion hypotheses through a deep neural network (DNN) model. Then, these hypotheses are linearly superimposed to approximate the final distortion through the minimization of mean square error (MMSE) between the filtered reconstruction and its original, uncompressed input. Finally, the linear superimposition coefficients are explicitly signaled in the compressed bitstream. As seen, the superimposition coefficients inherently generalize the MSR to various content. And using DNNs to generate distortion hypotheses essentially models the spatial priors of a local block and its underlying compression error distribution. As a result, the MSR using a small-scale convolutional neural network (CNN) model with only 5k parameters and 5 KMACs/pixel achieves 4.35% (Intra) and 2.49% (Inter) BD-Rate (Bjøntegaard Delta Rate) gains over the AV1 anchor. Ablation studies further demonstrate that the MSR can be generalized to diverse small-scale network structures, different standards (e.g., H.265/HEVC and H.266/VVC), and diverse video content (e.g., screen videos). Dandan Ding, Guangkun Zhen, Debargha Mukherjee, Urvang Joshi, Zhan Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Non-Separable Filtering with Side-Information and Contextually-Designed Filters for Next Generation Video CodecsabstractWe propose two new in-loop filtering tools to augment the loop restoration process in AV1. Targeting challenging future compression scenarios that will be faced by the emerging AV2 standard, the proposed tools are designed to improve performance at both high and low bit-rates. Our non-separable Wiener filtering proposal aims to increase quality especially over directional features and textures in the decoded picture with the help of side-information at higher rates. The proposed pixel-adaptive filter complements at lower rates by improving performance without side-information relying solely on finely characterized pixel contexts. Simulation results show the rate-distortion efficacy of the proposed tools. Onur G. Guleryuz, Debargha Mukherjee, Yue Chen 0040, Keng-Shih Lu, Urvang Joshi |
ICIP | 2 |
| 2022 | Switchable CNN-Based Same-Resolution and Super-Resolution In-Loop Restoration for Next Generation Video CodecsabstractWe present a common framework for in-loop same-resolution and super-resolution restoration for incorporation into a next-generation video codec. Building on the in-loop filtering pipeline in the AV1 video codec from the Alliance for Open Media (AOM), we first enhance it to support symmetric spatial down-up scaling with better down and upscaling filters, followed by adding switchable frame-level CNNs to restore the reconstructed frames. Furthermore, the architectures for the CNNs used are constrained to be very simple with a relatively small number of parameters and multiply-add operations per decoded pixel to make them practically feasible in hardware. Preliminary results are presented on test sets being used by AOM for their ongoing next-generation video codec (AV2) development effort. Urvang Joshi, Yue Chen 0040, Debargha Mukherjee, Onur G. Guleryuz, Shan Li 0001, In Suk Chong |
ICIP | 3 |
| 2022 | Intra Prediction of Regular and Near-Regular Textures Via Graph-Based InpaintingabstractIntra prediction is an important technique to improve coding efficiency by exploiting the spatial redundancy present in typical video sequences. In video coding standards such as H.264/AVC, HEVC and VVC, directional predictors are utilized to generate prediction along a single direction within a block to be coded. However, these predictors fail to generate an accurate prediction when the block contains complex patterns such as periodic textures. In this paper, we propose a graph-based inpainting method that can handle both regular and near-regular textures. The proposed inpainting method utilizes a total variation model associated with the Laplacian matrix of a graph, whose edge weights are a function of pixel patch distance. We evaluate the performance of our proposed method as an additional prediction mode combined with the H.264/AVC coding standard. Experimental results show that the proposed method can significantly outperform H.264/AVC predictors in areas with high frequency periodic patterns. Wen-Yang Lu, Eduardo Pavez, Antonio Ortega, Debargha Mukherjee, Onur G. Guleryuz, Keng-Shih Lu |
ICIP | 4 |
| 2022 | Quadtree-based Guided CNN for AV1 In-loop FilteringabstractRecently, learning-based in-loop filtering has attracted lots of attention. State-of-the-art works generally deploy computationally expensive, large-scale neural networks, which is unfriendly to practical applications. Besides, since these models are generally pre-trained using a limited dataset and applied to various videos, they may fail in video contents excluded in the training dataset. To address these issues, this paper develops a Guide CNN in-loop filtering framework to obtain the restored signal. Our basic idea is to construct a subspace and use the projection of the original signal into this subspace to approximate the original signal itself. Specifically, we employ CNN to transform the degraded signal into M subsignals to construct the optimal subspace since the training of CNN is essentially an optimization procedure. Furthermore, unlike existing CNN models that process all blocks uniformly, our method leverages a quadtree structure to implement the Guided CNN through R-D optimization. As such, the best partition to Guided CNN can be determined. We exemplify the proposed method in AV1 codec. Experimental results show that the Guided CNN framework achieves 2.19% and 1.31% BD-Rate gains over the AV1 anchor in intra and inter coding mode, respectively, while the normal CNN achieves only 1.64% and 1.04%. Gongchun Ding, Dandan Ding, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
ICIP | 4 |
| 2021 | Improved Chroma from Luma Prediction in AV1 Based On Virtual Chroma Block GenerationabstractChroma from Luma (CfL) prediction is an efficient coding tool in AV1 which builds chroma prediction by implementing a linear model on luma pixels. To avoid transmission, the off-set factor in the linear model is set to the average of neighboring chroma pixels. An improved CfL algorithm is proposed to derive the offset factor based on a virtual chroma block. Such a block is constructed by using the chroma of the matched pixel which is determined according to the luma difference of neighboring regions and the co-located luma. Compared with libaom, the proposed CfL algorithm provides 0.31% and 0.15% weighted PSNR BD-rate saving under AI and RA configuration, respectively. Experimental results show that over 1.00% and 0.80% BD-rate saving can be achieved for chroma components under AI and RA configuration. With the proposed algorithm, the percentage of pixels with CfL as the optimal coding mode is increased. Junyan Huo, Menglin Zhang, Wenhan Qiao, Fuzheng Yang 0001, Hui Su, Debargha Mukherjee |
ICME | 6 |
| 2021 | Study On Coding Tools Beyond AV1abstractThe Alliance for Open Media has recently initiated coding tool exploration activities towards the next-generation video coding beyond AV1. With this regard, this paper presents a package of coding tools that have been investigated, implemented and tested on top of the codebase, known as libaom, which is used for the exploration of next-generation video compression tools. The proposed tools cover several technical areas based on a traditional hybrid video coding structure, including block partitioning, prediction, transform and loop filtering. The proposed coding tools are integrated as a package, and a combined coding gain over AV1 is demonstrated in this paper. Furthermore, to better understand the behavior of each tool, besides the combined coding gain, the tool-on and tool-off tests are also simulated and reported for each individual coding tool. Experimental results show that, compared to libaom, the proposed methods achieve an average 8.0% (up to 22.0%) overall BD-rate reduction for All Intra coding configuration a wide range of image and video content. Xin Zhao 0003, Madhu Peringassery Krishnan, Yixin Du, Shan Liu 0001, Debargha Mukherjee, Yaowu Xu, Adrian Grange |
ICME | 6 |
| 2021 | Generalized Optical Flow based Motion Vector Refinement in AV1abstractBi-directional optical flow (BDOF) is a coding tool that has been recently adopted into Versatile Video Coding (VVC) standard. In BDOF, gradients of prediction samples are exploited to refine motion vector (MV) per subblock within a prediction block, and thus enhance inter prediction quality in a two-pass framework. In this work, we extend the concept of BDOF to a more general compound prediction framework and integrated this framework into the AV1 codec. In particular, we support MV refinement not only in bi-directional compound prediction, but also in uni-directional prediction, where the two reference blocks are both from the past or both from the future. Furthermore, two reference blocks are allowed to have arbitrary temporal distances to the current block. For implementation on the AV1 codec, five additional inter compound modes have been added, and the proposed method is performed on top of those modes. Experimental results show more than 2.8% BD rate saving on the Google test set. Keng-Shih Lu, Sarah Parker, Debargha Mukherjee |
PCS | 3 |
| 2021 | Block-based Learned Image Coding with Convolutional Autoencoder and Intra-Prediction Aided Entropy CodingabstractRecent works on learned image coding using autoencoder models have achieved promising results in rate-distortion performance. Typically, an autoencoder is used to transform an image into a latent tensor, which is then quantized and entropy coded. Based on a work by Ballé et al., we adapted the autoencoder with a hyperprior model to code images in a block-based approach. When the autoencoder model is directly applied to code small image blocks, spatial redundancy in the larger image cannot be fully utilized, resulting in a decrease in ratedistortion performance. We propose a method to utilize border information in the entropy coding of latent and hyper-latent tensors, which has achieved promising results. We show that using intra-prediction to help entropy coding is more effective than applying a convolutional autoencoder with hyper priors to intra-prediction residual blocks. Zhongzheng Yuan, Debargha Mukherjee, Balu Adsumilli, Yao Wang 0001 |
PCS | 3 |
| 2021 | A Technical Overview of AV1abstractThe AV1 video compression format is developed by the Alliance for Open Media consortium. It achieves more than a 30% reduction in bit rate compared to its predecessor VP9 for the same decoded video quality. This article provides a technical overview of the AV1 codec design that enables the compression performance gains with considerations for hardware feasibility. Jingning Han, Bohan Li 0006, Debargha Mukherjee, Ching-Han Chiang, Adrian Grange, Hui Su, Sarah Parker, Sai Deng, Urvang Joshi, Yue Chen 0040, Yunqing Wang, Paul Wilkins, Yaowu Xu, Jim Bankoski |
Proc. IEEE | 3 |
| 2020 | Guided CNN Restoration with Explicitly Signaled Linear CombinationabstractState-of-the-art Convolutional Neural Network (CNN) based loop restoration generally involves a CNN structure with a large number of parameters and applies the CNN model to those degraded frames uniformly to generate their restored version, even though the contents within these frames are different. By contrast, in this paper, we propose a Guided CNN Restoration (GNR) scheme, where a CNN is used in conjunction with explicitly signaled guide parameters, with an aim to adapt the CNN model to different input contents. Specifically, the CNN architecture is designed such that the final restoration is constrained within the subspace generated by various output channels of the CNN, and meanwhile the weighting parameters for a linear combination of the output channels to obtain the final restoration are explicitly signaled by the encoders. The proposed GNR is incorporated into an AV1 encoder to replace the anchor in-loop filters and the weighting parameters are written into the encoded bitstream. Experimental results show that given a small CNN with 3,312 parameters, the proposed approach achieves a BD-rate reduction of 3.06% over the AV1 anchor, while the traditional CNN-based method only achieves 1.39%. Lingyi Kong, Dandan Ding, Fuchang Liu, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
ICIP | 4 |
| 2020 | Perceptually Inspired Weighted MSE Optimization Using Irregularity-Aware Graph Fourier TransformabstractIn image and video coding applications, distortion has been traditionally measured using mean square error (MSE), which suggests the use of orthogonal transforms, such as the discrete cosine transform (DCT). Perceptual metrics such as Structural Similarity (SSIM) are typically used after encoding, but not tied to the encoding process. In this paper, we consider an alternative framework where the goal is to optimize a weighted MSE metric, where different weights can be assigned to each pixel so as to reflect their relative importance in terms of perceptual image quality. For this purpose, we propose a novel transform coding scheme based on irregularity-aware graph Fourier transform (IAGFT), where the induced IAGFT is orthogonal, but the orthogonality is defined with respect to an inner product corresponding to the weighted MSE. We propose to use weights derived from local variances of the input image, such that the weighted MSE aligns with SSIM. In this way, the associated IAGFT can achieve a coding efficiency improvement in SSIM with respect to conventional transform coding based on DCT. Our experimental results show a compression gain in terms of multi-scale SSIM on test images. Keng-Shih Lu, Antonio Ortega, Debargha Mukherjee, Yue Chen 0040 |
ICIP | 3 |
| 2020 | On Extended Transform Partitions For The Next Generation Video CODECabstractAV1, AOMedia's royalty free codec, has enjoyed a great amount of success since its 2018 release. It currently achieves 31% BDRATE gains over VP9, and is on its way to becoming YouTube's default codec. Although the industry is currently focused on the implementation and optimization of AV1, AOMedia Research continues to develop new coding tools that deliver higher coding gains within acceptable complexity bounds. Here, we focus on improving transform coding. While AV1 has made great strides in transform coding over VP9, the residue signal still consumes a large portion of the bitstream. In this paper, we describe a more flexible transform partitioning scheme, which will allow the next generation codec to more efficiently target areas in the residue signal with high energy, leading to better residue compression. Sarah Parker, Yue Chen 0040, Urvang Joshi, Elliott Karpilovsky, Debargha Mukherjee |
ICIP | 5 |
| 2019 | AV1 in-loop Filtering using a Wide-Activation Structured Residual NetworkabstractThe in-loop filter, which constitutes an important part in modern video coding, improves both subjective and objective quality of reconstructed frames. Lately, Convolutional Neural Network (CNN) has demonstrated its superiority over traditional methods in addressing in-loop filtering problem. In this paper, we develop a CNN-based in-loop filter, namely Wide Activation Residual Network (WARN), for AV1 encoder. On top of the plain Residual Network (ResNet), we introduce wide activation to each residual block, making a more reasonable allocation of network parameters. When incorporating WARN into video encoder, particular to inter coding, it is intricate to obtain the global optimum performance. After simplifying this as an end-to-end trainable problem, we propose a skipping method by taking advantage of the hierarchical reference structure in AV1. Experimental results show that our WARN achieves up to 14.42% and 9.64% BD-rate reduction in intra and inter coding, respectively. All the code and model of our approach are available at https://github.com/IVC-Projects/AV1_WARN. Dandan Ding, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
ICIP | 3 |
| 2019 | A CNN-based In-loop Filtering Approach for AV1 Video CodecabstractIn-loop filter using Convolutional Neural Network (CNN) has lately attracted lots of attention in video coding. CNN models may be trained to learn how to restore degradation introduced by compression in pictures, and hence effectively help improve the coding efficiency. State-of-the-art work in this field generally employs a single network to enhance reconstructed frames mainly in intra coding. In this paper, we develop a depth-variable network handling both intra and inter coding. The depth of our network is varied with the distortion levels of reconstructed frames. Moreover, we leverage a skip enhancing strategy for inter coding, which improves both the coding efficiency and the resulting visual quality, while maintaining low computational complexity. We apply our approach to AV1, a newly released video coding standard from AOM. Experimental results show that our approach achieves an average BD-rate reduction of 7.27% and 5.57% for intra and inter modes, respectively, compared to AV1 anchor. The code and model of our approach are published in our Github website [1]. Dandan Ding, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
PCS | 3 |
| 2019 | In-loop Frame Super-resolution in AV1abstractAV1 is a recently standardized royalty-free video codec from the industry consortium Alliance for Open Media. One of the most innovative coding tools supported in AV1 is an in-loop frame super-resolution mode, that allows an encoder to code any frame at a horizontally reduced spatial resolution by one of several levels, followed by upsampling and super-resolving to full resolution, before replacing reference buffers. This mode is partly enabled by a feature in AV1 that natively allows the motion compensated prediction loop to operate across scales between a coded frame and the available references, thereby allowing on-the-fly resolution change mid-stream within a sequence. For the actual super-resolving process a normative upscaler is followed by an in-loop restoration tool that recovers some of the high frequency information lost in the downsampling process. On-the-fly resolution change capability in conjunction with the frame-superresolution mode in AV1 opens up a new dimension for codec bitrate and quality optimization that has not been possible to explore in any prior standardized video codec with (soon expected) decoding hardware support. This paper provides an overview of how the relevant tools work in AV1, but unlocking them with intelligent encoder decisions to extract real-world benefit, is largely left as future work. Urvang Joshi, Debargha Mukherjee, Yue Chen 0040, Sarah Parker, Adrian Grange |
PCS | 2 |
| 2019 | Machine Learning Accelerated Transform Search For AV1abstractAV1 is the state-of-the-art open and royalty-free video compression format that achieves significant bitrate savings over previous generation of video codecs. One of AV1's major improvement over its predecessor VP9 is the support of more diverse and flexible transform size and kernel selection. However, it also drastically increases the search space for transform unit rate-distortion optimization in AV1 encoders. Unlike conventional encoder speed features that are based on heuristics, we propose a machine learning (ML) based approach to accelerate the transform size and kernel search for AV1. The ML models use input features extracted from the prediction residue block such as standard deviation, correlation and energy distribution. The output of the models indicates the estimated likelihood of which transform size and kernel would be selected as the optimal choice. Based on the ML models, the encoder can prune out the transform size and kernel candidates that are unlikely to be selected and save unnecessary computation to compute their rate-distortion cost. The proposed approach is implemented and tested on the AV1 reference library libaom. The experimental results show that satisfactory encoding speed improvement can be achieved with extremely low compression performance loss. The framework and methodology can also be easily migrated to other video codecs and implementations. Hui Su, Alexander Bokov, Debargha Mukherjee, Yunqing Wang, Yue Chen 0040 |
PCS | 4 |
| 2018 | Efficient AV1 Video Coding Using a Multi-layer FrameworkabstractThis paper proposes a multi-layer multi-reference prediction framework for effective video compression. Current AOM/AV1 baseline uses three reference frames for the inter prediction of each video frame. This paper first presents a new coding tool that extends the total number of reference frames in both forward and backward prediction directions. A multi-layer framework is then described, which suggests the encoder design and places different reference frames within one Golden Frame (GF) group to different layers. The multi-layer framework leverages the existing coding tools in the AV1 baseline, including the tool of "show_existing_frame" and the reference frame buffer update module of a wide flexibility. The use of extended ALTREF_FRAMEs is proposed, and multiple ALTREF_FRAME candidates are selected and widely spaced within one GF group. ALTREF_FRAME is a constructed, no-show reference obtained through temporal filtering of a look-ahead frame. In the multi-layer structure, one reference frame may serve different roles for the encoding of different frames through the virtual index manipulation. The experimental results have been collected over several video test sets of various resolutions and characteristics both texture- and motion-wise, which demonstrate that the proposed approach achieves a consistent coding gain compared to the AV1 baseline. For instance, using PSNR as the distortion metric, an average bitrate saving of 5.57+% in BDRate is obtained for the CIF-level resolution set, some of which has a gain of up to 13+%, and 4.47% on average for the VGA-level resolution set, some of which up to 18+%. Zoe Liu, Debargha Mukherjee, Jingning Han, Paul Wilkins, Yaowu Xu, Kenneth Rose |
DCC | 3 |
| 2018 | An Overview of Core Coding Tools in the AV1 Video CodecabstractAV1 is an emerging open-source and royalty-free video compression format, which is jointly developed and finalized in early 2018 by the Alliance for Open Media (AOMedia) industry consortium. The main goal of AV1 development is to achieve substantial compression gain over state-of-the-art codecs while maintaining practical decoding complexity and hardware feasibility. This paper provides a brief technical overview of key coding techniques in AV1 along with preliminary compression performance comparison against VP9 and HEVC. Yue Chen 0040, Debargha Mukherjee, Jingning Han, Adrian Grange, Yaowu Xu, Zoe Liu, Sarah Parker, Hui Su, Urvang Joshi, Ching-Han Chiang, Yunqing Wang, Paul Wilkins, Jim Bankoski, Luc N. Trudeau, Nathan E. Egge, Jean-Marc Valin, Thomas Davies 0002, Steinar Midtskogen, Andrey Norkin, Peter De Rivaz |
PCS | 2 |
| 2018 | Efficient Rate-distortion Approximation and Transform Type Selection using Laplacian OperatorsabstractRate-distortion (RD) optimization is an important tool in many video compression standards and can be used for transform selection. However, this is typically very computationally demanding because a full RD search involves the computation of transform co-efficients for each candidate transform. In this paper, we propose an approach that uses sparse Laplacian operators to estimate the RD cost by computing a weighted squared sum of transform coefficients, without having to compute the actual transform coefficients. We demonstrate experimentally how our method can be applied for transform selection. Implemented in the AV1 encoder, our approach yields a significant speed-up in encoding time with a small increase in bitrate. Keng-Shih Lu, Antonio Ortega, Debargha Mukherjee, Yue Chen 0040 |
PCS | 3 |
| 2017 | Variable block-size overlapped block motion compensation in the next generation open-source video codecabstractTraditional motion compensation is based on a simplistic assumption of a piecewise constant and pure translational motion field. Premised on the inefficiency of such block-copying-based algorithms in compressing real-life videos, this paper proposes overlapped block motion compensation for AV1, a next generation open-source video codec. Motions assigned to surrounding blocks will contribute to predicting a current block, via a well-defined overlapping scheme appropriately designed for advanced variable block-size partitioning frameworks. To efficiently incorporate the proposed tool, the encoder is optimized with proper early termination to manage complexity in conjunction with weighted motion search accounting for non-uniformly weighted distortion in overlapped prediction. Moreover as an extension, by re-arranging decoding operations, we enable non-causal overlapping which also utilizes motions of bottom and right neighbors. Being verified by integration in the AV1 codec, the tools achieve considerable coding gains on various test sets. Yue Chen 0040, Debargha Mukherjee |
ICIP | 2 |
| 2017 | Adaptive interpolation filter scheme in AV1abstractVideo codecs heavily depend on sub-pixel level motion compensation to achieve superior compression performance. Interpolation filters with both anti-aliasing and denoising properties play a critical role in producing high quality prediction at sub-pixel positions. Prior research has developed many adaptive filtering schemes to improve the prediction precision for compression gains. On the other hand, such filtering operations require intense computation and may lead to scattered cache footprints, therefore, account for a major portion of the overall decoding cost in both software and hardware implementations. An adaptive interpolation filtering scheme is proposed in this work to optimize the trade off between prediction quality and decoding performance. It employs a separable model and selects filter kernels independently for horizontal and vertical directions to better capture statistical variations. In order to obtain sharper transition and reduce the ripple effect in the passband in frequency domain, a 12-tap filter is introduced in conjunction with a complimentary operation design that minimizes its impact on the decoding performance. The scheme achieves on average 1.3% coding gains across a wide range of test settings, with fairly limited additional hardware cost. Ching-Han Chiang, Jingning Han, Stan Vitvitskyy, Debargha Mukherjee, Yaowu Xu |
ICIP | 4 |
| 2017 | A switchable loop-restoration with side-information framework for the emerging AV1 video codecabstractImage restoration schemes have traditionally been targeted only for use in a blind scenario, where the aim is to improve the quality of an image or video after it has suffered degradations in capture, processing, storage, coding or transmission, at a time when the undegraded source is no longer available. However, these schemes can also be adopted for compression, because any information that can be restored can be saved, thereby aiding compressibility. In this use case, the source is known at the time of compression, and the encoder can send additional side-information to the decoder with the compressed bit-stream to specify how the decoder is supposed to restore. In this paper, we describe three novel schemes that are applied in-loop to reconstructed frames after a conventional deblocking loop filter has been applied. These schemes are switchable within a frame per suitably sized tile. The specific schemes described are based on separable symmetric Wiener filters, dual self-guided filters with subspace projection, and domain transform recursive filters. Results with the new coding tool are presented on the emerging AV1 video codec on standard test sets, showing upwards of 2% bit-rate savings. Debargha Mukherjee, Shunyao Li, Yue Chen 0040, Aamir Anis, Sarah Parker, Jim Bankoski |
ICIP | 1 |
| 2017 | Global and locally adaptive warped motion compensation in video compressionabstractReal motion is often complex, and cannot always be sufficiently described by a translational model for the purpose of video coding. Translational motion models can only map a rectangular to a rectangular of the same size in a reference frame, but real motion observed in videos often go beyond that. For example, motion due to camera shake, panning and zoom might require transformations that support shearing, scaling, rotation and changes in aspect ratio. In order to handle these instances adequately for video encoding, we introduce two coding modes - a global motion mode and a locally adaptive warped motion mode, that make use of affine and/or homographic projections to more accurately describe and predict complex motion. These warped modes respectively capture global motion computed at the frame level, or local motion computed at the block level based on local motion statistics. We pay special attention to efficient SIMD friendly decoder-side warping operations to make these tools practically useful. Together, these techniques prove to yield considerable coding gains on standard test sets. Sarah Parker, Yue Chen 0040, David Barker, Peter De Rivaz, Debargha Mukherjee |
ICIP | 5 |
| 2017 | Learning separable transforms by inverse covariance estimationabstractOrthogonal transforms are one of the most important components of a video encoder system. They are applied to residual block images obtained as the difference between a target and its prediction. In this paper we propose a framework to design separable transforms from prediction residual statistics. We model the data as a 2D Gaussian Markov random field and approximate its inverse covariance by a matrix with a separable structure, thus explicitly constructing a separable orthonormal matrix that approximates the KLT. Our designed transforms can adapt to prediction residual statistics, have low complexity (compared to non separable transforms), require selecting few parameters and outperform hybrid DCT/ADST separable transform for intra coding of AV1 residuals. Eduardo Pavez, Antonio Ortega, Debargha Mukherjee |
ICIP | 3 |
| 2016 | Introduction of New Associate EditorsabstractPresents a listing of the new Associate Editors for this issue of the publication. Nikolaos V. Boulgouris, David Bull 0001, Marco Cagnazzo, Andrea Cavallaro, Gene Cheung, Amit K. Roy-Chowdhury, Pedro Comesaña Alfaro, Sarp Ertürk, Markus Flierl, Gian Luca Foresti, Gang Hua 0001, Zhu Li 0001, Weisi Lin, Siwei Ma 0001, Pramod Kumar Meher, Debargha Mukherjee, Aleksandra Pizurica, Andrea Prati 0001, Paolo Remagnino, Arun Ross, Shin'ichi Satoh 0001, Andreas E. Savakis, Heiko Schwarz, Ling Shao 0001, Shervin Shirmohammadi, Giuseppe Valenzise, Meng Wang 0001, Zhou Wang 0001, Yonggang Wen 0001, Dong Xu 0001, Junsong Yuan 0001, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 17 |
| 2015 | Super-resolution for inconsistent scalable video streamingabstractWe propose an example-based super-resolution method for a new framework of scalable video streaming. The proposed method is applicable to scalable video where the enhancement layer of some frames (e.g., within the same Group-of-Pictures) might be dropped due to changing network conditions. This leads to a streaming scenario that we call Inconsistent Scalable Video (ISV) streaming. In this paper, we focus on spatial ISV. In particular, at the decoder, the frames with the enhancement layer are used as the super-resolution dictionary for other video frames that their enhancement layers were dropped. Using the proposed method, the frames with dropped enhancement layer are scaled up to the original size. Our experimental results show a clear improvement over traditional interpolation based scaling. Abo Talib Mahfoodh, Debargha Mukherjee, Hayder Radha |
ICIP | 2 |
| 2014 | Joint inter-intra prediction based on mode-variant and edge-directed weighting approaches in video codingabstractMost modern video compression codecs, like VP9, HEVC and H.264, encode square or rectangular blocks either by inter prediction or intra prediction. A joint inter-intra predictor that combines motion compensation and intra extrapolation by two novel weighting schemes is proposed to improve compression quality. Prior work on joint prediction employs inter-intra weights that only rely on the pixel locations. As an enhancement, we design a weighting approach by also considering the angle of intra prediction, which is the actual direction that the intra prediction errors evolve. Moreover, our second approach, inspired by prior work on geometric-partition-based motion compensation, breaks the limitation of traditional quad-tree partition by jointly using different predictors that implies soft step weighting functions for new and existing objects co-occurring around irregular motion edges. The proposed joint prediction approaches deliver consistent coding gains, as shown by extensive experiments on the experimental branch of VP9, Google's open source video compression tool. Yue Chen 0040, Debargha Mukherjee, Jingning Han, Kenneth Rose |
ICASSP | 2 |
| 2014 | A pre-filtering approach to exploit decoupled prediction and transform block structures in video codingabstractRecent video coding techniques allow for decoupling of the transform block partition from that employed for prediction. For example, HEVC allows a transform block to overlap multiple prediction blocks. This paper is premised on the observation that in order to truly realize the potential of such enhanced flexibility, it is necessary to account for and mitigate considerable side effects due to stitching together independently predicted blocks, including the emergence of spurious high frequency components from sharp transitions across boundaries, which undermine the transform efficacy. The proposed solution involves an appropriately designed pre-filtering approach to mitigate boundary transition effects whenever a transform spans data from multiple prediction blocks. Moreover, this filtering technique enables extending the flexibility in decoupling prediction and transform structures, as various restrictions may now be eliminated. In particular, it makes it possible and beneficial to allow a transform block to span residual data from both inter and intra predicted blocks, whereas HEVC necessarily forces a single type of prediction in each coding unit. The method is further extended to include motion refinement that accounts for the pre-filtering approach. Experiments provide evidence for consistent coding gains over HEVC and VP9. Yue Chen 0040, Kenneth Rose, Jingning Han, Debargha Mukherjee |
ICIP | 4 |
| 2013 | A butterfly structured design of the hybrid transform coding schemeabstractThe hybrid transform coding scheme that alternates amongst the asymmetric discrete sine transform (ADST) and the discrete cosine transform (DCT) depending on the boundary prediction conditions, is an efficient tool for video and image compression. It optimally exploits the statistical characteristics of prediction residual, thereby achieving significant coding performance gains over the conventional DCT-based approach. A practical concern lies in the intrinsic conflict between transform kernels of ADST and DCT, which prevents a butterfly structured implementation for parallel computing. Hence the hybrid transform coding scheme has to rely on matrix multiplication, which presents a speed-up barrier due to under-utilization of the hardware, especially for larger block sizes. In this work, we devise a novel ADST-like transform whose kernel is consistent with that of DCT, thereby enabling butterfly structured computation flow, while largely retaining the performance advantages of hybrid transform coding scheme in terms of compression efficiency. A prototype implementation of the proposed butterfly structured hybrid transform coding scheme is available in the VP9 codec repository. Jingning Han, Yaowu Xu, Debargha Mukherjee |
PCS | 3 |
| 2013 | The latest open-source video codec VP9 - An overview and preliminary resultsabstractGoogle has recently finalized a next generation open-source video codec called VP9, as part of the libvpx repository of the WebM project (http://www.webmproject.org/). Starting from the VP8 video codec released by Google in 2010 as the baseline, various enhancements and new tools were added, resulting in the next-generation VP9 bit-stream. This paper provides a brief technical overview of VP9 along with comparisons with other state-of-the-art video codecs H.264/AVC and HEVC on standard test sets. Results show VP9 to be quite competitive with mainstream state-of-the-art codecs. Debargha Mukherjee, Jim Bankoski, Adrian Grange, Jingning Han, John Koleszar, Paul Wilkins, Yaowu Xu, Ronald Bultje |
PCS | 1 |
| 2013 | Learning-Based, Automatic 2D-to-3D Image and Video ConversionabstractDespite a significant growth in the last few years, the availability of 3D content is still dwarfed by that of its 2D counterpart. To close this gap, many 2D-to-3D image and video conversion methods have been proposed. Methods involving human operators have been most successful but also time-consuming and costly. Automatic methods, which typically make use of a deterministic 3D scene model, have not yet achieved the same level of quality for they rely on assumptions that are often violated in practice. In this paper, we propose a new class of methods that are based on the radically different approach of learning the 2D-to-3D conversion from examples. We develop two types of methods. The first is based on learning a point mapping from local image/video attributes, such as color, spatial position, and, in the case of video, motion at each pixel, to scene-depth at that pixel using a regression type idea. The second method is based on globally estimating the entire depth map of a query image directly from a repository of 3D images ( image+depth pairs or stereopairs) using a nearest-neighbor regression type idea. We demonstrate both the efficacy and the computational efficiency of our methods on numerous 2D images and discuss their drawbacks and benefits. Although far from perfect, our results demonstrate that repositories of 3D content can be used for effective 2D-to-3D image conversion. An extension to video is immediate by enforcing temporal continuity of computed depth maps. Janusz Konrad, Prakash Ishwar, Debargha Mukherjee |
IEEE Trans. Image Process. | 5 |
| 2012 | Video Super-Resolution Using Codebooks Derived From Key-FramesabstractExample-based super-resolution (SR) is an attractive option to Bayesian approaches to enhance image resolution. We use a multiresolution approach to example-based SR and discuss codebook construction for video sequences. We match a block to be super-resolved to a low-resolution version of the reference high-resolution image blocks. Once the match is found, we carefully apply the high-frequency contents of the chosen reference block to the one to be super-resolved. In essence, the method relies on “betting” that if the low-frequency contents of two blocks are very similar, their high-frequency contents also might match. In particular, we are interested in scenarios where examples can be picked up from readily available high-resolution images that are strongly related to the frame to be super-resolved. Hence, they constitute an excellent source of material to construct a dynamic codebook. Here, we propose a method to super-resolve a video using multiple overlapped variable-block-size codebooks. We implemented a mixed-resolution video coding scenario, where some frames are encoded at a higher resolution and can be used to enhance the other lower-resolution ones. In another scenario, we consider the framework where the camera captures a video at a lower resolution and also takes periodic snapshots at a higher resolution. Results indicate substantial gains over interpolation and fixed-codebook SR, and significant gains over previous works as well. Edson M. Hung, Ricardo L. de Queiroz, Fernanda Brandi, Karen França de Oliveira, Debargha Mukherjee |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2010 | Multi-resolution redundancy for error-resilient video transmissionabstractIn this paper we advocate a multi-resolution mechanism for redundant information generation and transmission with low overheads in bit-rate, in order to enable reliable video communication over challenging lossy networks with low latency. Our previous work entitled RECAP, transmitted a low resolution redundant version of a parent video stream, in order to achieve effective concealment of isolated and burst losses by clever multi-frame super-resolution processing at the decoder. A reference picture selection mechanism was used to transmit the low-resolution redundant video with high guarantee and stop drift in case of losses. In this work, we extend the framework to incorporate an additional distributed Wyner-Ziv coding layer on the low-resolution information to further correct errors, and consequently improve the error-concealed picture quality. The error-concealed super-resolved frame using the low-resolution information alone, now acts as side-information at the decoder to correct additional errors by decoding the Wyner-Ziv layer. Preliminary results are presented to demonstrate the efficacy of the proposed approach. Debargha Mukherjee, Wai-tian Tan |
ICASSP | 1 |
| 2010 | Complexity-scalable H.264/AVC in an IPP-based video encoderabstractReal-time high-definition video encoding is a computation-hungry task that challenges software-based solutions. For that, in this work we adopted an Intel software implementation of an H.264 video encoder and optimized its prediction stage in the complexity sense (C). Thus, besides looking for the coding options which lead to the best coded representation in terms of rate and distortion, we constrain the process to fit within a certain time budget. We present an RDC-optimized framework which allows for real-time HD video compression. Tiago A. da Fonseca, Ricardo L. de Queiroz, Debargha Mukherjee |
ICIP | 3 |
| 2010 | Mixed-resolution distributed video codec without motion estimation at the encoderabstractInspired by recent results showing that Wyner-Ziv coding using a combination of source and channel coding may be more efficient than pure channel coding, we have applied coset codes for the source coding part in the transform domain for Wyner-Ziv coding of video. The framework is a mixed-resolution approach where reduced encoding complexity is achieved by low resolution encoding of non-reference frames and regular encoding of the reference frames. Different from our previous works, no motion estimation is carried at the encoder, neither for the low resolution frames nor the reference frames. The entropy coders of the H.264/AVC codec were tuned to improve their performance for encoding the cosets. An encoding mode with lowest encoding complexity than H.264/AVC intra mode is achieved. Experimental results show a competitive rate-distortion performance especially at low bit rates. Bruno Macchiavello, Edson M. Hung, Ricardo L. de Queiroz, Debargha Mukherjee |
ICIP | 4 |
| 2010 | A Wyner-Ziv Video TranscoderabstractWyner-Ziv (WZ) coding of video utilizes simple encoders and highly complex decoders. A transcoder from a WZ codec to a traditional codec can potentially increase the range of applications for WZ codecs. We present a transcoder scheme from the most popular WZ codec architecture to a differential pulse code modulation/discrete cosine transform codec. As a proof of concept, we implemented this transcoder using a simple pixel-domain WZ codec and the standard H.263+. The transcoder design aims at reducing complexity as a large amount of computation is saved by reusing the motion estimation, calculated at the side information generation process, and the l-frame streams. New approaches are used to generate side information and to map motion vectors for the transcoder. Results are presented to demonstrate the transcoder performance. Eduardo Peixoto, Ricardo L. de Queiroz, Debargha Mukherjee |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Distributed Image Coding for Digital Image Recovery From the Print-Scan ChannelabstractA printed digital photograph is difficult to reuse because the digital information that generated the print may no longer be available. This paper describes a method for approximating the original digital image by combining a scan of the printed photograph with digital auxiliary information kept together with the print. We formulate and solve the approximation problem using a Wyner-Ziv coding framework. During encoding, the Wyner-Ziv auxiliary information consists of a small amount of digital data composed of a number of sampled luminance pixel blocks and a number of sampled color pixel values to enable subsequent accurate registration and color-reproduction during decoding. The registration and color information is augmented by an additional amount of digital data encoded using Wyner-Ziv coding techniques that recovers residual errors and lost high spatial frequencies. The decoding process consists of scanning the printed photograph, together with a two step decoding process. The first decoding step, using the registration and color auxiliary information, generates a side-information image which registers and color corrects the scanned image. The second decoding step uses the additional Wyner-Ziv layer together with the side-information image to provide a closer approximation of the original, reducing residual errors and restoring the lost high spatial frequencies. The experimental results confirm the reduced digital storage needs when the scanned print assists in the digital reconstruction. Ramin Samadani, Debargha Mukherjee |
IEEE Trans. Image Process. | 2 |
| 2009 | Tables for practical Wyner-Ziv coding of Laplacian sourcesabstractMany practical coding scenarios deal with sources with transform coefficients that are well modeled as Laplacians. For the Wyner-Ziv coding problem for such sources when correlated side-information is available at the decoder, the side-information is modeled as obtained by independent additive Laplacian or Gaussian innovation on the source. This paper deals with the optimal choice of encoding parameters for practical Wyner-Ziv coding in such scenarios, using the same quantizer family as in the regular codec to cover a range of rate-distortion trade-offs, given the variances of the source and innovation. Using our prior analysis of a general encoding model based on multi-level coset codes combining source and channel coding, we present comprehensive tables with optimal encoding parameters. These tables can be readily incorporated into a practical codec to read off the encoding parameters. Debargha Mukherjee |
ICASSP | 1 |
| 2009 | Receiver error concealment using acknowledge preview (RECAP) - An approach to resilient video streamingabstractHigh-quality and low-latency video streaming is essential to providing a natural user experience in video conferencing. This is challenging over lossy networks since compressed video is highly fragile while the low-latency requirement limits the effectiveness of traditional error control approaches such as retransmission and forward error correction. In this paper, we advocate a practical solution for low-latency video communications over best-effort networks that employs an additional low-quality, low-resolution but robustly coded copy of the video. This approach, called RECAP, incurs minimal rate overhead, and can be combined with previously decoded frames to achieve effective concealment of isolated and burst losses even under tight delay constraints. RECAP achieves PSNR gains of 2-6 dB against complete frame loss. Chuohao Yeo, Wai-tian Tan, Debargha Mukherjee |
ICASSP | 3 |
| 2009 | Efficiency improvements for a geometric-partition-based video coderabstractH.264/AVC has brought an important increase in coding efficiency in comparison to previous video coding standards. One of its features is the use of macroblock partitioning in a tree-based structure. The use of macroblock partitions based in arbitrary line segments, like wedge partitions, has been reported to increase coding gains. The main problem of these non-standard new partitions schemes is the increase in computational complexity. Thus, our work proposes improvements in this extension to the H.264/AVC standard. First, we present a motion vector prediction scheme based on directional partitions. Second, we present a method for complexity reduction based on the most frequent partitions. The results show that it is possible to still produce good coding gains with lower complexity than previous approaches. Renan U. Ferreira, Edson M. Hung, Ricardo L. de Queiroz, Debargha Mukherjee |
ICIP | 4 |
| 2009 | Design of high capacity 3D print codes aiming for robustness to the PS channel and external distortionsabstractThe process of adding high-density information onto printed material enables and improves interesting hardcopy document applications, such as: security, authentication, physical-electronic round tripping, item-level tagging as well as consumer/product interaction. This investigation on robust and high capacity print codes aims to maximize information payload in a given printed page area, subject to robustness to distortions originated by printing and scanning processes and also to degradations introduced by user manipulation of printed documents. The novel approach includes statistical print-and-scan channel characterization, designing of robust segmentation, unsupervised Bayesian color classification with expectation-maximization algorithm for parameters estimation of a mixture of Gaussians model and design of error correction codes. Results illustrate the performance evaluated under real channel and distortions conditions. High payload is achieved with sufficient robustness to distortions resulting of regular office hardcopy document handling: print-and-scan channel and user manipulation. Joceli Mayer, José Carlos M. Bermudez, Andrei Piccinini Legg, Bartolomeu F. Uchôa Filho, Debargha Mukherjee, Amir Said, Ramin Samadani, Steven J. Simske |
ICIP | 5 |
| 2009 | Mapping motion vectors for Awyner-Ziv video transcoderabstractWyner-Ziv (WZ) coding of video utilizes simple encoders and highly complex decoders. A transcoder from a WZ codec to a traditional codec can potentially increase the range of applications for WZ codecs. We present a transcoder scheme from the most popular WZ codec architecture to a DPCM/DCT codec. As a proof of concept, we implemented this transcoder using a simple pixel domain WZ codec and the standard H.263+. The transcoder design aims at reducing complexity, since the transcoder has to perform both WZ decoding and DPCM/DCT encoding, including motion estimation. New approaches are used to map motion vectors for such a transcoder. Results are presented to demonstrate the transcoder performance. Eduardo Peixoto, Ricardo L. de Queiroz, Debargha Mukherjee |
ICIP | 3 |
| 2009 | Design of high capacity 3D print codes with visual cues aiming for robustness to the PS channel and external distortionsabstractAdding high-density information to printed materials enables and improves interesting hardcopy document applications involving security, authentication, physical-electronic round tripping, item-level tagging, and consumer/product interaction. This investigation of robust and high capacity print codes aims to maximize information payload in a given printed page area, subject to robustness to channel errors including distortions introduced by the printing and scanning processes and also due to the usual degradations introduced by user manipulation of printed documents. The novel approach includes statistical print-and-scan channel characterization, designing of robust segmentation using visual cues, unsupervised Bayesian color classification with expectation-maximization algorithm for parameters estimation of a mixture of Gaussians model and design of error correction codes. Results illustrate the performance evaluated under real channel and distortions conditions. High payload is achieved with sufficient robustness to distortions resulting of regular office hardcopy document handling: print-and-scan channel and user manipulation. Joceli Mayer, José Carlos M. Bermudez, Andrei Piccinini Legg, Bartolomeu F. Uchôa Filho, Debargha Mukherjee, Amir Said, Ramin Samadani, Steven J. Simske |
MMSP | 5 |
| 2009 | Iterative Side-Information Generation in a Mixed Resolution Wyner-Ziv FrameworkabstractWe propose a mixed resolution framework based on full resolution key frames and spatial-reduction-based Wyner-Ziv coding of intermediate nonreference frames. Improved rate-distortion performance is achieved by enabling better side-information generation at the decoder side and better rate-allocation at the encoder side. The framework enables reduced encoding complexity by low resolution encoding of the nonreference frames, followed by Wyner-Ziv coding of the Laplacian residue. The quantized transform coefficients of the residual frame are mapped to cosets without the use of a feedback channel. A study to select optimal coding parameters in the creation of the memoryless cosets is made. Furthermore, a correlation estimation mechanism that guides the parameter choice process is proposed. The decoder first decodes the low resolution base layer and then generates a super-resolved side-information frame at full resolution using past and future key frames. Coset decoding is carried using side-information to obtain a higher quality version of the decoded frame. Implementation results are presented for the H.264/AVC codec. Bruno Macchiavello, Debargha Mukherjee, Ricardo L. de Queiroz |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Super-resolution of video using key frames and motion estimationabstractMany scalable video coding systems use frame down- sampling in order to reduce complexity and to enable enhancement layers. Super-resolution (SR) can be used to help the up-sampling and recovering processes of those frames. We are interested in reversed-complexity (distributed) coding methods, wherein few key frames are encoded at normal resolution, while the rest are down- sampled and encoded at reduced resolution along with the enhancement layers. We are only interested on the decoder side, wherein we carry motion estimation of the down- sampled frames using the key frames as references. When a match is made, the high-frequency components of the key frames (KF) are used to super-resolve the non-key frames (NKF). The motion estimation process is performed using blocks of band-pass versions of the frames, rather than low- pass ones. Results indicate the improved performance of the proposed super-resolution algorithm. Fernanda Brandi, Ricardo L. de Queiroz, Debargha Mukherjee |
ICIP | 3 |
| 2008 | Rate-distortion analysis of weighted prediction for error resilienceabstractIn this paper, we extend our work in [1] and address an approach to theoretically analyze the rate-distortion (R-D) performance of the weighted prediction feature provided within the scope of H.264/AVC. We consider the weighted prediction as a standard-compatible leaky prediction approach for the purpose of error resilience. We adopt a quantization noise model that explicitly formulates the relationship between the data rate and the distortion in the mean-square-error (MSE) sense. We derive a comprehensive rate-distortion function for both the error-free scenario and the one with error drift. Through adjusting the weight coefficients in H.264/AVC, we also simulate H.264/AVC video streaming over error-prone networks and obtain the operational rate-distortion results using various leaky factors for both error-free and error-drift scenarios. We compare our theoretical results with the operational R-D curves and demonstrate that the theoretical results conform with the operational results. Ragip Kurceren, Debargha Mukherjee |
ICIP | 3 |
| 2008 | Parameter estimation for an H.264-based distributed video coderabstractIn this paper we present a statistical model used to select coding parameters for a mixed resolution Wyner-Ziv framework implemented using the H.264/AVC standard. This paper extends the results of a previous work for the H.263+ case to the H.264/AVC coder, since the parameters need to be recalculated for the H.264 case. The proposed correlation estimation mechanism guides the parameter choice process, and also yields the statistical model used for decoding. This mechanism is proposed based on extracting edge information and residual error rate in co-located blocks from the low resolution base layer that is available at both ends. Bruno Macchiavello, Ricardo L. de Queiroz, Debargha Mukherjee |
ICIP | 3 |
| 2008 | Distributed image coding for imformation recovery from the print-scan channelabstractA printed photograph is difficult to reuse because the digital pixels that generated it are no longer available. This paper describes approximating the original digital pixels by combining a scan of the printed photograph with small amounts of digital auxiliary information kept with the print, using an encoding such as a dense barcode. The auxiliary information enables accurate registration and color-reproduction, and the recovery of information lost due to the print-scan channel, by distributed source coding techniques. Approximating the original digital image enables many uses, including making high-fidelity reprints, as well as robust watermark detection. A block entropy coder using terminated tree codes for cosets of DCT coefficients is proposed for distributed image coding. Debargha Mukherjee, Ramin Samadani |
ICIP | 1 |
| 2008 | Exploiting patterns of data magnitude for efficient image codingabstractIn order to increase the computational efficiency of compression methods we have to consider that new hardware architectures increasingly rely more on wider data paths and parallel processing (e.g., SIMD and multi-core), than on faster clocks. Higher data throughputs are achieved with entropy coding methods that process larger amounts of information each time, and use context dependencies that are less complicated and that can be quickly updated. We propose a coding method with properties more suited to the new processors, that achieves better compression by exploiting patterns of data magnitude. We present experimental results on image coding implementations that take advantage of the fast decay of transform coefficient variance with frequency. Amir Said, Debargha Mukherjee |
ICIP | 2 |
| 2008 | Super resolution of video using key framesabstractIn many video compression systems, the frames are down-sampled before transmission. Also, in many scalable systems, the residual after down- and up-sampling is encoded and transmitted. Sometimes, a few frames are encoded at normal resolution (key frames) while the other frames are encoded at reduced resolution. Super resolution can be used to enhance the up-sampling process, using motion information to improve traditional interpolation. In this paper, we propose to use a super resolution method to up-sample the non-key frames using the key frames as reference. We build dictionaries on-the- fly using the key frames instead of the traditional off-line training images. The high-frequency data of matching blocks are added to the low-resolution blocks. Since the key frames are very similar to the non-key frames, the method is robust enough to allow successful super resolution of highly compressed (severely degraded) sequences. Results are presented for many predefined block sizes, key frames frequencies, and compression parameters. Fernanda Brandi, Ricardo L. de Queiroz, Debargha Mukherjee |
ISCAS | 3 |
| 2007 | Format-Independent Authentication of Arbitrary Scalable Bit-Streams using One-Way AccumulatorsabstractWe propose a mechanism for authentication of general scalable bit-streams, based on quasi-commutative one-way accumulator functions. Such functions allow flexible partitioning between an auxiliary hash computed for removed parts of an original bit-stream and the hash that a receiver can compute from the received bit-stream, and yet allow generation of a common root hash for the original bit-stream against which the authentication may be conducted. Unlike prior work using Merkle hash trees, the number of auxiliary hashes to be transmitted is always one, independent of the actual version of the bit-stream to be authenticated, and the mechanism is independent of the order of hash accumulation. Further, the method readily lends itself to format-independent authentication and adaptation mechanism by use of appropriate standardized metadata. Debargha Mukherjee |
ICASSP (2) | 1 |
| 2007 | Complexity Control for Real-Time Video CodingabstractA methodology for complexity scalable video encoding and complexity control within the framework of the H.264/AVC video encoder is presented. To yield good rate-distortion performance under strict complexity/time constraints for instance in real-time communication, a framework for optimal complexity allocation at the macroblock level is necessary. We developed a macroblock level fast motion estimation based complexity scalable motion/mode search algorithm where the complexity is adapted jointly by parameters that determine the aggressiveness of an early stop criteria, the number of ordered modes searched, and the accuracy of motion estimation steps for the INTER modes. Next, these complexity parameters are adapted per macroblock based on a control loop to approximately satisfy an encoding frame rate target. The optimal manner of adapting the parameters is derived from prior training. Results using the developed scalable complexity H.264/AVC encoder demonstrate the benefit of adaptive complexity allocation over uniform complexity scaling. Emrah Akyol, Debargha Mukherjee |
ICIP (1) | 2 |
| 2007 | Motion-Based Side-Information Generation for a Scalable Wyner-Ziv Video CoderabstractA motion-based side-information generation scheme with semi super-resolution for a scalable Wyner-Ziv coder framework is introduced. It is known that the performance of any Wyner-Ziv coder is heavily dependent on the efficiency of the side-information generation. We propose an iterative block based scheme to generate a semi super-resolution frame using the past and future reference frames which should be coded at full-resolution. To enable this side-information generation the framework should allow for low encoding complexity, reducing the spatial resolution only in the non-reference frames. The enhancement layer is produced using a residual frame of the reduced resolution encoded frame. The decoder first decodes the low resolution base layer and iteratively generates the side-information, along with channel decoding, to obtain a higher quality version of the decoded frame. Results of the implementation of the framework using the motion-based side-information in the H.263+ and H.264 standards are presented. Bruno Macchiavello, Ricardo L. de Queiroz, Debargha Mukherjee |
ICIP (6) | 3 |
| 2007 | A simple reversed-complexity Wyner-Ziv video coding mode based on a spatial reduction frameworkabstractA spatial-resolution reduction based framework for incorporation of a Wyner-Ziv frame coding mode in existing video codecs is presented, to enable a mode of operation with low encoding complexity. The core Wyner-Ziv frame coder works on the Laplacian residual of a lower-resolution frame encoded by a regular codec at reduced resolution. The quantized transform coefficients of the residual frame are mapped to cosets to reduce the bit-rate. A detailed rate-distortion analysis and procedure for obtaining the optimal parameters based on a realistic statistical model for the transform coefficients and the side information is also presented. The decoder iteratively conducts motion-based side-information generation and coset decoding, to gradually refine the estimate of the frame. Preliminary results are presented for application to the H.263+ video codec. Debargha Mukherjee, Bruno Macchiavello, Ricardo L. de Queiroz |
VCIP | 1 |
| 2007 | Pet fur color and texture classificationabstractObject segmentation is important in image analysis for imaging tasks such as image rendering and image retrieval. Pet owners have been known to be quite vocal about how important it is to render their pets perfectly. We present here an algorithm for pet (mammal) fur color classification and an algorithm for pet (animal) fur texture classification. Per fur color classification can be applied as a necessary condition for identifying the regions in an image that may contain pets much like the skin tone classification for human flesh detection. As a result of the evolution, fur coloration of all mammals is caused by a natural organic pigment called Melanin and Melanin has only very limited color ranges. We have conducted a statistical analysis and concluded that mammal fur colors can be only in levels of gray or in two colors after the proper color quantization. This pet fur color classification algorithm has been applied for peteye detection. We also present here an algorithm for animal fur texture classification using the recently developed multi-resolution directional sub-band Contourlet transform. The experimental results are very promising as these transforms can identify regions of an image that may contain fur of mammals, scale of reptiles and feather of birds, etc. Combining the color and texture classification, one can have a set of strong classifiers for identifying possible animals in an image. Jonathan Yen, Debargha Mukherjee, Suk Hwan Lim, Daniel Tretter |
VCIP | 2 |
| 2006 | On Macroblock Partition for Motion CompensationabstractIn the H.264/AVC video coding standard, motion compensation can be performed by partitioning macroblocks into square or rectangular sub-macroblocks in a quadtree decomposition. This paper studies a motion compensation method using wedges, i.e. partitioning macroblocks or sub-macroblocks into two regions by an arbitrary line segment. This technique allows the shapes of the divided regions to better match the boundaries between moving objects. However, there are a large number of ways to slice a block and searching exhaustively over all of them would be an extremely computer-intensive task. Thus, we propose a fast algorithm which detects the predominant edge orientations within a block in order to pre-select candidate wedge lines. Finally a comparison among macroblock partition methods is performed, which points to the higher performance of the wedge partition method. Edson M. Hung, Ricardo L. de Queiroz, Debargha Mukherjee |
ICIP | 3 |
| 2005 | Format independent encryption of generalized scalable bit-streams enabling arbitrary secure adaptations [multimedia communication applications]abstractSecure format-independent adaptation of bit-streams during delivery is becoming increasingly important to cope with content piracy while still accommodating diverse networks, terminals and formats. Generalized scalable bit-streams are particularly advantageous in this regard, since they enable a variety of efficient and secure adaptations. Further, by associating such a bit-stream with appropriate metadata, such as those standardized in MPEG-21 Part 7, entitled digital item adaptation (DIA), the adaptation process can be fully format-independent. In this paper, to maximally secure a generalized scalable bit-stream while allowing arbitrary encrypted domain adaptations, strong progressive encryption methods are extended to multiple dimensions. Further it is shown that by appropriate modeling of such bit-streams and re-use of some DIA descriptions, the encryption and decryption engines themselves can be entirely metadata-driven and format-independent. This leads to end-to-end format-independent secure and adaptive delivery architectures for scalable bit-streams. Debargha Mukherjee, Huisheng Wang, Amir Said, Sam Liu |
ICASSP (2) | 1 |
| 2005 | Compact dependent key generation methods for encryption-based subscription differentiation for scalable bit-streamsabstractSubscription differentiation can be easily supported by deploying a scalable (prioritized) bit-stream and encrypting different parts of the bit-stream with various keys. Based on the subscription level, only the keys corresponding to the authorized part of the bit-stream are transmitted to the receiver. Since state-of-the-art scalable bit-streams offer a plethora of spatial, temporal and SNR resolutions, the number of keys to communicate for the various resolutions can be very large and variable. In this paper, two methods are proposed to significantly reduce the number of keys to be communicated by using dependent keys related by one-way functions. The first method is based on using separate one-way function chains for the keys for each scalability dimension. The second method allows simultaneous key progression along multiple dimensions by using special types of quasi-commutative one-way functions called one-way accumulators. The resultant one-way structure is named the accumulator mesh. Debargha Mukherjee, Mihaela van der Schaar |
ICIP (2) | 1 |
| 2005 | A framework for fully format-independent adaptation of scalable bit streamsabstractThis paper presents a framework for modeling and adaptation of arbitrary scalable multimedia bit streams in a manner that is fully format agnostic ("universal transcoding"). It is entitled structured scalable metaformats (SSM), and is based on establishing a universal model for all scalable bit streams, which in turn allows a compact specification of the set of allowed adaptations, and also how the bit stream adaptation is performed. In order to enable format-agnostic adaptation, two elements are standardized: 1) metadata associated with a scalable bit stream conveying the parameters of its SSM model, and how it is to be manipulated to obtain various adapted versions, as well as information that allows making appropriate adaptation decisions using only the model parameters and 2) a specification of outbound network and terminal constraints, for real-time decisions based on both user preferences and network conditions. By interpreting the descriptions, a universal adaptation engine can adapt the content to suit the specified needs and preferences of recipients, without knowledge of the specifics of the content, its encoding and encryption. With universal adaptation, different adaptation infrastructures are no longer needed for different types of scalable media, eliminating the high costs of traditional transcoding. Many of the SSM concepts have been adopted into the MPEG-21 Part 7 standard entitled Digital Item Adaptation. Debargha Mukherjee, Amir Said, Sam Liu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2005 | Optimal adaptation decision-taking for terminal and network quality-of-serviceabstractIn order to cater to the diversity of terminals and networks, efficient, and flexible adaptation of multimedia content in the delivery path to end consumers is required. To this end, it is necessary to associate the content with metadata that provides the relationship between feasible adaptation choices and various media characteristics obtained as a function of these choices. Furthermore, adaptation is driven by specification of terminal, network, user preference or rights based constraints on media characteristics that are to be satisfied by the adaptation process. Using the metadata and the constraint specification, an adaptation engine can take an appropriate decision for adaptation, efficiently and flexibly. MPEG-21 Part 7 entitled Digital Item Adaptation standardizes among other things the metadata and constraint specifications that act as interfaces to the decision-taking component of an adaptation engine. This paper presents the concepts behind these tools in the standard, shows universal methods based on pattern search to process the information in the tools to make decisions, and presents some adaptation use cases where these tools can be used. Debargha Mukherjee, E. Delfosse, Jae-Gon Kim |
IEEE Trans. Multim. | 1 |
| 2004 | Format-independent scalable bit-stream adaptation using mpeg-21 dia
Debargha Mukherjee, Geraldine Kuo, Shih-Ta Hsiang, Sam Liu, Amir Said |
ICIP | 1 |
| 2004 | Mediabeads: an architecture for path-enhanced media applicationsabstractTagging digital media, such as photos and videos, with capture time and location information has previously been proposed to enhance its organization and presentation. We believe that the full path traveled during media capture, rather than just the media capture locations, provides a much richer context for understanding and "re-living" a trip experience, and offers many possibilities for novel applications. We introduce the concept of path-enhanced media, in which media is associated and stored together with a densely sampled path in time and space and we present the MediaBeads architecture for capturing, representing, browsing, editing, presenting, and searching this data. The architecture includes, among other things, novel data representations, new algorithms for automatically building movie-like presentations of trips, and novel search applications. Michael Harville, Ramin Samadani, Daniel Tretter, Debargha Mukherjee, Ullas Gargi, N. Chang |
ICME | 4 |
| 2004 | PathMarker: systems for capturing tripsabstractCentral to capturing a trip is knowing where you were, and when you were there. Combining continuous path data with media (path-enhanced media or PEM) offers substantial advantages over the previous approach of tagging individual media with time and location. Prototype systems, collectively called PathMarker, are used for gathering, editing, presenting and browsing PEM. We have developed: (1) a methodology for gathering PEM with off-the-shelf hardware; (2) software for automatic conversion of the raw path data and media into an application independent XML representation; (3) two example PEM applications. The first application provides map-overlaid trip editing, presentation and browsing. The second application provides a 3D immersive environment with digital elevation maps for automatic trip flybys and for browsing. Experience with a number of recorded trips confirms that PathMarker systems seem to capture the essence of a trip. Ramin Samadani, Debargha Mukherjee, Ullas Gargi, N. Chang, Daniel Tretter, Michael Harville |
ICME | 2 |
| 2003 | Fully scalable video transmission using the SSM adaptation framework
Debargha Mukherjee, Peisong Chen, Shih-Ta Hsiang, John W. Woods, Amir Said |
VCIP | 1 |
| 2003 | Vector SPIHT for embedded wavelet video and image codingabstractThe set partitioning in hierarchical trees (SPIHT) approach for still-image compression proposed by Said and Pearlman (1996) is one of the most efficient embedded monochrome image compression schemes known to date. The algorithm relies on a very efficient scanning and bit-allocation scheme for quantizing the coefficients obtained by a wavelet decomposition of an image. In this paper, we adopt this approach to scan groups (vectors) of wavelet coefficients, and use successive refinement vector quantization (VQ) techniques with staggered bit-allocation to quantize the groups at once. The scheme is named vector SPIHT (VSPIHT). We present discussions on possible models for the distributions of the coefficient vectors, and show how trained classified tree-multistage VQ techniques can be used to efficiently quantize them. Extensive coding results comparing VSPIHT to scalar SPIHT in the mean-squared-error sense, are presented for monochrome images. VSPIHT is found to yield superior performance for most images, especially those with high detail content. The method is also applied to color video coding, where a partially scalable bitstream is generated. We present the coding results on QCIF sequences as compared against H.263. Debargha Mukherjee, Sanjit K. Mitra |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Low complexity guaranteed fit compound document compressionabstractWe propose a new, very low complexity, single-pass algorithm for compression of continuous tone compound documents, known as GRAFIT (GuaRAnteed FIT), that can guarantee a minimum compression ratio of as much as 12:1 and even more, for all images in a single pass, while maintaining visually lossless quality when reproduced at resolution 300 dpi or more. The compression ratio is guaranteed in a single pass irrespective of the image being compressed. The complexity of the proposed encoder and decoder is orders of magnitude smaller than all image compression algorithms known today. For electronic compound documents, text is always compressed losslessly, and depending on the type of image, the actual compression ratio achieved may be as high as 200:1 or more. For photographic images, while the compression performance is inferior to DCT or wavelet coders, for documents at resolution 300 dpi and above, the quality is still visually lossless. Overall, performance of GRAFIT is highly competitive with the more expensive algorithms like JPEG2000, JPEG, or JPEG-LS and this performance is achieved in a single pass at much lower cost in both software and hardware. Debargha Mukherjee, Christos Chrysafis, Amir Said |
ICIP (1) | 1 |
| 2002 | JPEG2000-matched MRC compression of compound documentsabstractThe mixed raster content (MRC) ITU document compression standard (T.44) specifies a multilayer decomposition model for compound documents into two contone image layers and a binary mask layer for independent compression. While T.44 does not recommend any procedure for decomposition, it does specify a set of allowable layer codecs to be used after decomposition. While T.44 only allows older standardized codecs such as JPEG/JBIG/G3/G4, higher compression could be achieved if newer contone and bi-level compression standards such as JPEG2000/JBIG2 were used instead. We present an MRC compound document codec using JPEG2000 as the image layer codec and a layer decomposition scheme matched to JPEG2000 for efficient compression. JBIG still codes the mask. Noise removal routines enable efficient coding of scanned documents along with electronic ones. Resolution scalable decoding features are also implemented. The segmentation mask, obtained from layer decomposition, serves to separate text and other features. Debargha Mukherjee, Christos Chrysafis, Amir Said |
ICIP (3) | 1 |
| 2002 | Image resizing in the compressed domain using subband DCTabstractResizing of digital images is needed in various applications, such as transmission of images over communication channels varying widely in their bandwidths, display at different resolutions depending on the resolution of a display device, etc. In this work, we propose a modification of a recently proposed elegant image resizing algorithm by Dugad and Ahuja (2001). We have also extended their approach and our modified versions to color images and studied their performance at different levels of compression for an image. Our proposed modified algorithms, in general, perform better than the earlier method in most cases. Though there is a marginal increase in the computation required in image-halving, the computation overhead of the proposed modification is higher compared to the Dugad-Ahuja algorithm in the case of doubling the images. Debargha Mukherjee, Sanjit K. Mitra |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Successive refinement lattice vector quantizationabstractLattice Vector quantization (LVQ) solves the complexity problem of LBG based vector quantizers, yielding very general codebooks. However, a single stage LVQ, when applied to high resolution quantization of a vector, may result in very large and unwieldy indices, making it unsuitable for applications requiring successive refinement. The goal of this work is to develop a unified framework for progressive uniform quantization of vectors without having to sacrifice the mean- squared-error advantage of lattice quantization. A successive refinement uniform vector quantization methodology is developed, where the codebooks in successive stages are all lattice codebooks, each in the shape of the Voronoi regions of the lattice at the previous stage. Such Voronoi shaped geometric lattice codebooks are named Voronoi lattice VQs (VLVQ). Measures of efficiency of successive refinement are developed based on the entropy of the indices transmitted by the VLVQs. Additionally, a constructive method for asymptotically optimal uniform quantization is developed using tree-structured subset VLVQs in conjunction with entropy coding. The methodology developed here essentially yields the optimal vector counterpart of scalar "bitplane-wise" refinement. Unfortunately it is not as trivial to implement as in the scalar case. Furthermore, the benefits of asymptotic optimality in tree-structured subset VLVQs remain elusive in practical nonasymptotic situations. Nevertheless, because scalar bitplane- wise refinement is extensively used in modern wavelet image coders, we have applied the VLVQ techniques to successively refine vectors of wavelet coefficients in the vector set-partitioning (VSPIHT) framework. The results are compared against SPIHT and the previous successive approximation wavelet vector quantization (SA-W-VQ) results of Sampson, da Silva and Ghanbari. Debargha Mukherjee, Sanjit K. Mitra |
IEEE Trans. Image Process. | 1 |
| 2001 | Linear-translate constrained storage VQ for VSPIHT wavelet image compressionabstractA new constrained storage VQ (CSVQ) structure based on linear transforms and translates of a common root codebook is proposed. The new VQ structure, named LT-CSVQ (linear translate CSVQ) acts as a building block for multistage VQ implementations (LT-CS-MSVQ), and significantly reduces storage requirements from that required in tree-multistage VQ implementations. LT-CS-MSVQ is most appropriate for medium rate multistage VQ implementations, and is applied to the previously proposed vector enhancement of Said and Pearlman's (1996) set partitioning in hierarchical trees (SPIHT) image coder named VSPIHT. Debargha Mukherjee, Sanjit K. Mitra |
ICASSP | 1 |
| 2001 | A new algorithm based on saturation and desaturation in the xy chromaticity diagram for enhancement and re-rendition of color imagesabstractThis paper presents a new algorithm for color contrast enhancement which adopts the xy chromaticity diagram and consists of two steps. All the "chromatic" colors are first maximally saturated within a certain gamut; the opportune desaturation operation which follows is based on the center of gravity law for color mixture. This prevents the introduction of unnatural colors and, in combination with a suitable manipulation of the brightness value, turns out to increase the image sharpness and to provide more appealing results Our technique can also be applied to re-rendition a color image since it can easily simulate the change of the illuminant(s) of the scene portrayed in the image, this could find useful applications in image and video editing. Some examples of the performance of the algorithm are reported and discussed. Luca Lucchese, Debargha Mukherjee, Sanjit K. Mitra |
ICIP (2) | 2 |
| 2001 | JPEG-matched MRC compression of compound documentsabstractMixed raster content (MRC) is an ITU document compression standard (T.44) specifying both a model for multilayer representation of a compound document, and a set of allowable standardized coders for the individual layers. The model requires decomposition of a document into two image layers and a binary mask layer, but the standard does not recommend any procedure for this task. For best compression results, the decomposition method should be optimized for the layer encoders. In this paper, a high performance MRC compound document codec is presented, where the layer decomposition scheme is matched to the JPEG encoder with arithmetic coding for the foreground and background image layers. JBIG is used to code the mask layer. Integrated noise removal routines enable handling of scanned documents along with electronic ones. Resolution scalable decoding features are also implemented. The page segmenter yields a segmentation mask, which serves to separate text and other features. Debargha Mukherjee, Nasir Memon, Amir Said |
ICIP (3) | 1 |
| 2000 | A vector set partitioning noisy channel image coder with unequal error protectionabstractA vector enhancement of Said and Pearlman's set partitioning in hierarchical trees (SPIHT) methodology, named VSPIHT, has been proposed for embedded wavelet image compression. A major advantage of vector-based embedded coding with fixed length VQs over scalar embedded coding is its superior robustness to noise. We show that vector set partitioning can effectively alter the balance of bits in the bit steam so that significantly fewer critical bits carrying significant information are transmitted, thereby improving inherent noise resilience. For low noise channels, the critical bits are protected, while the degradation in reconstruction quality caused by errors in noncritical quantization information, can be reduced by appropriate VQ indexing, or designing channel optimized VQs for the successive refinement systems. For very noisy channels, unequal error protection to the critical and noncritical bits with either block codes or convolution codes are used. Additionally, the error protection for the critical bits is changed from pass to pass. A buffering mechanism is used to produce an unequally protected bit stream without sacrificing the embedding property. Extensive simulation results are presented for noisy channels, including bursty channels. Debargha Mukherjee, Sanjit K. Mitra |
IEEE J. Sel. Areas Commun. | 1 |
| 2000 | A source and channel-coding framework for vector-based data hiding in videoabstractDigital data hiding is a technology being developed for multimedia services, where significant amounts of secure data is invisibly hidden inside a host data source by the owner, for retrieval only by those authorized. The hidden data should be recoverable even after the host has undergone standard transformations, such as compression. In this paper, we present a source and channel coding framework for data hiding, allowing any tradeoff between the visibility of distortions introduced, the amount of data embedded, and the degree of robustness to noise. The secure data is source coded by vector quantization, and the indices obtained in the process are embedded in the host video using orthogonal transform domain vector perturbations. Transform coefficients of the host are grouped into vectors and perturbed using noise-resilient channel codes derived from multidimensional lattices. The perturbations are constrained by a maximum allowable mean-squared error that can be introduced in the host. Channel-optimized VQ can be used for increased robustness to noise. The generic approach is readily adapted to make retrieval possible for applications where the original host is not available to the retriever. The secure data in our implementations are low spatial and temporal resolution video, and sampled speech, while the host data is QCIF video. The host video with the embedded data is H.263 compressed, before attempting retrieval of the hidden video and speech from the reconstructed video. The quality of the extracted video and speech is shown for varying compression ratios of the host video. Debargha Mukherjee, Jong Jin Chae, Sanjit K. Mitra |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | Generalized variable dimensional set partitioning for embedded wavelet image compressionabstractA vector enhancement of Said and Pearlman's (1996) set partitioning in hierarchical trees (SPIHT) methodology, named VSPIHT, has recently been proposed for embedded wavelet image compression. While the VSPIHT algorithm works better than scalar SPIHT for most images, a common vector dimension to use for coding an entire image may not be optimal. Since statistics vary widely within an image, a greater efficiency can be achieved if different vector dimensions are used for coding the wavelet coefficients from different portions of the image. We present a generalized methodology for developing a variable dimensional set partitioning coder, where different parts of an image may be coded in different vectoring modes, with different scale factors, and up to different number of passes. A Lagrangian rate-distortion criterion is used to make the optimum coding choices. Coding passes are made jointly for the vectoring modes to produce an embedded bitstream. Debargha Mukherjee, Sanjit K. Mitra |
ICASSP | 1 |
| 1998 | Vector set partitioning with classified successive refinement VQ for embedded wavelet image and video codingabstractThe set partitioning in hierarchical trees (SPIHT) approach for still image compression proposed by Said and Pearlman (see IEEE Transactions on Circuits and Systems for Video Technology, vol.6, no.3, 1996) is one of the most efficient embedded gray image compression schemes to date. The algorithm relies on a very efficient scanning cum bit-allocation scheme for quantizing the coefficients obtained by a wavelet decomposition of an image. We adopt this scheme to scan vectors of wavelet coefficients, and use successive refinement VQ techniques with staggered bit-allocation to quantize several wavelet coefficients at once. The new scheme is named VSPIHT (vector SPIHT). We present some coding results comparing VSPIHT to the scalar counterpart in the mean-squared-error sense. The method readily generalizes to color images and video where the vector-based approach makes more sense. We present the coding results on intra frames of QCIF sequences as compared against H.263. Debargha Mukherjee, Sanjit K. Mitra |
ICASSP | 1 |
| 1998 | Color Image Embedding using Multidimensional Lattice StructuresabstractThis paper describes a robust data embedding scheme which uses noise resilient channel codes based on a multidimensional lattice structure. Compared to prior work in digital watermarking, the proposed scheme can handle a significantly larger quantity of signature data such as gray-scale or color images. A trade-off between the quantity of hidden data and the quality of the watermarked image is achieved by varying the number of quantization levels for the signature, and a scale factor for data embedding. Experimental results on signature recovery from JPEG compressed watermarked images show that good quality reconstruction is possible even when the images are lossy compressed by as much as 85%. Potential applications of this method include, in addition to watermarking, digital data hiding for security and for bit stream control and manipulation. Jong Jin Chae, Debargha Mukherjee, B. S. Manjunath |
ICIP (1) | 2 |
| 1998 | A Source and Channel Coding Approach to Data Hiding with Application to Hiding Speech in VideoabstractDigital data hiding is a technology being developed for multimedia services, where non-trivial amounts of signature data is invisibly hidden inside a host data source by the owner before the latter is freely distributed. Only those authorized can recover the hidden data from the host, even after the latter has undergone standard transformations such as compression. We adopt a quantitative source and channel coding approach to hiding large amounts of compressible signature data inside the raw host. The signature data is source coded by vector quantization, and the indices are embedded in the host by perturbing it using orthogonal transform domain vector perturbations. The transform coefficients of the parent data are grouped into vectors, and the vectors are perturbed using noise-resilient channel codes derived from multidimensional lattices. The perturbations are constrained by a maximum allowable mean-squared error that can be introduced in the host. The generic approach is readily adapted to make retrieval possible even for applications where the original host is not available to the retriever. This scheme is applied to hiding speech in video. The host video is wavelet transformed frame by frame, and vectors of coefficients are perturbed using lattice channel codes to represent hidden vector quantized speech. The embedded video is subjected to H.263 compression before retrieving the hidden speech from it. The retrieved speech is intelligible even with large compression ratios of the host video. Debargha Mukherjee, Jong Jin Chae, Sanjit K. Mitra |
ICIP (1) | 1 |
| 1998 | Vector Set-Partitioning with Successive Refinement Voronoi Lattice VQ for Embedded Wavelet Image CodingabstractWhile lattice vector quantization (LVQ) can solve the complexity problem of LBG based vector quantizers, and also yield very general codebooks, a single stage lattice VQ, when applied to high variance vectors result in very large and unwieldy indices, making it unsuitable for applications requiring successive refinement. The goal of this work is to develop a unified framework for progressive uniform quantization of vectors, without having to sacrifice the mean-squared-error advantage of lattice quantization. A successive refinement uniform vector quantization paradigm is developed, where the codebooks in successive stages are all lattice codebooks, each in the shape of the Voronoi region of the lattice at the previous stage. The Voronoi shaped lattice codebook at each stage is called Voronoi lattice VQ (VLVQ). Measures of efficiency of successive refinement are developed. The developed methodology is applied to successively refine vectors of wavelet coefficients in the vector set-partitioning (VSPIHT) framework to obtain an embedded bitstream. The results are compared against the previous successive approximation wavelet vector quantization (SA-W-VQ) results of Sampson, da Silva, and Ghanbari (see IEEE Trans. Image Processing, vol.5, no.2, p.299-310, 1996) for image coding. Debargha Mukherjee, Sanjit K. Mitra |
ICIP (1) | 1 |
| 1998 | Arithmetic coded vector SPIHT with classified tree-multistage VQ for color image codingabstractA vector extension of the set partitioning in hierarchical trees (SPIHT) algorithm, named vector-SPIHT (VSPIHT), using trained classified successive refinement VQ, has recently been proposed. In this work, vector set-partitioning is applied to multispectral image compression, in particular to 24-bit color images. Since the individual spectral components are sufficiently correlated, VSPIHT can effectively exploit both the inter-component redundancy as well as the spatial redundancy within each subband of each component, to yield performance superior to separate scalar SPIHT coding of each component. Adaptive arithmetic coding of the first stage VQ index for each class, as well as the significance information, further improves the performance. Coding results demonstrate that the vector-based approach for color images significantly outperforms the scalar counterpart in the mean-squared-error sense. Debargha Mukherjee, Sanjit K. Mitra |
MMSP | 1 |
| 1997 | Combined Mode Selection and Macroblock Quantization Step Adaptation for the H.263 Video EncoderabstractThe low bitrate video coding standard H.263 allows four modes of coding, and 31 quantization parameters for each macroblock in a sequence. A new quantization parameter can be set at the beginning of each group of blocks (GOB), and then adapted from macroblock to macroblock within the GOB in a constrained manner. In this work, for every GOB in a frame, an m-best search scheme is employed through a trellis to find the best modes and the best quantizer step changes, in a rate-distortion sense, for both when the starting quantization parameter is predetermined, and when it is to be selected optimally as well. Thus the algorithm performs the most generic kind of optimization at the macroblock level for the P-frames of H.263. It is demonstrated that a relatively low search depth is sufficient to produce coding results superior to established standard-compatible coders. Debargha Mukherjee, Sanjit K. Mitra |
ICIP (2) | 1 |
| 1996 | Subband DCT: definition, analysis, and applicationsabstractThe discrete cosine transform (DCT) is well known for its highly efficient coding performance and is widely used in many image compression applications. However, in low bit rate coding, it produces undesirable block artifacts that are visually not pleasing. In addition, in many practical applications, faster computation and easier VLSI implementation of DCT coefficients are also important issues. The removal of the block artifacts and faster DCT computation are therefore of practical interest. In this paper, we investigate a modified DCT computation scheme, to be called the subband DCT (SB-DCT), that provides a simple, efficient solution to the reduction of the block artifacts while achieving faster computation. We have applied the new approach for the low bit rate coding and decoding of images. Simulation results on real images have verified the improved performance obtained using the proposed method over the standard JPEG method. Sung-Hwan Jung, Sanjit K. Mitra, Debargha Mukherjee |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1996 | Rate-distortion optimized mode selection for very low bit rate video coding and the emerging H.263 standardabstractThis paper addresses the problem of encoder optimization in a macroblock-based multimode video compression system. An efficient solution is proposed in which, for a given image region, the optimum combination of macroblock modes and the associated mode parameters are jointly selected so as to minimize the overall distortion for a given bit-rate budget. Conditions for optimizing the encoder operation are derived within a rate-constrained product code framework using a Lagrangian formulation. The instantaneous rate of the encoder is controlled by a single Lagrange multiplier that makes the method amenable to mobile wireless networks with time-varying capacity. When rate and distortion dependencies are introduced between adjacent blocks (as is the case when the motion vectors are differentially encoded and/or overlapped block motion compensation is employed), the ensuing encoder complexity is surmounted using dynamic programming. Due to the generic nature of the algorithm, it can be successfully applied to the problem of encoder control in numerous video coding standards, including H.261, MPEG-1, and MPEG-2. Moreover, the strategy is especially relevant for very low bit rate coding over wireless communication channels where the low dimensionality of the images associated with these bit rates makes real-time implementation very feasible. Accordingly, in this paper, the method is successfully applied to the emerging H.263 video coding standard with excellent results at rates as low as 8.0 Kb per second. Direct comparisons with the H.263 test model, TMN5, demonstrate that gains in peak signal-to-noise ratios (PSNR) are achievable over a wide range of rates. Thomas Wiegand 0001, Michael Lightstone, Debargha Mukherjee, T. George Campbell, Sanjit K. Mitra |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1995 | Adaptive Neighborhood Extended Contrast Enhancement and Its Modifications
Debargha Mukherjee, Biswanath N. Chatterji |
CVGIP Graph. Model. Image Process. | 1 |