Kiran M. Misra

dblp:17/10699 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0002-0551-2705ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 9 first-author · 8 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Ultra-Low Complexity Neural Networks for Next Generation Video Decoding
abstract
We consider the problem of embedding a neural network directly into a video decoder. This requires a design with complexity suitable for implementation on mobile and power constrained devices. To achieve this goal, we explored Multi-scale CNN (MSCNN) design in [1]. In this paper, we improve the design to support super resolution spatial scale factors SF==(1.5×, 2×, 3×, 4×, 6×) by modifying the polyphase filter (Figure 1a) that generates an upsampled output using g(scale) phases and stride of Sscale. When SF= 1.5 ×, g(scale, Sscale) = (9,2); Otherwise it is (scale2,1). gsG, kK, and sS denote channel group size of G, kernel size of K×K, and stride of S. To reduce per-pixel Multiply-Accumulates (MACs), the 3×1 and 1×3 convolutional layers use Canonical Polyadic (CP) decomposition and reduced channel count. These changes reduce MACs/pixel from 1,924 in [1] to 1,192 to 584. Figure 1b, shows the placement of MSCNN in AVM [2]. We code 4K video, using AOMedia's Adaptive Streaming (AS) test conditions and compare MSCNN versus following resampler combinations: Downsampling - [L5: Lanczos(5), L6: Lanczos(6)]; Upsampling - [L5, L6, BL: Bilinear, BC: Bicubic]. We observe MSCNN provides on average 30.4% rate saving.
Kiran M. Misra, Shashwat Ranjan Chaurasia, C. Andrew Segall, Byeongdoo Choi
DCC1
2025 Efficient Random Access Method Using Seed and Inter-Key Frames for Next Generation Video Codec
abstract
Random access points are a key property of a video coding system. These points indicate where a decoder can start decoding, and they traditionally correspond to frames that are not predicted from previous data. In this paper, we revisit the random access problem in the context of modern video streaming and over-the-top transport systems. We propose that these systems employ an alternative approach that relies on "seed" frames that are periodically provided to the decoder. These frames persist in decoder memory and can be used for prediction of each random access point. Experimental results show the efficacy of the proposed approach. Specifically, we observe a 6.50% reduction in bit-rate when using the AOMedia common test conditions, an 18.38% reduction when emulating live sports events, and a 35.03% improvement for security applications when measured using VMAF.
Byeongdoo Choi, C. Andrew Segall, Kiran M. Misra
ICIP3
2023 Multiscale convolutional neural networks for in-loop video restoration
abstract
Incorporating neural networks into a video codec as an in-loop filter has been shown to provide significant improvements in coding efficiency. Unfortunately, the computational complexity associated with the neural network, specifically the number of multiply-accumulate (MAC) operations, makes these approaches intractable in practice. In this paper, we consider using a multiscale approach to reduce complexity while maintaining coding efficiency. Experimental results demonstrate a 5.4× reduction in MAC operations while achieving an average bit rate savings of 6.4% and 6.3% for all intra and random access coding, respectively, when compared to the evolving AV2 standard. Ablation studies are also provided and show that the approach achieves all but 0.2% of the coding efficiency of full resolution processing.
Kiran M. Misra, C. Andrew Segall, Byeongdoo Choi
DCC1
2023 Reduced Complexity Multiscale CNN for in-Loop Video Restoration
abstract
Convolutional neural networks (CNNs) have shown promising improvements in video coding efficiency when included in traditional block-based codecs as a loop filter. Unfortunately, these coding gains are often accompanied by significant increases in complexity, measured by the number of multiply-accumulate (MAC) operations, that make them intractable in practice. As a result, there is considerable interest in reducing complexity for these CNN-based approaches. In previous work, we have shown that multiscale CNNs provide a path to reduce the associated MAC count. In this paper, we extend our work to consider channel grouping, spatial support limitations and shallower network depths to further reduce the MAC count of these multi-scale architectures. We demonstrate that the method can achieve an average VMAF bitrate reduction of 6.1% and 2.6% for all intra and random-access coding respectively, when compared to the evolving AV2 standard. Complexity is reduced to 1.85k MACs per pixel, which is a 390× reduction over previously published results.
Kiran M. Misra, C. Andrew Segall, Byeongdoo Choi
ICIP1
2022 Efficient Feature Compression for the Object Tracking Task
abstract
In object tracking systems, often clients capture video, encode it and transmit it to a server that performs the actual machine task. In this paper we propose an alternative architecture, where we instead transmit features to the server. Specifically, we partition the Joint Detection and Embedding (JDE) person tracking network into client and server side sub-networks and code the intermediate tensors i.e. features. The features are compressed for transmission using a Deep Neural Network (DNN) we design and train specifically for carrying out the tracking task. The DNN uses trainable non-uniform quantizers, conditional probability estimators, hierarchical coding; concepts that have been used in the past for neural networks based image and video compression. Additionally, the DNN includes a novel parameterized dual-path layer that comprises of an autoencoder in one path and a convolution layer in the other. The tensor output by each path is added before being consumed by subsequent layers. The parameter value for this dual-path layer controls the output channel count and correspondingly the bitrate of transmitted bitstream. We demonstrate that our model improves coding efficiency by 43.67% over state-of-the-art Versatile Video Coding standard that codes the source video in pixel domain.
Robert Henzel, Kiran M. Misra, Tianying Ji
ICIP2
2022 Video Feature Compression for Machine Tasks
abstract
We consider the problem of transmitting video from a remote device to a cloud-based classification system in a bandwidth limited network. Our focus is on developing an end-to-end system that extracts features from the video data and compresses these features for transmission. In this paper, we consider approaches that operate on each video frame independently as well as exploiting the temporal correlation between frames. In both cases, the transmitted features can be used for object detection and instance segmentation tasks using existing, pre-trained networks. Results show the efficacy of the approach with improvements in coding efficiency ranging from 46.3% to 92.8% when compared to compressing the video data using state-of-the-art video compression standards.
Kiran M. Misra, Tianying Ji, C. Andrew Segall, Frank Bossen
ICME1
2021 Deblocking filtering in VVC
abstract
The paper describes the novel aspects of deblocking filter in VVC. We demonstrate the insufficiency of HEVC deblocking filter when employing larger transform block sizes and describe the design changes made to enable reduction of the resulting blocking artifacts for both luma and chroma, while also allowing for parallel friendly processing. Additionally, VVC deblocking includes filtering based on local luma level to address blocking artifacts observed in high dynamic range content. The paper also describes modifications made in VVC deblocking to address blocking artifacts introduced by additional coding tools present in VVC (but not in HEVC). The additional coding tools result in prediction and transform block boundaries at disparate locations in a coded picture. Similar to previous generation of video coding standards, the deblocking in VVC, adaptively filters samples located at the block boundaries based on coding modes and local spatial activity.
Kenneth Andersson, Kiran M. Misra, Masaru Ikeda, Dmytro Rusanovskyy, Shunsuke Iwamura
PCS2
2021 VVC In-Loop Filters
abstract
This paper presents an overview of the technologies for in-loop processing and filtering in the Versatile Video Coding (VVC) standard. These processes comprise luma mapping with chroma scaling, deblocking filter, sample adaptive offset, adaptive loop filter and cross-component adaptive loop filter. They are qualified as “in-loop” because they are applied inside the encoding and decoding loops, before storing the pictures in the decoded picture buffer. The filters are complementary and address different purposes. Luma mapping with chroma scaling aims at adaptively modifying the coded samples distribution for improved coding efficiency. The deblocking filter aims at reducing blocking discontinuities. Sample adaptive offset mostly aims at reducing artifacts resulting from the quantization of transform coefficients. Adaptive loop filter and cross-component adaptive loop filter are adaptive filters enabling to enhance the reconstructed signal, using for instance Wiener-filter encoding approaches. The paper provides an overview of the in-loop filtering process and a detailed description of the filtering algorithms. Objective compression efficiency results are provided for each filter, with indication of cumulative coding gains. Subjective benefits are illustrated. Implementation issues considered during the design of the VVC in-loop filters are also discussed.
Marta Karczewicz, Jonathan Taquet, Ching-Yeh Chen, Kiran M. Misra, Kenneth Andersson, Peng Yin 0002, Taoran Lu, Edouard François, Jie Chen 0006
IEEE Trans. Circuits Syst. Video Technol.5
2020 Tools for Video Coding Beyond HEVC: Flexible Partitioning, Motion Vector Coding, Luma Adaptive Quantization, and Improved Deblocking
abstract
This paper provides a description and analysis of technology included in two contributions to the Call for Proposals for Video Compression with Capability beyond HEVC. This Call for Proposals was issued jointly by the Moving Picture Experts Group (MPEG) and the Video Coding Experts Group (VCEG). The contribution emphasized a flexible and rectangular partitioning structure, which was combined with both new and existing coding tools. New coding tools included methods for improved motion vector coding, quantization signaling, and deblocking; while existing tools were largely methods studied in the Joint Exploration Model software. Results show the efficacy of the approach. Using the sequences and test conditions defined in the Call for Proposals, the described approach provided an average bitrate reduction, relative to an HEVC anchor, of 41.2% and 35.7% for 4K and HD test sequences, respectively. Moreover, the method achieved a compression performance of 34.3% and 32.2% for high dynamic range content using the perceptual quantizer (PQ) and Hybrid-Log Gamma transfer functions, respectively.
Kiran M. Misra, C. Andrew Segall, Frank Bossen
IEEE Trans. Circuits Syst. Video Technol.1
2019 Enhanced Compression beyond HEVC for Next Generation Content
abstract
The Joint Video Experts Team recently evaluated technology in response to a Call for Proposals for Video Compression with Capability beyond HEVC. A number of proposed solutions were evaluated, with a sub-set demonstrating the potential to reduce bit-rates by over 40% compared to HEVC. This paper presents the author's contributions to one of these proposals. The proposal emphasized a flexible, rectangular partitioning structure that was combined with new coding tools, including improved motion vector coding and quantization signaling methods. Results show the efficacy of the approach. Using the evaluation procedure defined in the Call, the described approach provides coding gains relative to an HEVC anchor of 41.2% and 35.7% for 4K-SDR and HD-SDR sequences, respectively, using the random access configuration; 29.0% for HD-SDR sequences using a low delay configuration, and gains of 34.3% and 32.2% for PQ-HDR and HLG-HDR sequences, respectively, using a random access configuration.
Kiran M. Misra, C. Andrew Segall, Weijia Zhu, Byeongdoo Choi, Frank Bossen, Phil Cowan
DCC1
2019 On Cross Component Adaptive Loop Filter for Video Compression
abstract
This paper introduces a video coding tool called the Cross Component Adaptive Loop Filter (CC-ALF). The goal of the tool is to improve chroma fidelity, and this is achieved by exploiting correlations between the luma and chroma channels during the loop filtering process. The tool analyzes the luma channel to determine a chroma residual that is added to a previously reconstructed chroma signal, where the analysis consists of applying a linear filter to the luma information. Adaptation of the filter is achieved by signaling the filter coefficients in the bit-stream and, furthermore, selectively enabling or disabling the process spatially across a picture. Results show the efficacy of the approach, where the performance is evaluated using the Common Test Conditions (CTC) defined by Joint Video Experts Team (JVET). The evaluations show that CCALF provides average chroma Bjøntegaard Delta improvement of 7.6%, 13.8% and 17.8% for all intra, random access, and low delay configurations. Additionally, an approach to shift these chroma gains to luma gains is considered and results in an average luma BD bitrate improvement of 0.9%.
Kiran M. Misra, Frank Bossen, C. Andrew Segall
PCS1
2017 Spatially Scalable HEVC for Layered Division Multiplexing in Broadcast
abstract
Recent broadcast standards support Layered Division Multiplexing (LDM) to achieve graceful degradation as signal quality degrades at the receiver. LDM is accomplished by using different constellations within the same Radio Frequency (RF) spectrum. LDM thus enables delivering multiple service tiers in a single broadcast channel. LDM when used in conjunction with scalable source coding codecs such as the Scalable extension of High Efficiency Video Coding (SHVC), further helps improve overall spectrum utilization and efficiency. In this paper we investigate a 2-tier broadcast LDM based service with one service tier aimed at lower video resolution such as 540p, 720p, 1080p for a mobile receiver (smaller/indoor antenna) and the other service tier targeting twice the video resolution of the lower tier, for stationary receivers (larger/outdoor antenna). The primary contribution of this paper is to identify 2-tier transmission configurations of interest to broadcasters and compare the spectrum and bitrate coding efficiency gains of an SHVC-based multi-tier service versus a simulcast (single layer) based multi-tier service for an Advanced Television Systems Committee (ATSC) 3.0 transmission system. Bitrate savings ranging from 38% and 57% is observed for the SHVC based layered system. For large coverage and pedestrian with a receiver test scenarios, channel utilization savings ranging from 23% to 46% is observed. For mobile and tablet in bedroom scenarios a smaller broadcast bandwidth savings ranging from 6% to 9% is observed.
Kiran M. Misra, C. Andrew Segall, Jie Zhao 0007, Seung-Hwan Kim 0001, Joan Llach, Alan Stein, John Stewart, Hendry, Ye-Kui Wang, Yan Ye 0003
DCC1
2016 Quantizer noise equalization between electro-optical transfer functions
abstract
We consider the problem of equalizing the quantization noise between two video encoders operating with video stored using different transfer function. Here, a transfer function is the mapping between linear light and the quantized code words and includes, for example, the well-known “gamma curve” used for most video applications. Our motivation for studying this problem is that the solution enables the direct conversation of a quantization strategy developed for a legacy, standard dynamic range transfer function to be easily mapped to a quantization strategy for a new, high dynamic range transfer function. This is especially timely given the beginning deployment of high dynamic range services throughout the world that use the BT-2020/SMPTE ST 2084 transfer function and not the BT.709/BT.1886 transfer function historically employed for coding high definition digital video. The paper derives a model for the needed relationship. Results show the benefit of the approach.
C. Andrew Segall, Jie Zhao 0007, Seung-Hwan Kim 0001, Kiran M. Misra
PCS4
2016 High quality HDR video compression using HEVC main 10 profile
abstract
This paper describes high-quality compression of high dynamic range (HDR) video using existing tools such as the HEVC Main 10 profile, the SMPTE ST 2084 (PQ) transfer function, and the BT.2020 non-constant luminance Y'CbCr color representation. First, we present novel mathematical bounds that reduce complexity of luminance-preserving subsampling (luma adjustment). A nested look-up table allows for further speedup. Second, an adaptive QP scheme is presented that obtains a better bit allocation balance between dark and bright areas of the picture. Third, a method to control the bit allocation balance between chroma and luma by adjusting the chroma QP offset is presented. The result is a considerable increase in perceptual quality compared to the anchors used in the 2015 MPEG High Dynamic Range/Wide Color Gamut Call for Evidence. All techniques are encoder-side-only, making them compatible with a regular decoder capable of supporting HEVC Main10/PQ/BT.2020, which is already available in some TV sets on the market.
Jacob Ström, Kenneth Andersson, Martin Pettersson, Per Hermansson, Jonatan Samuelsson, C. Andrew Segall, Jie Zhao 0007, Seung-Hwan Kim 0001, Kiran M. Misra, Alexis M. Tourapis, Yeping Su, David Singer
PCS9
2013 Content-adaptive upsampling for scalable video coding
abstract
In the developing scalable extension of the HEVC/H.265 standard, a low-resolution baselayer picture may be used for predicting a higher resolution enhancement layer. This requires an upsampling process to generate the prediction, and this upsampling process is traditionally a linear and time invariant interpolator. In this paper, we consider an upsampling design that is both non-linear and content adaptive. This choice is motivated by the compression noise in the baselayer. We propose a novel approach to include content-aware filtering into the upsampling process. It has low complexity. More importantly, it has low latency and uses the same number of line buffers as a linear interpolator. Results show the efficacy of the method. Specifically, using the test model for the scalable extension of HEVC/H.265, we observe an average bit-rate reduction of 1.1%, when the change in resolution between layers is a factor of two in each dimension.
Jie Zhao 0007, Kiran M. Misra, C. Andrew Segall
PCS2
2012 On transform dynamic range in high efficiency video coding
abstract
We consider the problem of determining the dynamic range of an inverse transform. An analytic approach is described, and the resulting bound is a function of the bit-depth of the input data and the maximum L1norm of the transformation matrix. Furthermore, the analytic approach is refined to account for nonlinearities within a transform due to intermediate, integer conversions. This refinement considers the maximum error introduced by the nonlinearity on the initial, linear bound. The efficacy of the proposed tool is demonstrated by analyzing the inverse transform in the high efficiency video coding (HEVC) standard developed by the JCT-VC.
Louis Kerofsky, Kiran M. Misra, C. Andrew Segall
ICIP2
2012 Tiles for managing computational complexity of video encoding and decoding
abstract
In this paper, we introduce the concept of tiles. Tiles are incorporated into the current design of the High Efficiency Video Coding (HEVC) standard being developed by the Joint Collaborative Team on Video Coding (JCT-VC). In the design, tiles are introduced to support high-level parallelism and also to reduce on-chip memory requirements. This paper describes the tile concept and reports results due to the technique.
Arild Fuldseth, Michael Horowitz, Kiran M. Misra, C. Andrew Segall, Minhua Zhou
PCS4
2011 Periodic entropy coder initialization for wavefront decoding of video bitstream
abstract
In this paper, we outline a coding strategy that initializes the entropy coding engine of a video codec at pre-defined locations within a bit-stream. Coupled with the causal dependencies of state-of-the-art video coding systems, this enables wavefront processing of the entropy decoding and reconstruction process simultaneously. Approaches to wavefront processing have been considered by others, and those methods either address the reconstruction process solely or require transmitting image data in non-raster scan order. Here, our key contribution is that we enable simultaneous entropy/reconstruction wavefront processing while still preserving a raster scan strategy. In this paper, we describe the system, as well as different strategies for initializing the context models. The performance of the proposed methods is evaluated, and the bitrate increase is shown to be nominal.
Kiran M. Misra, Jie Zhao 0007, C. Andrew Segall
ICIP1