EDBT 2026 Demo / reviewers in the wild / expert
Urvang Joshi
dblp:182/3486
· DBLP profile ↗
16ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-9590-9505ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive AV2 In-loop Filtering via Guided Neural Model with Vectorized Quantization
Kequan Mao, Dandan Ding, Urvang Joshi, Debargha Mukherjee |
ISCAS | 3 |
| 2025 | Super Resolution-Based Video Coding via Lightweight Implicit Neural ModelingabstractThe super-resolution (SR)-based coding tool is widely employed in modern video coding standards. By encoding video frames at a reduced resolution and then restoring them to their original resolution during the in-loop filtering stage, this tool helps to further reduce the bitrate and improve the coding performance. Current video coding standards typically devise rule-based SR methods in their codecs, compromising the coding efficiency to maintain low computational complexity. As deep neural network (DNN)-based SR methods are proving more effective than rule-based approaches, this paper proposes integrating the neural SR into video codecs to enhance coding performance while minimizing the computational cost. To this end, we propose a Lightweight Implicit Neural Model (LIM). Specifically, our LIM, consisting of Lightweight Feature Aggregation Network (LFANet) and Coordinate Upsampling Network Based on B-spline Representation (CURNet), is developed to support SR-based coding at an arbitrary scale. We exemplify the proposed method on the ongoing AVM reference software and conduct extensive experiments to demonstrate its effectiveness. Compared with anchored AVM, our method improves the BD-Rate by 5.52%, which significantly outperforms state-of-the-art works. Meanwhile, its computational complexity is much lower than others, having only 22.5k parameters and 18.7k FLOPs/pixel complexity, which is attractive to real-world applications. Xianlu Bian, Dandan Ding, Urvang Joshi, Debargha Mukherjee |
DCC | 4 |
| 2025 | Extension of Semi-Decoupled Partitioning in Inter FramesabstractThe Alliance for Open Media (AOMedia) has been exploring new coding tools to enhance AV1 capabilities. Semi-Decoupled Partitioning (SDP), originally designed for intra frames in research-v2.0.0, improves coding by decoupling luma and chroma block partitioning. This study extends SDP to inter frames by introducing intra region coding, where the root node is explicitly signaled in the bitstream. Within the intra region, luma components of the intra-coded blocks can be further split, while chroma components remain unsplit. The experiments are implemented on the 8thanchor, research-v8.0.0, of AVM reference software with CTCv7, and experimental results show that the proposed method can achieve 0.12%, 2.39%, 2.58% coding gain for Y, U, and V component separately with random access configuration and 5% encoding time increase and almost no decoding time increase. Madhu Peringassery Krishnan, Shan Liu 0001, Jayasingam Adhuran, Minhao Tang, Jianle Chen, Urvang Joshi, Mohammed Golam Sarwer, Debargha Mukerjee |
ICIP | 7 |
| 2025 | AVM In-loop Filtering using Adaptive Hardware-friendly Neural Networks
Urvang Joshi, Akshaya Purohit, Debargha Mukherjee, Shan Li 0001, Randy Hsin, In Suk Chong |
PCS | 1 |
| 2024 | ELIM: Extremely Low-Complexity Implicit Neural Model for Super Resolution-Based CodingabstractThe super-resolution (SR)-based coding, which encodes a frame at a reduced resolution to achieve a lower bitrate, is a prevalent tool used in modern video coding standards. Accordingly, the low-resolution frame is restored to the full resolution at the reconstruction stage for subsequent reference. Therefore, the resolution restoration algorithm significantly affects the coding performance. This paper devises a highly efficient and extremely low complexity implicit neural model (ELIM) for SR-based encoding to support arbitrary scale factors. Specifically, ELIM consists of two stages: Feature Aggregation and Coordinate Upsampling. In Feature Aggregation, we embed a simplified attention block to the U-Net style framework to collect valuable information while reducing computational complexity through downsampling. In Coordinate Upsampling, in addition to the extracted content features, information including coordinate relative location and pixel cell size is fused to achieve better performance. We exemplify ELIM on the AV2 codec (the next generation of AV1). Extensive experiments demonstrate its superior performance: it achieves 4.17% BD-Rate gains over the anchor AV2 reference software with only 5,185 flops/pixel, significantly surpassing existing methods. The low complexity of ELIM is attractive to real applications. Dandan Ding, Urvang Joshi, Debargha Mukherjee |
PCS | 4 |
| 2023 | Neural Adaptive Loop Filtering for Video Coding: Exploring Multi-Hypothesis Sample RefinementabstractAdaptive loop filtering (ALF) is extensively investigated for lossy video coding to mitigate compression noise. Numerous learning-based ALFs have emerged recently and improved the coding efficiency significantly through the use of complexity-intensive, large-scale models trained on excessive samples, making it impractical for real-life applications. By contrast, lightweight, small-scale ALF models cannot promise convincing performance and model generalization. In principle, the ALF estimates the sample distortion for restoration. Instead of directly approximating the distortion as in existing solutions, we reformulate it as a Multi-hypothesis Sample Refinement (MSR) problem. To this end, we first generate multiple distortion hypotheses through a deep neural network (DNN) model. Then, these hypotheses are linearly superimposed to approximate the final distortion through the minimization of mean square error (MMSE) between the filtered reconstruction and its original, uncompressed input. Finally, the linear superimposition coefficients are explicitly signaled in the compressed bitstream. As seen, the superimposition coefficients inherently generalize the MSR to various content. And using DNNs to generate distortion hypotheses essentially models the spatial priors of a local block and its underlying compression error distribution. As a result, the MSR using a small-scale convolutional neural network (CNN) model with only 5k parameters and 5 KMACs/pixel achieves 4.35% (Intra) and 2.49% (Inter) BD-Rate (Bjøntegaard Delta Rate) gains over the AV1 anchor. Ablation studies further demonstrate that the MSR can be generalized to diverse small-scale network structures, different standards (e.g., H.265/HEVC and H.266/VVC), and diverse video content (e.g., screen videos). Dandan Ding, Guangkun Zhen, Debargha Mukherjee, Urvang Joshi, Zhan Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Non-Separable Filtering with Side-Information and Contextually-Designed Filters for Next Generation Video CodecsabstractWe propose two new in-loop filtering tools to augment the loop restoration process in AV1. Targeting challenging future compression scenarios that will be faced by the emerging AV2 standard, the proposed tools are designed to improve performance at both high and low bit-rates. Our non-separable Wiener filtering proposal aims to increase quality especially over directional features and textures in the decoded picture with the help of side-information at higher rates. The proposed pixel-adaptive filter complements at lower rates by improving performance without side-information relying solely on finely characterized pixel contexts. Simulation results show the rate-distortion efficacy of the proposed tools. Onur G. Guleryuz, Debargha Mukherjee, Yue Chen 0040, Keng-Shih Lu, Urvang Joshi |
ICIP | 5 |
| 2022 | Switchable CNN-Based Same-Resolution and Super-Resolution In-Loop Restoration for Next Generation Video CodecsabstractWe present a common framework for in-loop same-resolution and super-resolution restoration for incorporation into a next-generation video codec. Building on the in-loop filtering pipeline in the AV1 video codec from the Alliance for Open Media (AOM), we first enhance it to support symmetric spatial down-up scaling with better down and upscaling filters, followed by adding switchable frame-level CNNs to restore the reconstructed frames. Furthermore, the architectures for the CNNs used are constrained to be very simple with a relatively small number of parameters and multiply-add operations per decoded pixel to make them practically feasible in hardware. Preliminary results are presented on test sets being used by AOM for their ongoing next-generation video codec (AV2) development effort. Urvang Joshi, Yue Chen 0040, Debargha Mukherjee, Onur G. Guleryuz, Shan Li 0001, In Suk Chong |
ICIP | 1 |
| 2022 | Quadtree-based Guided CNN for AV1 In-loop FilteringabstractRecently, learning-based in-loop filtering has attracted lots of attention. State-of-the-art works generally deploy computationally expensive, large-scale neural networks, which is unfriendly to practical applications. Besides, since these models are generally pre-trained using a limited dataset and applied to various videos, they may fail in video contents excluded in the training dataset. To address these issues, this paper develops a Guide CNN in-loop filtering framework to obtain the restored signal. Our basic idea is to construct a subspace and use the projection of the original signal into this subspace to approximate the original signal itself. Specifically, we employ CNN to transform the degraded signal into M subsignals to construct the optimal subspace since the training of CNN is essentially an optimization procedure. Furthermore, unlike existing CNN models that process all blocks uniformly, our method leverages a quadtree structure to implement the Guided CNN through R-D optimization. As such, the best partition to Guided CNN can be determined. We exemplify the proposed method in AV1 codec. Experimental results show that the Guided CNN framework achieves 2.19% and 1.31% BD-Rate gains over the AV1 anchor in intra and inter coding mode, respectively, while the normal CNN achieves only 1.64% and 1.04%. Gongchun Ding, Dandan Ding, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
ICIP | 5 |
| 2021 | A Technical Overview of AV1abstractThe AV1 video compression format is developed by the Alliance for Open Media consortium. It achieves more than a 30% reduction in bit rate compared to its predecessor VP9 for the same decoded video quality. This article provides a technical overview of the AV1 codec design that enables the compression performance gains with considerations for hardware feasibility. Jingning Han, Bohan Li 0006, Debargha Mukherjee, Ching-Han Chiang, Adrian Grange, Hui Su, Sarah Parker, Sai Deng, Urvang Joshi, Yue Chen 0040, Yunqing Wang, Paul Wilkins, Yaowu Xu, Jim Bankoski |
Proc. IEEE | 10 |
| 2020 | Guided CNN Restoration with Explicitly Signaled Linear CombinationabstractState-of-the-art Convolutional Neural Network (CNN) based loop restoration generally involves a CNN structure with a large number of parameters and applies the CNN model to those degraded frames uniformly to generate their restored version, even though the contents within these frames are different. By contrast, in this paper, we propose a Guided CNN Restoration (GNR) scheme, where a CNN is used in conjunction with explicitly signaled guide parameters, with an aim to adapt the CNN model to different input contents. Specifically, the CNN architecture is designed such that the final restoration is constrained within the subspace generated by various output channels of the CNN, and meanwhile the weighting parameters for a linear combination of the output channels to obtain the final restoration are explicitly signaled by the encoders. The proposed GNR is incorporated into an AV1 encoder to replace the anchor in-loop filters and the weighting parameters are written into the encoded bitstream. Experimental results show that given a small CNN with 3,312 parameters, the proposed approach achieves a BD-rate reduction of 3.06% over the AV1 anchor, while the traditional CNN-based method only achieves 1.39%. Lingyi Kong, Dandan Ding, Fuchang Liu, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
ICIP | 5 |
| 2020 | On Extended Transform Partitions For The Next Generation Video CODECabstractAV1, AOMedia's royalty free codec, has enjoyed a great amount of success since its 2018 release. It currently achieves 31% BDRATE gains over VP9, and is on its way to becoming YouTube's default codec. Although the industry is currently focused on the implementation and optimization of AV1, AOMedia Research continues to develop new coding tools that deliver higher coding gains within acceptable complexity bounds. Here, we focus on improving transform coding. While AV1 has made great strides in transform coding over VP9, the residue signal still consumes a large portion of the bitstream. In this paper, we describe a more flexible transform partitioning scheme, which will allow the next generation codec to more efficiently target areas in the residue signal with high energy, leading to better residue compression. Sarah Parker, Yue Chen 0040, Urvang Joshi, Elliott Karpilovsky, Debargha Mukherjee |
ICIP | 3 |
| 2019 | AV1 in-loop Filtering using a Wide-Activation Structured Residual NetworkabstractThe in-loop filter, which constitutes an important part in modern video coding, improves both subjective and objective quality of reconstructed frames. Lately, Convolutional Neural Network (CNN) has demonstrated its superiority over traditional methods in addressing in-loop filtering problem. In this paper, we develop a CNN-based in-loop filter, namely Wide Activation Residual Network (WARN), for AV1 encoder. On top of the plain Residual Network (ResNet), we introduce wide activation to each residual block, making a more reasonable allocation of network parameters. When incorporating WARN into video encoder, particular to inter coding, it is intricate to obtain the global optimum performance. After simplifying this as an end-to-end trainable problem, we propose a skipping method by taking advantage of the hierarchical reference structure in AV1. Experimental results show that our WARN achieves up to 14.42% and 9.64% BD-rate reduction in intra and inter coding, respectively. All the code and model of our approach are available at https://github.com/IVC-Projects/AV1_WARN. Dandan Ding, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
ICIP | 4 |
| 2019 | A CNN-based In-loop Filtering Approach for AV1 Video CodecabstractIn-loop filter using Convolutional Neural Network (CNN) has lately attracted lots of attention in video coding. CNN models may be trained to learn how to restore degradation introduced by compression in pictures, and hence effectively help improve the coding efficiency. State-of-the-art work in this field generally employs a single network to enhance reconstructed frames mainly in intra coding. In this paper, we develop a depth-variable network handling both intra and inter coding. The depth of our network is varied with the distortion levels of reconstructed frames. Moreover, we leverage a skip enhancing strategy for inter coding, which improves both the coding efficiency and the resulting visual quality, while maintaining low computational complexity. We apply our approach to AV1, a newly released video coding standard from AOM. Experimental results show that our approach achieves an average BD-rate reduction of 7.27% and 5.57% for intra and inter modes, respectively, compared to AV1 anchor. The code and model of our approach are published in our Github website [1]. Dandan Ding, Debargha Mukherjee, Urvang Joshi, Yue Chen 0040 |
PCS | 4 |
| 2019 | In-loop Frame Super-resolution in AV1abstractAV1 is a recently standardized royalty-free video codec from the industry consortium Alliance for Open Media. One of the most innovative coding tools supported in AV1 is an in-loop frame super-resolution mode, that allows an encoder to code any frame at a horizontally reduced spatial resolution by one of several levels, followed by upsampling and super-resolving to full resolution, before replacing reference buffers. This mode is partly enabled by a feature in AV1 that natively allows the motion compensated prediction loop to operate across scales between a coded frame and the available references, thereby allowing on-the-fly resolution change mid-stream within a sequence. For the actual super-resolving process a normative upscaler is followed by an in-loop restoration tool that recovers some of the high frequency information lost in the downsampling process. On-the-fly resolution change capability in conjunction with the frame-superresolution mode in AV1 opens up a new dimension for codec bitrate and quality optimization that has not been possible to explore in any prior standardized video codec with (soon expected) decoding hardware support. This paper provides an overview of how the relevant tools work in AV1, but unlocking them with intelligent encoder decisions to extract real-world benefit, is largely left as future work. Urvang Joshi, Debargha Mukherjee, Yue Chen 0040, Sarah Parker, Adrian Grange |
PCS | 1 |
| 2018 | An Overview of Core Coding Tools in the AV1 Video CodecabstractAV1 is an emerging open-source and royalty-free video compression format, which is jointly developed and finalized in early 2018 by the Alliance for Open Media (AOMedia) industry consortium. The main goal of AV1 development is to achieve substantial compression gain over state-of-the-art codecs while maintaining practical decoding complexity and hardware feasibility. This paper provides a brief technical overview of key coding techniques in AV1 along with preliminary compression performance comparison against VP9 and HEVC. Yue Chen 0040, Debargha Mukherjee, Jingning Han, Adrian Grange, Yaowu Xu, Zoe Liu, Sarah Parker, Hui Su, Urvang Joshi, Ching-Han Chiang, Yunqing Wang, Paul Wilkins, Jim Bankoski, Luc N. Trudeau, Nathan E. Egge, Jean-Marc Valin, Thomas Davies 0002, Steinar Midtskogen, Andrey Norkin, Peter De Rivaz |
PCS | 10 |