Anustup Choudhury

dblp:09/1120 · DBLP profile ↗
← Back
18ranked-venue papers
11as first author
8since 2021 · last 2026
0000-0001-6618-9211ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 11 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Gaussian Representations for Video
abstract
We introduce Gaussian representations for videos (GaRV), a novel video encoding and decoding scheme based upon 3D Gaussians. Unlike traditional representations, which encode videos as sequences of frames, or neural representations, which encode videos within the weights of a neural network, we encode videos as a collection of 3D Gaussians within a space-time volume. The key advantage of our approach is that it enables efficient and flexible rasterization-based video decoding. With a slight drop in overall compression rate, GaRV offers an 8-50× improvement in decoding time and 2.5-15× reduction in GPU memory compared with neural counterparts. Existing Gaussian video techniques require 2-30× more disk space, while also using more GPU resources than GaRV. Moreover, GaRV offers unique flexibility in how and when pixels are decoded: One can non-sequentially decode frames/regions without penalty and can selectively decode regions at high-resolution to enable low-cost foveated video decoding.
Sachin Shah, Anustup Choudhury, Guan-Ming Su, Jaclyn Pytlarz, Christopher A. Metzler, Trisha Mittal
WACV2
2025 Quanta-Slomo: Single Photon Camera Guided 100x Video Frame Interpolation
abstract
Video Frame Interpolation (VFI) enhances video quality by computationally increasing video frame rates. Conventional VFI methods rely on simplified motion assumptions (e.g. linear motion) between adjacent frames to generate intermediate frames, which often causes error in highly dynamic scenarios. The Single Photon Avalanche Diode (SPAD) array is an emerging type of quanta image sensor that can capture videos at very high frame rates but can be heavily corrupted by Poisson noise. We present Quanta-SloMo, a novel VFI method that leverages the high frame rate but noisy video from a SPAD sensor to guide the VFI of high quality, low frame rate video captured by a CMOS sensor. Our framework utilizes residual learning, feature alignment through deformable convolution, multi-frame merging, and handles the noise-to-blur trade-off between SPAD video frames through a virtual exposure stack. We demonstrate that Quanta-SloMo outperforms state-of-the-art VFI methods by significant margins, especially for scenes with large motion.
Anustup Choudhury, Guan-Ming Su, Andreas Velten
ICIP2
2025 Long-Short Exposure Fusion With Event Data For Low-Light Video Enhancement
abstract
Traditional frame-based imaging systems face trade-offs especially under low-light conditions: long-exposure frames have high signal-to-noise ratio (SNR) but suffer from motion blur and low frame rate, while short-exposure frames are sharp but suffer from noise. Event cameras can complement these modalities by capturing brightness changes and exposure-invariant information at very high temporal resolution. We propose a novel multi-modal framework for lowlight enhancement by integrating information from long-exposure, short-exposure, and event data. We introduce an enhanced lightup mechanism to improve short-exposure frames using illumination from long-exposure frames, and an SNR-guided fusion strategy for noise suppression and detail preservation. We use event information to improve feature selection, especially in the low SNR regions. Experimental results on public datasets demonstrate that our method can effectively handle noise and motion blur. We obtain 14dB improvement in PSNR over state-of-the-art methods based on conventional frames and 4dB improvement over methods based on frames and events.
Daizong Tian, Anustup Choudhury, Walt Husak
ICIP2
2024 NeRVA: Joint Implicit Neural Representations for Videos and Audios
abstract
Neural fields also known as implicit neural representations (INR) have recently been shown to be quite effective at representing video content. However, video content typically contains audio and existing works on INR do not represent both video and audio. Since INR for audio is not well-explored, we propose a novel neural representation for representing audio (NeRA) that can represent audio using a neural network. We also propose a novel neural representation for jointly representing both videos and audios (NeRVA) using a neural network. We represent multimedia as a neural network that takes timestamp as input and outputs the corresponding RGB image frame and the audio samples. The proposed neural network architecture uses a combination of Multi-Layer Perceptron (MLP) and convolutional blocks. We also demonstrate that our joint representation of multimedia content has better performance than individually representing the components of multimedia (An improvement of +2dB PSNR for video and 10× FAD for audio).
Anustup Choudhury, Praneet Singh, Guan-Ming Su
ICME1
2024 Outdoor Scene Relighting with Diffusion Models
Jinlin Lai, Anustup Choudhury, Guan-Ming Su
ICPR (6)2
2024 A 'deep' review of video super-resolution
Subhadra Gopalakrishnan, Anustup Choudhury
Signal Process. Image Commun.2
2023 Neural Field Real-Time Transmission Using Multiple Description Coding with Random Position Sampling
abstract
Neural fields are a new signal representation and are widely used now to represent various forms of multimedia. However, due to their size, they could introduce latency when delivered over the network. A way to mitigate that is to use Multiple description coding (MDC), that provides multiple descriptions of the same content which can then be transmitted along different paths to improve reliability/efficiency. In this paper, we introduce a novel MDC framework that is based on neural fields. We first apply a randomly sampled neural field (with small model size) to generate multiple descriptions based on random initializations. We leverage the fact that neural fields are continuous functions (constructed using a multi-layer perceptron (MLP)) and use that to reconstruct the entire original content and progressively improve the quality as more descriptions are obtained. We validate the effectiveness of the proposed method by showing results on public data sets.
Anustup Choudhury, Guan-Ming Su
ICIP1
2023 Progressive Coding for Neural Field Transmission
abstract
Neural fields are a new signal representation and are widely used now to represent various forms of multimedia. However, due to their size, they could introduce latency when delivered over the network. A way to mitigate that is to use Progressive Coding which encodes the content into a single bitstream that can be decoded at various bitrates. In this paper, we introduce a novel progressive coding framework for neural fields. We propose three progressive scanning mechanisms - first is based on the layers of the model, the second is based on the bit-planes across all layers of the model, and the third is based on the block of coefficients across all layers of the model. We leverage the fact that neural fields are continuous functions (constructed using a multi-layer perceptron (MLP)) and use that to reconstruct the entire original content and progressively improve the quality as more bits are obtained. We validate the effectiveness of the proposed method by showing results on public data sets.
Anustup Choudhury, Guan-Ming Su
ISM1
2020 Robust HDR image quality assessment using combination of quality metrics
Anustup Choudhury
Multim. Tools Appl.1
2019 HDR Display Quality Evaluation by incorporating Perceptual Component Models into a Machine Learning framework
Anustup Choudhury, Scott Daly
Signal Process. Image Commun.1
2018 Machine learning as applied intrinsically to individual dimensions of HDR Display Quality
abstract
This study builds on previous work exploring machine learning and perceptual transforms in predicting overall display quality as a function of image quality dimensions that correspond to physical display parameters. Previously, we found that the use of perceptually transformed parameters or machine learning exceeded the performance of predictors using just physical parameters and linear regression. Further, the combination of perceptually transformed parameters with machine learning allowed for robustness to parameters outside of the data set, both for cases of interpolation and extrapolation. Here we apply machine learning at a more intrinsic level. We first evaluate how well the machine learning can develop predictors of the individual dimensions of the overall quality, and then how well those individual predictors can be consolidated across themselves to predict the overall display quality. Having predictions of individual dimensions of quality that are closely related to specific hardware design choices enables more nimble cost trade-off design options.
Anustup Choudhury, Scott Daly
PCS1
2016 Boosting performance and speed of single-image super-resolution based on partitioned linear regression
abstract
In this paper, we show how to boost the performance and speed of the Simple Functions (SF) algorithm for single-image superresolution [1]. This method partitions the low-resolution patch space and learns linear regressors to map low-resolution to highresolution patches. We optimize the partitioning of the patch feature space by first employing dimensionality reduction and then explicitly minimizing the overall super-resolution reconstruction error during training. We also improve selection of training patches. In the super-resolution stage, we use a k-d tree for fast nearest neighbor search of partitions, and combine multiple regression models from neighboring partitions. Experimental results on benchmark data sets show improvements in both image quality and speed over SF. Also, our method outperforms state-of-the-art super-resolution methods in image quality.
Anustup Choudhury, Peter van Beek
ICIP1
2015 Facial video super resolution using semantic exemplar components
abstract
We present a method for video super resolution using exemplar images of semantic components. In previous work, we proposed a novel super resolution framework based on semantic components and applied it to still images of human faces. In this paper, we extend the approach to video sequences and propose several methods to overcome temporal jitter that results from standard single frame processing. To achieve consistent selection of facial components from a database of exemplars, we introduce a weighted histogram constructed over a temporal window. We then use pixel-based alignment between the exemplar and input image to reduce temporal jitter of the selected component. To further improve temporal stability, we include a temporal constraint into a final optimization stage that blends high resolution exemplar image data into the upscaled input image. We compare our results on face video clips to those of several state-of-the-art super resolution methods, demonstrating the efficacy of the proposed approach.
Anustup Choudhury, Peter van Beek, C. Andrew Segall
ICIP2
2012 Image detail enhancement using a dictionary technique
abstract
We present a novel approach to detail enhancement using a dictionary-based technique. For each low-resolution input image patch, we seek a sparse representation from an over-complete dictionary and use that to estimate the high-resolution patch. We modify an existing dictionary-based super-resolution method in several ways to achieve enhancement of fine detail without introduction of new artifacts. These modifications include adaptive enhancement of reconstructed detail patches based on edge analysis to avoid halo artifacts and using an adaptive regularization term to enable noise suppression while enhancing detail. We compare with state-of-the-art methods and show better results in terms of enhancement with suppression of noise.
Anustup Choudhury, Peter van Beek, C. Andrew Segall
ICIP1
2012 A Framework for Robust Online Video Contrast Enhancement Using Modularity Optimization
abstract
We address the problem of video contrast enhancement. Existing techniques either do not exploit temporal information at all or do not exploit it correctly. This results in inconsistency that causes undesirable flash and flickering artifacts. Our method analyzes video streams and cluster frames that are similar to each other. Our method does not have omniscient information about the entire video sequence. It is an online process with a fixed delay. A sliding window mechanism successfully detects shot boundaries “on-the-fly” in a video. A graph-based technique called “modularity” performs automatic clustering of video frames without a priori information about clusters. For every cluster in the video, we extract key frames belonging to each cluster using eigen analysis and estimate enhancement parameters for only the key frame, then use these parameters to enhance frames belonging to that cluster, thus making our method robust. We evaluate the clustering method on video sequences from the TRECVid 2001 dataset and compare it with existing methods. We show reduction of flash artifacts in enhanced videos. We show statistically significant improvement in perceived video quality and validate that by conducting experiments on human observers. We show application of our clustering process to perform robust video segmentation.
Anustup Choudhury, Gérard G. Medioni
IEEE Trans. Circuits Syst. Video Technol.1
2010 Color Constancy Using Standard Deviation of Color Channels
abstract
We address here the problem of color constancy and propose a new method to achieve color constancy based on the statistics of images with color cast. Images with color cast have standard deviation of one color channel significantly different from that of other color channels. This observation is also applicable to local patches of images and ratio of the maximum and minimum standard deviation of color channels of local patches is used as a prior to select a pixel color as illumination color. We provide extensive validation of our method on commonly used datasets having images under varying illumination conditions and show our method to be robust to choice of dataset and at least as good as current state-of-the-art color constancy approaches.
Anustup Choudhury, Gérard G. Medioni
ICPR1
2009 Color constancy using denoising methods and cepstral analysis
abstract
We address here the problem of color constancy and propose two new methods for achieving color constancy-the first method uses denoising techniques such as a Gaussian filter, Median filter, Bilateral filter and Non-local means filter to smooth the image for illuminant estimation, while the second method acts in the frequency domain by doing a cepstral analysis of the image. We provide extensive validation tests for our illuminant estimation on commonly used datasets having images under different illumination conditions, and the results show that both new methods outperform current state-of-the-art color constancy approaches, at a very low computational cost.
Anustup Choudhury, Gérard G. Medioni
ICIP1
2006 Using Semantic Features for Scene Classification: how Good do they Need to Be?
abstract
Semantic scene classification is a useful, yet challenging problem in image understanding. Most existing systems are based on low-level features, such as color or texture, and succeed to some extent. Intuitively, semantic features, such as sky, water, or foliage, which can be detected automatically, should help close the so-called semantic gap and lead to higher scene classification accuracy. To answer the question of how accurate the detectors themselves need to be, we adopt a generally applicable scene classification scheme that combines semantic features and their spatial layout as encoded implicitly using a block-based method. Our scene classification results show that although our current detectors collectively are still inadequate to outperform low-level features under the same scheme, semantic features hold promise as simulated detectors can achieve superior classification accuracy once their own accuracies reach above a nontrivial 90%
Matthew R. Boutell, Anustup Choudhury, Jiebo Luo 0001
ICME2