Somdyuti Paul

dblp:204/1738 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0002-2762-7263ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Towards efficient real-time video motion transfer via generative time series modeling
abstract
Motion Transfer is an Artificial Intelligence (AI) technique that synthesizes videos by transferring motion dynamics from a driving video to a source image. In this work we propose a deep learning-based framework to enable real-time video motion transfer which is critical for enabling bandwidth-efficient applications such as video conferencing, remote health monitoring, virtual reality interaction, and vision-based anomaly detection. This is done using keypoints which serve as semantically meaningful, compact representations of motion across time, and are extracted from every video frame via a self-supervised detector. To enable bandwidth savings during video transmission we perform forecasting of keypoints using two generative time series models–Variational Recurrent Neural Networks (VRNN) and Gated Recurrent Units with Normalizing Flows (GRU-NF)–enabling both single and diverse future prediction modes. The predicted keypoints are transformed into realistic video frames using an optical flow-based module paired with a generator network, thereby facilitating accurate video forecasting and enabling efficient, low-frame-rate video transmission. Based on the application this allows the framework to either generate a deterministic future sequence or sample a diverse set of plausible futures. Experimental results across three benchmark video datasets using state-of-the-art quality and diversity metrics for video animation and reconstruction tasks demonstrate that VRNN achieves the best point-forecast fidelity (lowest MAE) in the majority of evaluated settings in applications requiring stable and accurate multi-step forecasting (e.g., video conferencing, remote patient monitoring) and is particularly competitive in higher-uncertainty, multi-modal settings. This is achieved by utilizing the superior reconstruction property of the Variational Autoencoder and by introducing recurrently conditioned stochastic latent variables that carry past contexts to capture uncertainty and temporal variation. On the other hand the GRU-NF model enables richer diversity of generated videos while maintaining high visual quality to better support tasks like AI-driven anomaly detection. This is realized by learning an invertible, exact-likelihood mapping between the keypoints and their latent representations which supports rich and controllable sampling of diverse yet coherent keypoint sequences. Our work lays the foundation for next-generation AI systems that require real-time, bandwidth-efficient, and semantically controllable video generation, with broad implications for communication, health, and manufacturing applications. The code is available at: https://github.com/Tasmiah1408028/RealtimeVideoMotionTransfer
Tasmiah Haque, Md. Asif Bin Syed, Byungheon Jeong, Sumit Mohan, Somdyuti Paul, Imtiaz Ahmed 0002, Srinjoy Das
Multim. Tools Appl.6
2024 Convex Hull Prediction for Adaptive Video Streaming by Recurrent Learning
abstract
Adaptive video streaming relies on the construction of efficient bitrate ladders to deliver the best possible visual quality to viewers under bandwidth constraints. The traditional method of content dependent bitrate ladder selection requires a video shot to be pre-encoded with multiple encoding parameters to find the optimal operating points given by the convex hull of the resulting rate-quality curves. However, this pre-encoding step is equivalent to an exhaustive search process over the space of possible encoding parameters, which causes significant overhead in terms of both computation and time expenditure. To reduce this overhead, we propose a deep learning based method of content aware convex hull prediction. We employ a recurrent convolutional network (RCN) to implicitly analyze the spatiotemporal complexity of video shots in order to predict their convex hulls. A two-step transfer learning scheme is adopted to train our proposed RCN-Hull model, which ensures sufficient content diversity to analyze scene complexity, while also making it possible to capture the scene statistics of pristine source videos. Our experimental results reveal that our proposed model yields better approximations of the optimal convex hulls, and offers competitive time savings as compared to existing approaches. On average, the pre-encoding time was reduced by 53.8% by our method, while the average Bjøntegaard delta bitrate (BD-rate) of the predicted convex hulls against ground truth was 0.26%, and the mean absolute deviation of the BD-rate distribution was 0.57%.
Somdyuti Paul, Andrey Norkin, Alan C. Bovik
IEEE Trans. Image Process.1
2023 Self-Supervised Learning of Perceptually Optimized Block Motion Estimates for Video Compression
abstract
Block based motion estimation is integral to inter prediction processes performed in hybrid video codecs. Prevalent block matching based methods that are used to compute block motion vectors (MVs) rely on computationally intensive search procedures. They also suffer from the aperture problem, which tends to worsen as the block size is reduced. Moreover, the block matching criteria used in typical codecs do not account for the resulting levels of perceptual quality of the motion compensated pictures that are created upon decoding. Towards achieving the elusive goal of perceptually optimized motion estimation, we propose a search-free block motion estimation framework using a multi-stage convolutional neural network, which is able to conduct motion estimation on multiple block sizes simultaneously, using a triplet of frames as input. This composite block translation network (CBT-Net) is trained in a self-supervised manner on a large database that we created from publicly available uncompressed video content. We deploy the multi-scale structural similarity (MS-SSIM) loss function to optimize the perceptual quality of the motion compensated predicted frames. Our experimental results highlight the computational efficiency of our proposed model relative to conventional block matching based motion estimation algorithms, for comparable prediction errors. Further, when used to perform inter prediction in AV1, the MV predictions of the perceptually optimized model result in average Bjøntegaard-delta rate (BD-rate) improvements of -1.73% and -1.31% with respect to the MS-SSIM and Video Multi-Method Assessment Fusion (VMAF) quality metrics, respectively, as compared to the block matching based motion estimation system employed in the SVT-AV1 encoder.
Somdyuti Paul, Andrey Norkin, Alan C. Bovik
IEEE Trans. Image Process.1
2022 A Subjective and Objective Study of Space-Time Subsampled Video Quality
abstract
Video dimensions are continuously increasing to provide more realistic and immersive experiences to global streaming and social media viewers. However, increments in video parameters such as spatial resolution and frame rate are inevitably associated with larger data volumes. Transmitting increasingly voluminous videos through limited bandwidth networks in a perceptually optimal way is a current challenge affecting billions of viewers. One recent practice adopted by video service providers is space-time resolution adaptation in conjunction with video compression. Consequently, it is important to understand how different levels of space-time subsampling and compression affect the perceptual quality of videos. Towards making progress in this direction, we constructed a large new resource, called the ETRI-LIVE Space-Time Subsampled Video Quality (ETRI-LIVE STSVQ) database, containing 437 videos generated by applying various levels of combined space-time subsampling and video compression on 15 diverse video contents. We also conducted a large-scale human study on the new dataset, collecting about 15,000 subjective judgments of video quality. We provide a rate-distortion analysis of the collected subjective scores, enabling us to investigate the perceptual impact of space-time subsampling at different bit rates. We also evaluated and compare the performance of leading video quality models on the new database. The new ETRI-LIVE STSVQ database is being made freely available at (https://live.ece.utexas.edu/research/ETRI-LIVE_STSVQ/index.html).
Dae Yeol Lee, Somdyuti Paul, Christos G. Bampis, Hyunsuk Ko, Seyoon Jeong, Blake Homan, Alan C. Bovik
IEEE Trans. Image Process.2
2021 On visual masking estimation for adaptive quantization using steerable filters
Somdyuti Paul, Andrey Norkin, Alan C. Bovik
Signal Process. Image Commun.1
2020 Speeding Up VP9 Intra Encoder With Hierarchical Deep Learning-Based Partition Prediction
abstract
In VP9 video codec, the sizes of blocks are decided during encoding by recursively partitioning 64×64 superblocks using rate-distortion optimization (RDO). This process is computationally intensive because of the combinatorial search space of possible partitions of a superblock. Here, we propose a deep learning based alternative framework to predict the intra-mode superblock partitions in the form of a four-level partition tree, using a hierarchical fully convolutional network (H-FCN). We created a large database of VP9 superblocks and the corresponding partitions to train an H-FCN model, which was subsequently integrated with the VP9 encoder to reduce the intra-mode encoding time. The experimental results establish that our approach speeds up intra-mode encoding by 69.7% on average, at the expense of a 1.71% increase in the Bjøntegaard-Delta bitrate (BD-rate). While VP9 provides several built-in speed levels which are designed to provide faster encoding at the expense of decreased rate-distortion performance, we find that our model is able to outperform the fastest recommended speed level of the reference VP9 encoder for the good quality intra encoding configuration, in terms of both speedup and BD-rate.
Somdyuti Paul, Andrey Norkin, Alan C. Bovik
IEEE Trans. Image Process.1
2019 Image Statistic Models Characterize Well Log Image Quality
abstract
Assessing the image quality of well logs is essential to ensure the accuracy of their digitization and subsequent processing. Currently, the suitability of well logs for information retrieval is solely determined on the basis of subjective judgments of their image quality by human experts. The success of natural scene statistics (NSS)-based models that are used to conduct no-reference (NR) quality assessment of photographic images motivates us to try to exploit them to characterize the quality of nonphotographic images, such as well logs. Accordingly, we develop a scheme to characterize the quality of a well log as “acceptable” or “unacceptable” for subsequent processing based on the natural image quality evaluator (NIQE), a successful NR image quality assessment model based on the NSS. Our experimental results show that the objective quality scores thus obtained can be reliably used to eliminate well logs of inferior quality from the processing pipeline, which can serve as a beneficial step to reduce the human hours spent in examining well logs and to improve the rate of information retrieval as well as the accuracy of retrieved information. Source code for the trained well log image quality predictor is available at https://github.com/Somdyuti2/Well_log_IQA.
Somdyuti Paul, Alan C. Bovik
IEEE Geosci. Remote. Sens. Lett.1
2017 Spatiotemporal Colorization of Video Using 3D Steerable Pyramids
abstract
We propose a new technique for video colorization based on spatiotemporal color propagation in the 3D video volume, utilizing the dominant orientation response obtained from the steerable pyramid decomposition of the video. The volumetric color diffusion from the sources that are marked by scribbles occurs across spatiotemporally smooth regions, and the prevention of leakage is facilitated by the spatiotemporal discontinuities in the output of steerable filters, representing the object boundaries and motion boundaries. Unlike most existing methods, our approach dispenses with the need of motion vectors for interframe color transfer and provides a general framework for image and video colorization. The experimental results establish the effectiveness of the proposed approach in colorizing videos having different types of motion and visual content even in the presence of occlusion, in terms of accuracy and computational requirements.
Somdyuti Paul, Saumik Bhattacharya, Sumana Gupta
IEEE Trans. Circuits Syst. Video Technol.1