Onkar Susladkar

dblp:321/1077 · also Onkar Kishor Susladkar · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0003-4511-1858ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Pyramidal Spectrum: Frequency-based Hierarchically Vector Quantized VAE for Videos
abstract
Variational Autoencoders (VAEs) form the foundation of modern video generation models. In particular, discrete latent VAEs with vector quantization have gained prominence for their superior perceptual sharpness, ability to model long-range dynamics, and efficient adaptation to downstream tasks. However, existing discrete VAEs face two key limitations: (i) a lack of frequency-domain modeling to enhance global spatiotemporal understanding, and (ii) fixed-resolution quantization schemes, preventing effective modeling of coarse-to-fine spatiotemporal hierarchies essential for video generation. To address these limitations, we propose a Pyramidal Vector Quantized Variational Autoencoder (PVQ-VAE) for videos. PVQ-VAE’s encoder–decoder leverages Fast Fourier Transform and Discrete Wavelet Transform to capture global semantics and multi-scale local details jointly. We introduce Pyramidal Vector Quantization (PVQ), a hierarchical quantization scheme that discretizes features at multiple resolutions to better capture multi-scale information. To further boost fidelity, we introduce a cross-modal contrastive loss guided by a pretrained high-resolution image VAE. PVQ-VAE achieves state-of-the-art performance on WebVid-val, COCO-val, and MCL-JCV, reconstructing videos with high perceptual quality at up to 32× spatial and 16× temporal compression. Project page of PVQ-VAE.
Tushar Prakash, Onkar Susladkar, Inderjit S. Dhillon, Sparsh Mittal
WACV2
2026 Confidence Through Parallel Attention for Depth and Uncertainty Estimation in Dynamic Environments
abstract
Monocular depth estimation is crucial for robotics, offering a lightweight and scalable alternative to stereo or LiDAR-based systems. While recent methods have achieved high accuracy, their efficacy degrades under real-world conditions such as occlusion, and domain shifts. We introduce ConFiDeNet, a unified framework that jointly predicts metric depth and associated aleatoric uncertainty, enabling risk-aware robotic perception. ConFiDeNet employs a lightweight parallel attention module that efficiently fuses semantic cues from DINOv2 dense descriptors and SAM2-based segmentation for densely occluded objects, enhancing structural understanding without sacrificing real-time performance. Further, we explicitly condition the model on environment type, improving generalization across diverse indoor and outdoor scenes without retraining. Our method achieves state-of-the-art results across six datasets under both supervised and zero-shot settings, outperforming nine prior techniques, including Marigold, ZoeDepth, PatchFusion, and MonoProb. With significantly faster inference and high prediction confidence, ConFiDeNet is readily deployable for embodied AI, self-driving applications, and robotic manipulation tasks. Keywords: Monocular depth estimation, Feature fusion, Robotic Vision, point cloud estimation. For code and weights Check the project page here
Onkar Susladkar, Rohit Pawar, Chirag Sehgal, Samaksh Ujjawal, Sparsh Mittal
WACV1
2025 ViCTr: Vital Consistency Transfer for Pathology Aware Image Synthesis
abstract
Synthesizing medical images remains challenging due to limited annotated pathological data, modality domain gaps, and the complexity of representing diffuse pathologies such as liver cirrhosis. Existing methods often struggle to maintain anatomical fidelity while accurately modeling pathological features, frequently relying on priors derived from natural images or inefficient multi-step sampling. In this work, we introduce ViCTr (Vital Consistency Transfer), a novel two-stage framework that combines a rectified flow trajectory with a Tweedie-corrected diffusion process to achieve high-fidelity, pathology-aware image synthesis. First, we pretrain ViCTr on the ATLAS-8k dataset using Elastic Weight Consolidation (EWC) to preserve critical anatomical structures. We then fine-tune the model adversarially with Low-Rank Adaptation (LoRA) modules for precise control over pathology severity. By reformulating Tweedie's formula within a linear trajectory framework, ViCTr supports one-step sampling, reducing inference from 50 steps to just 4, without sacrificing anatomical realism. We evaluate ViCTr on BTCV (CT), AMOS (MRI), and CirrMRI600+ (cirrhosis) datasets. Results demonstrate state-of-the-art performance, achieving a Medical Frechet Inception Distance (MFID) of 17.01 for cirrhosis synthesis 28% lower than existing approaches and improving nnUNet segmentation by +3.8% mDSC when used for data augmentation. Radiologist reviews indicate that ViCTr-generated liver cirrhosis MRIs are clinically indistinguishable from real scans. To our knowledge, ViCTr is the first method to provide fine-grained, pathology-aware MRI synthesis with graded severity control, closing a critical gap in AI-driven medical imaging research.
Onkar Susladkar, Gayatri Deshmukh, Yalcin Tur, Gorkem Durak, Ulas Bagci
ICCV1
2025 Historic Scripts to Modern Vision: A Novel Dataset and A VLM Framework for Transliteration of Modi Script to Devanagari
Harshal Kausadikar, Tanvi Kale, Onkar Susladkar, Sparsh Mittal
ICDAR (5)3
2025 MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion
abstract
The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video processing. First, we introduce the 3D Mobile Inverted Vector-Quantization Variational Autoencoder (3D-MBQ-VAE), which combines Variational Autoencoders (VAEs) with masked modeling to enhance spatiotemporal video compression. The model achieves superior temporal consistency and state-of-the-art (SOTA) reconstruction quality by employing a novel training strategy with full frame masking. Second, we present MotionAura, a text-to-video generation framework that utilizes vector-quantized diffusion models to discretize the latent space and capture complex motion dynamics, producing temporally coherent videos aligned with text prompts. Third, we propose a spectral transformer-based denoising network that processes video data in the frequency domain using the Fourier Transform. This method effectively captures global context and long-range dependencies for high-quality video generation and denoising. Lastly, we introduce a downstream task of Sketch Guided Video Inpainting. This task leverages Low-Rank Adaptation (LoRA) for parameter-efficient fine-tuning. Our models achieve SOTA performance on a range of benchmarks. Our work offers robust frameworks for spatiotemporal modeling and user-driven video content manipulation.
Onkar Susladkar, Jishu Sen Gupta, Chirag Sehgal, Sparsh Mittal, Rekha Singhal
ICLR1
2025 Large-scale multi-center CT and MRI segmentation of pancreas with deep learning
abstract
Automated volumetric segmentation of the pancreas on cross-sectional imaging is needed for diagnosis and follow-up of pancreatic diseases. While CT-based pancreatic segmentation is more established, MRI-based segmentation methods are understudied, largely due to a lack of publicly available datasets, benchmarking research efforts, and domain-specific deep learning methods. In this retrospective study, we collected a large dataset (767 scans from 499 participants) of T1-weighted (T1 W) and T2-weighted (T2 W) abdominal MRI series from five centers between March 2004 and November 2022. We also collected CT scans of 1,350 patients from publicly available sources for benchmarking purposes. We introduced a new pancreas segmentation method, called PanSegNet , combining the strengths of nnUNet and a Transformer network with a new linear attention module enabling volumetric computation. We tested PanSegNet ’s accuracy in cross-modality (a total of 2,117 scans) and cross-center settings with Dice and Hausdorff distance (HD95) evaluation metrics. We used Cohen’s kappa statistics for intra and inter-rater agreement evaluation and paired t-tests for volume and Dice comparisons, respectively. For segmentation accuracy, we achieved Dice coefficients of 88.3% (±7.2%, at case level) with CT, 85.0% (±7.9%) with T1 W MRI, and 86.3% (±6.4%) with T2 W MRI. There was a high correlation for pancreas volume prediction with R 2 of 0.91, 0.84, and 0.85 for CT, T1 W, and T2 W, respectively. We found moderate inter-observer (0.624 and 0.638 for T1 W and T2 W MRI, respectively) and high intra-observer agreement scores. All MRI data is made available at https://osf.io/kysnj/ . Our source code is available at https://github.com/NUBagciLab/PaNSegNet . • We develop a first-ever cross-platform compatible (T1 W, T2 W, and CT) pancreas segmentation tool, named PanSegNet . • PaNSegNet has innovative “linear self-attention” blocks to reduce computational cost significantly while operating on 3D. • We shared our both source code and multi-center multi-contrast MRI datasets with ground truths. • PaNSegNet underwent rigorous validation, including cross-domain and multi-center comparisons between CT and MRI scans.
Zheyuan Zhang 0001, Elif Keles, Gorkem Durak, Yavuz Taktak, Onkar Susladkar, Vandan Gorade, Debesh Jha, Asli C. Ormeci, Alpay Medetalibeyoglu, Lanhong Yao, Bin Wang 0068, Ilkin Isler, Linkai Peng, Hongyi Pan, Camila Lopes Vendrami, Amir Bourhani, Yury Velichko, Boqing Gong, Concetto Spampinato, Ayis Pyrros, Pallavi Tiwari, Derk C. F. Klatte, Megan Engels, Sanne Hoogenboom, Candice W. Bolan, Emil Agarunov, Nassier Harfouch, Chenchan Huang, Marco J. Bruno, Ivo Schoots, Rajesh Keswani, Frank H. Miller, Tamas Gonda, Cemal Yazici, Temel Tirkes, Baris Turkbey, Michael B. Wallace, Ulas Bagci
Medical Image Anal.5
2024 GRIZAL: Generative Prior-guided Zero-Shot Temporal Action Localization
abstract
Zero-shot temporal action localization (TAL) aims to temporally localize actions in videos without prior training examples.To address the challenges of TAL, we offer GRIZAL, a model that uses multimodal embeddings and dynamic motion cues to localize actions effectively.GRIZAL achieves sample diversity by using large-scale generative models such as GPT-4 for generating textual augmentations and DALL-E for generating image augmentations.Our model integrates vision-language embeddings with optical flow insights, optimized through a blend of supervised and self-supervised loss functions.On Activi-tyNet, Thumos14 and Charades-STA datasets, GRIZAL vastly outperforms state-of-the-art zero-shot TAL models, demonstrating its robustness and adaptability across a wide range of video content.The code and models are available on https://github.com/CandleLabAI/ GRIZAL-EMNLP2024.
Onkar Susladkar, Gayatri Deshmukh, Vandan Gorade, Sparsh Mittal
EMNLP1
2024 D2Styler: Advancing Arbitrary Style Transfer with Discrete Diffusion Methods
Onkar Susladkar, Gayatri Deshmukh, Sparsh Mittal, Parth Shastri
ICPR (6)1
2024 Textual Alchemy: CoFormer for Scene Text Understanding
abstract
The paper presents CoFormer (Convolutional Fourier Transformer), a robust and adaptable transformer architecture designed for a range of scene text tasks. CoFormer integrates convolution and Fourier operations into the transformer architecture. Thus, it leverages convolution properties such as shared weights, local receptive fields, and spatial subsampling, while the Fourier operation emphasizes composite characteristics from the frequency domain. The research further proposes two new pretraining datasets, named Textverse10M-E and Textverse10M-H. Using these datasets, we demonstrate the efficacy of pretraining for scene text understanding. CoFormer achieves state-of-the-art results with and without pretraining on two downstream tasks: scene text recognition (STR) and scene text editing (STE). The paper further proposes LISTNet (Language Invariant Style Transfer), a novel framework for bi-lingual STE. It also introduces three datasets, viz., TST500K for STE, CSTR2.5M and Akshara550 for STR. The source-code of CoFormer is available at https://github.com/CandleLabAI/CoFormer-WACV-2024.
Gayatri Deshmukh, Onkar Susladkar, Dhruv Makwana, Sparsh Mittal, R. Sai Chandra Teja
WACV2
2024 LIVENet: A novel network for real-world low-light image denoising and enhancement
abstract
Low-light image enhancement (LLIE) is the process of improving the quality of images taken in low-light conditions while striking a balance between enhancing image illumination and maintaining their natural appearance. This involves reducing noise, enhancing details, and correcting colors, all while avoiding artifacts such as halo effects or color distortions. We propose LIVENet, a novel deep neural network that jointly performs noise reduction on lowlight images and enhances illumination and texture details. LIVENet has two stages: the image enhancement stage and the refinement stage. For the image enhancement stage, we propose a Latent Subspace Denoising Block (LSDB) that uses a low-rank representation of low-light features to suppress the noise and predict a noise-free grayscale image. We propose enhancing an RGB image by eliminating noise. This is done by converting it into YCbCr color space and replacing the noisy luminance (Y) channel with the predicted noise-free grayscale image. LIVENet also predicts the transmission map and atmospheric light in the image enhancement stage. LIVENet produces an enhanced image with rich color and illumination by feeding them to an atmospheric scattering model. In the refinement stage, the texture information from the grayscale image is incorporated into the improved image using a Spatial Feature Transform (SFT) layer. Experiments on different datasets demonstrate that LIVENet’s enhanced images consistently outperform previous techniques across various quality metrics. The source code can be obtained from https://github.com/CandleLabAI/LiveNet.
Dhruv Makwana, Gayatri Deshmukh, Onkar Susladkar, Sparsh Mittal, R. Sai Chandra Teja
WACV3
2024 MOVES: Movable and moving LiDAR scene segmentation in label-free settings using static reconstruction
Dhruv Makwana, Onkar Susladkar, Anurag Mittal, Prem Kumar Kalra
Pattern Recognit.3
2023 SLBERT: A Novel Pre-Training Framework for Joint Speech and Language Modeling
abstract
We propose SLBERT (Speech and Language pre-training framework for BERT), an end-to-end trainable framework for learning joint representations of speech and language modalities. We enhance the well-known BERT architecture to provide a dual-stream multimodal architecture that processes both speech and language input. To enable effective information exchange between the two modalities, we introduce a novel attention fusion mechanism via AF-Blocks. To acquire robust contrastive representations for speech and language processing applications, we pre-train SLBERT on three auxiliary tasks: Masked Language Modeling, Masked Speech Modeling, and Speech-Language Matching. We evaluate our proposed model on two well-known multimodal tasks: intent classification and sentiment analysis. Our model achieves state-of-the-art results on both benchmarks while surpassing even larger baselines.
Onkar Susladkar, Prajwal Gatti, Santosh Kumar Yadav
ICASSP1
2023 TPFNet: A Novel Text In-painting Transformer for Text Removal
Onkar Susladkar, Dhruv Makwana, Gayatri Deshmukh, Sparsh Mittal, R. Sai Chandra Teja, Rekha Singhal
ICDAR (6)1
2023 GAFNet: A Global Fourier Self Attention Based Novel Network for multi-modal downstream tasks
abstract
In "vision and language" problems, multimodal inputs are simultaneously processed for combined visual and textual understanding for image-text embedding. In this paper, we discuss the necessity of considering the difference between the feature space and the distribution when performing multimodal learning. We deal with this problem through deep learning and a generative model approach. We introduce a novel network, GAFNet (Global Attention Fourier Net), which learns through large-scale pre-training over three image-text datasets (COCO, SBU, and CC-3M), for achieving high performance on downstream vision and language tasks. We propose a GAF (Global Attention Fourier) module, which integrates multiple modalities into one latent space. GAF module is independent of the type of modality, and it allows combining shared representations at each stage. Various ways of thinking about the relationships between different modalities directly affect the model’s design. In contrast to previous research, our work considers visual grounding as a pretrainable and transferable quality instead of something that must be trained from scratch. We show that GAFNet is a versatile network that can be used for a wide range of downstream tasks. Experimental results demonstrate that our technique achieves state-of-the-art performance on multimodal classification on the CrisisMD dataset and image generation on the COCO dataset. For image-text retrieval, our technique achieves competitive performance.
Onkar Susladkar, Gayatri Deshmukh, Dhruv Makwana, Sparsh Mittal, R. Sai Chandra Teja, Rekha Singhal
WACV1
2022 ClarifyNet: A high-pass and low-pass filtering based CNN for single image dehazing
Onkar Susladkar, Gayatri Deshmukh, Subhrajit Nag, Ananya Mantravadi, Dhruv Makwana, Sujitha Ravichandran, R. Sai Chandra Teja, Gajanan H. Chavhan, C. Krishna Mohan, Sparsh Mittal
J. Syst. Archit.1
2021 Study of Similarity Measures as Features in Classification for Answer Sentence Selection Task in Hindi Question Answering: Language-Specific v/s Other Measures
Devika Verma, Ramprasad S. Joshi, Shubhamkar Joshi, Onkar Susladkar
PACLIC4