Ashutosh Kulkarni

dblp:80/1984 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-7265-3540ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DTMIR-Pro: Domain Translation with Prompt-based Latent-Space Generalization for Multi-Weather Image Restoration
abstract
Multi-weather image restoration seeks to recover scene visibility under rainy, snowy, and hazy conditions, thereby enhancing high-level vision tasks. Existing methods typically train on combined datasets with single-type weather degradations, limiting their generalization to real-world scenarios involving mixed degradations. Domain translation has emerged as a viable solution by generating diverse weather-degraded variants of the same scene. However, current approaches require separate models for each degradation type, resulting in increased system complexity. To address this, we propose DTMIR-Pro, a prompt-based domain translation framework with latent space generalization for multi-weather image restoration. A single trainable network performs multi-domain translation using domain-adaptive prompts and dynamic kernel selection via a proposed Dynamic Multi-Head Attention block to learn diverse degradation patterns. The restoration network takes translated outputs and employs a Multi-Weather Fusion Block with global-local feature streams to capture complex degradations. Furthermore, we introduce a Similarity-Based Encoder Routing mechanism to transfer domain-specific features from the translation encoder to the restoration stage. Extensive experiments on both synthetic and real-world weather-degraded datasets demonstrate the effectiveness and generalizability of the proposed method. The code is made available at https://github.com/AshutoshKulkarni4998/DTMIR-Pro.
Ashutosh Kulkarni, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Balasubramanian Raman
WACV1
2025 Phaseformer: Phase-Based Attention Mechanism for Underwater Image Restoration and Beyond
abstract
Quality degradation is observed in underwater images due to the effects of light refraction and absorption by water, leading to issues like color cast, haziness, and limited visibility. This degradation negatively affects the performance of autonomous underwater vehicles used in marine applications. To address these challenges, we propose a lightweight phase-based transformer network with 1.77M parameters for underwater image restoration (UIR). Our approachfocuses on effectively extracting non-contaminated features using a phase-based self-attention mechanism. We also introduce an optimized phase attention block to restore structural information by propagating prominent attentive features from the input. We evaluate our method on both synthetic (UIEB, UFO-120) and real-world (UIEB, U45, UCCS, SQUID) underwater image datasets. Additionally, we demonstrate its effectiveness for low-light image enhancement using the LOL dataset. Through extensive ab-lation studies and comparative analysis, it is clear that the proposed approach outperforms existing state-of-the-art (SOTA) methods. Code is available at Phaseformer.
Md Raqib Khan, Anshul Negi, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001
WACV3
2024 Zero Reference based Low-light Enhancement with Wavelet Optimization
abstract
Images captured in low light conditions usually suffer from poor visibility, a high amount of noise, and little information stored in the dark image, which has a negative impact on subsequent processing for outdoor computer vision applications. Presently, numerous deep learning based methods achieved superior performance with multi-exposure paired training data or additional information. However, obtaining multi-exposure data samples is a tedious task in real-time scenarios. To mitigate this challenge, we propose a zero reference based learnable wavelet approach without multi-exposure paired training data requirement for low-light image enhancement. Our proposed approach generates the low light image and learns to project an image into noise free similar looking image, then we enhance the image using retinex theory. Further, we have proposed learnable wavelet block to remove the hidden noise amplified while enhancement. We introduce Gaussian-based supervision to improve the smoothness of the image. Extensive experimental analysis on synthetic as well as real-world images, along with thorough ablation study demonstrate the effectiveness of our proposed method over the existing state-of-the-art methods for low-light image enhancement. The code is provided at https://github.com/vision-lab-sggsiet/Zero-Reference-based-Low-light-Enhancement-with-Wavelet-Optimization.
Vivek Deshmukh, Adinath Madhavrao Dukre, Ashutosh Kulkarni, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Anil Balaji Gonde
AVSS3
2024 Frequency Modulated Deformable Transformer for Underwater Image Enhancement
Adinath Madhavrao Dukre, Vivek Deshmukh, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, Anil Balaji Gonde, M. Subrahmanyam 0001
ICPR (32)3
2024 Attentive Color Fusion Transformer Network (ACFTNet) for Underwater Image Enhancement
Mohd Ubaid Wani, Md Raqib Khan, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001
ICPR (21)3
2024 C2AIR: Consolidated Compact Aerial Image Haze Removal
abstract
Aerial image haze removal deals with improving the visibility and quality of images captured from aerial platforms, such as drones and satellites. Aerial images are commonly used in various applications such as environmental monitoring, and disaster response. These applications usually require cleaner data for accurate functioning. However, atmospheric conditions such as haze or fog can significantly degrade the quality of these images, reducing their contrast, color saturation, and sharpness, making it difficult to extract meaningful information from them. Existing methods rely on computationally heavy and haze density (light, moderate, dense) specific architectures for aerial image dehazing. In light of these limitations, we propose a novel lightweight and consolidated approach for aerial image dehazing. In this approach, we propose Density Aware Query Modulated Block for learning weather degradations in input features and guiding the restoration process. Further, we propose Cross Collaborative Feed-Forward Block for learning to restore varying sizes of the structures in the input images. Finally, we propose Gated Adaptive Feature Fusion block to achieve inter-scale and intra-feature attentive fusion, effective for aerial image restoration. Extensive analysis on benchmark aerial image dehazing datasets and real-world images, along with detailed ablation studies validate the effectiveness of the proposed approach. Further, we have analysed our method for other restoration task such as underwater image enhancement to experiment its wide applicability. The code is available at https://github.com/AshutoshKulkarni4998/C2AIR.
Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001
WACV1
2023 Underwater Image Enhancement with Phase Transfer and Attention
abstract
Underwater pictures typically suffer from substantial deterioration due to the refraction and absorption of light by water, including color cast, hazy blur, and limited visibility. Such degradation in visibility eventually reduces the effectiveness of marine applications installed on autonomous underwater vehicles. Hence, an efficient pre-processing step is required for the significant performance of these applications. As a solution, underwater image enhancement (UIE) mainly focuses on enhancing the visibility of degraded images along with restoring crucial details. Existing methods generally utilize (a) complex cascaded architectures, (b) different degradation-prone color spaces, and (c) direct skip connections that pass irrelevant content. In light of this, we propose a lightweight transformer network with 1.7M parameters (1/6thof the existing state-of-the-art method) consisting of the proposed gray-scale attention and phase transformer block for UIE. A gray-scale attention block is proposed for the effective extraction of non-contaminated features. Further, a phase transfer block is proposed for effectively restoring the structural information in the outputs by propagating most relevant and undegraded features from the inputs. A comprehensive evaluation of the proposed method on synthetic (EUVP, UIEB) and real-world (UIEB, UCCS) image datasets as well as extensive ablation studies confirm its effectiveness over existing state-of-the-art approaches. The source code is provided at: https://github.com/Mdraqibkhan/UIEPTA.
Md Raqib Khan, Ashutosh Kulkarni, Shruti S. Phutke, M. Subrahmanyam 0001
IJCNN2
2023 Aerial Image Dehazing with Attentive Deformable Transformers
abstract
Aerial imagery is widely utilized in visual data dependent applications such as military surveillance, earthquake assessment, etc. For these applications, minute texture in the aerial image are essential as any disturbance can cause inaccurate prediction. However, atmospheric haze severely reduces the visibility of the scene to be analysed, and hence takes a toll on accuracy of higher level applications. Existing methods either utilize additional prior while training, or produce sub-optimal outputs on different densities of haze degradation, due to absence of local and global dependencies in the extracted features. Therefore, it is essential to have a texture preserving algorithm for aerial image dehazing. In light of this, we propose a work that introduces a novel deformable multi-head attention with spatially attentive offset extraction based solution for aerial image dehazing. Here, the deformable multi-head attention is introduced to reconstruct fine level texture in the restored image. We also introduce spatially attentive offset extractor in the deformable convolution for focusing on relevant contextual information. Further, edge boosting skip connections are proposed for effectively passing edge features from shallow layers to deeper layers of the network. Thorough experimentation on synthetic as well as real-world data, along with extensive ablation study, demonstrate that the proposed method outperforms the prevailing works on aerial image dehazing. The code is provided at https://github.com/AshutoshKulkarni4998/AIDTransformer.
Ashutosh Kulkarni, M. Subrahmanyam 0001
WACV1
2023 Unified Multi-Weather Visibility Restoration
abstract
Automated surveillance is widely opted for appli- cations such as traffic monitoring, vehicle identification, etc. But, various weather degradation factors such as rain and snow streaks, along with atmospheric veil severely affect the perceptual quality of an image, eventually affecting the performance of these applications. There exist weather specific (rain, haze, snow, etc.) methods focusing on respective restoration task. As image restoration is a preprocessing step for high level surveillance applications, it is practically inapplicable to have different architectures for different weather restoration. In this paper, we propose a lightweight unified network, having 1.1 M parameters (1/40th and 1/6th of the existing rain with veil removal, and snow with veil removal methods respectively) for removal of rain and snow along with the veiling effect present in the images. In this network, we propose two parallel streams to handle the degradations and restoration: First, degradation removal stream (DRS) focuses mainly on removing randomly repeating degradations i.e., rain and snow streaks, through the proposed adaptive multi-scale feature sharing block (AMFSB) and stage-wise subtractive block (SSB). Second, feature corrector stream (FCS) mainly focuses on refining the partial outputs of the first stream, reducing the veiling effect and acts supplementary to the first stream. Finally, we leverage contrastive regularization for better convergence of the proposed network. Substantial experiments on synthetic as well as real-world images, along with extensive ablation studies, demonstrate that the proposed method performs competitively with the existing methods for multi-weather image restoration. The code is available at:https://github.com/AshutoshKulkarni4998/UVRNet.
Ashutosh Kulkarni, Prashant W. Patil, M. Subrahmanyam 0001, Sunil Gupta 0001
IEEE Trans. Multim.1
2022 Consolidated Adversarial Network for Video De-raining and De-hazing
abstract
The performance of recent video enhancement methods is superior in specific hazy, rainy, snowy, and foggy weather conditions. However, these approaches can handle degradation rendered by single weather. We propose an integrated lightweight adversarial learning network to handle the degradations induced by different weather conditions. This is a unique approach to mitigate the problem of video restoration for multi- weather degraded videos using single network. The proposed architecture combines the idea of multi-resolution analysis with a multi-scale encoder and domain-specific feature learning is achieved using domain-aware filtering modules. The architecture provides recurrent feature sharing for temporal consistency, achieved by feeding the previous frame output as feedback. Substantial experiments on various datasets demonstrate that the proposed method performs competitively with the existing state-of-the-art approaches for video restoration in multi-weather conditions.
Vijay M. Galshetwar, Ashutosh Kulkarni, Sachin Chaudhary
AVSS2
2022 Robust Unseen Video Understanding for Various Surveillance Environments
abstract
Automated video-based applications are a highly demanding technique from a security perspective, where detection of moving objects i.e., moving object segmentation (MOS) is performed. Therefore, we have proposed an effective solution with a spatio-temporal squeeze excitation mechanism (SqEm) based multi-level feature sharing encoder-decoder network for MOS. Here, the SqEm module is proposed to get prominent foreground edge information using spatio-temporal features. Further, a multi-level feature sharing residual decoder module is proposed with respective SqEm features and previous output features for accurate and consistent foreground segmentation. To handle the foreground or background class imbalance issue, we propose a region of interest-based edge loss. The extensive experimental analysis on three databases is conducted. Result analysis and ablation study proved the robustness of the proposed network for unseen video understanding over SOTA methods.
Prashant W. Patil, Jasdeep Singh, Praful Hambarde, Ashutosh Kulkarni, Sachin Chaudhary, M. Subrahmanyam 0001
AVSS4
2022 Progressive Subtractive Recurrent Lightweight Network for Video Deraining
abstract
Presence of rainy artifacts severely degrade the overall visual quality of a video and tend to overlap with the useful information present in the video frames. This degraded video affects the effectiveness of many automated applications like traffic monitoring, surveillance,etc.As video deraining is a pre-processing step for automated applications, it is highly demanded to have a lightweight deraining module. Therefore, in this paper, a“Progressive Subtractive Recurrent Lightweight Network”is proposed for video deraining. Initially, the Multi-Kernel feature Sharing Residual Block (MKSRB) is designed to learn different sizes of rain streaks which facilitates the complete removal of rain streaks through progressive subtractions. These MKSRB features are merged with previous frame output recurrently to maintain the temporal consistency. Further, multi-receptive feature subtraction is performed through Multi-scale Multi-Receptive Difference Block (MMRDB) to avoid loss of details and extract high-frequency information. Finally, progressively learned features through MKSRB and recurrent feature merging are aggregated with fused MMRDB features which outputs the rain-free frame. Substantial experiments on prevailing synthetic datasets and real-world videos verify the superior performance of the proposed method over the existing state-of-the-art methods for video deraining.
Ashutosh Kulkarni, Prashant W. Patil, M. Subrahmanyam 0001
IEEE Signal Process. Lett.1
2022 WiperNet: A Lightweight Multi-Weather Restoration Network for Enhanced Surveillance
abstract
Adherent raindrops, rainstreaks and snow severely degrade the perceptual quality of an image, eventually affecting the performance of several computer vision based applications which are applied in outdoor scenarios, e.g., traffic monitoring, autonomous driving, etc. Due to the complex appearance properties, removal of such degradations from an image is a challenging task. Working towards mitigating this problem, in this paper, a lightweight network named as WiperNet is proposed which tackles the problem of raindrops, rain streaks and snow removal present in an image. The WiperNet makes use of the proposed Dual Restoration (DR) mechanism, where the input features are processed twice through the network. In the network, Multi-scale Context Aware Residual Block (MCARB) is proposed for integrating contextual information from various scales. Also, Adaptive Varying Receptive Fusion Block (AVRFB) is proposed for adaptively fusing the information acquired through different dilation rates. Finally, we propose a Feature Refinement Stream which makes use of multiple kernel sizes of convolution filters and spatio-channel attention blocks for focusing on relevant information for effective removal of the degradations while using the coarse outputs of the features from the initial layers of the network. Substantial experiments and ablation study scrutinize that the proposed lightweight WiperNet outperforms the existing state-of-the-art methods for raindrop, rain streak and snow removal. The code is provided athttps://github.com/AshutoshKulkarni4998/WiperNet.
Ashutosh Kulkarni, M. Subrahmanyam 0001
IEEE Trans. Intell. Transp. Syst.1
2021 An Unified Recurrent Video Object Segmentation Framework for Various Surveillance Environments
abstract
Moving object segmentation (MOS) in videos received considerable attention because of its broad security-based applications like robotics, outdoor video surveillance, self-driving cars, etc. The current prevailing algorithms highly depend on additional trained modules for other applications or complicated training procedures or neglect the inter-frame spatio-temporal structural dependencies. To address these issues, a simple, robust, and effective unified recurrent edge aggregation approach is proposed for MOS, in which additional trained modules or fine-tuning on a test video frame(s) are not required. Here, a recurrent edge aggregation module (REAM) is proposed to extract effective foreground relevant features capturing spatio-temporal structural dependencies with encoder and respective decoder features connected recurrently from previous frame. These REAM features are then connected to a decoder through skip connections for comprehensive learning named as temporal information propagation. Further, the motion refinement block with multi-scale dense residual is proposed to combine the features from the optical flow encoder stream and the last REAM module for holistic feature learning. Finally, these holistic features and REAM features are given to the decoder block for segmentation. To guide the decoder block, previous frame output with respective scales is utilized. The different configurations of training-testing techniques are examined to evaluate the performance of the proposed method. Specifically, outdoor videos often suffer from constrained visibility due to different environmental conditions and other small particles in the air that scatter the light in the atmosphere. Thus, comprehensive result analysis is conducted on six benchmark video datasets with different surveillance environments. We demonstrate that the proposed method outperforms the state-of-the-art methods for MOS without any pre-trained module, fine-tuning on the test video frame(s) or complicated training.
Prashant W. Patil, Akshay Dudhane, Ashutosh Kulkarni, M. Subrahmanyam 0001, Anil Balaji Gonde, Sunil Gupta 0001
IEEE Trans. Image Process.3
1998 Modeling and Analysis of The Difference-Bit Cache
abstract
Advances in VLSI technology and processor architectures have resulted in a tremendous increase in processor speeds and memory capacities. However, memory latencies have failed to improve as rapidly, making memory systems the performance bottlenecks in most high performance processor architectures. Caching is a time-tested mechanism to solve this speed disparity. Among the different cache mapping strategies, direct mapping is the only configuration where the critical path is merely the time required to access a RAM. Although direct mapped caches are preferable considering hit-access times, they have poor hit ratios compared to associative caches. The difference-bit cache proposed by Juan, Lang and Navarro (1996), is functionally equivalent to a two-way set-associative cache but tries to achieve an access time smaller than that of a conventional two-way set-associative cache and close to that of a direct-mapped cache. We modeled and analyzed the difference-bit cache to prove the hypothesis of its small access time. We have also tried to prove that the access time advantage of the difference-bit cache improves over the conventional two-way set-associative cache with an increase in the cache size. Finally we have tried to analyze the trade-off involved in applying these techniques to a higher associativity cache.
Ashutosh Kulkarni, Navin Chander, Soumya Pillai, Lizy Kurian John
Great Lakes Symposium on VLSI1