EDBT 2026 Demo / reviewers in the wild / expert
Prashant W. Patil
dblp:254/4281
· DBLP profile ↗
28ranked-venue papers
14as first author
18since 2021 · last 2026
0000-0003-2604-6501ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DTMIR-Pro: Domain Translation with Prompt-based Latent-Space Generalization for Multi-Weather Image RestorationabstractMulti-weather image restoration seeks to recover scene visibility under rainy, snowy, and hazy conditions, thereby enhancing high-level vision tasks. Existing methods typically train on combined datasets with single-type weather degradations, limiting their generalization to real-world scenarios involving mixed degradations. Domain translation has emerged as a viable solution by generating diverse weather-degraded variants of the same scene. However, current approaches require separate models for each degradation type, resulting in increased system complexity. To address this, we propose DTMIR-Pro, a prompt-based domain translation framework with latent space generalization for multi-weather image restoration. A single trainable network performs multi-domain translation using domain-adaptive prompts and dynamic kernel selection via a proposed Dynamic Multi-Head Attention block to learn diverse degradation patterns. The restoration network takes translated outputs and employs a Multi-Weather Fusion Block with global-local feature streams to capture complex degradations. Furthermore, we introduce a Similarity-Based Encoder Routing mechanism to transfer domain-specific features from the translation encoder to the restoration stage. Extensive experiments on both synthetic and real-world weather-degraded datasets demonstrate the effectiveness and generalizability of the proposed method. The code is made available at https://github.com/AshutoshKulkarni4998/DTMIR-Pro. Ashutosh Kulkarni, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Balasubramanian Raman |
WACV | 2 |
| 2026 | MemeTAG: Keyword-Driven Meme Classification through Tag Embedding ReconstructionabstractThe proliferation of harmful internet memes poses a significant societal threat, yet their automated classification remains a formidable algorithmic challenge due to the nuanced, multimodal nature of their content. To address this, we introduce MemeTAG, a novel dual-objective framework that pioneers a keyword-aware approach to meme classification. Our core innovation is a two-part semantic guidance mechanism: first, we leverage a pretrained Vision-Language Model to generate a set of descriptive keywords, that capture the high-level semantics. Second, we introduce the Aggregated Tag Inference Network (ATIN), an attention-based module that distills these keywords into a single, rich semantic embedding. This embedding serves as a target for a novel auxiliary reconstruction loss, which compels the model to learn deeply aligned visual and textual features. This approach, combined with an efficient three-stage training strategy, establishes a new state-of-the-art on the HarMeme, Hateful Memes Challenge (HMC), and PrideMM datasets, decisively outperforming existing state-of-the-art methods. Akshit Sharma, Prashant W. Patil |
WACV | 2 |
| 2026 | Clear Roads, Clear Vision: Advancements in Multi-Weather Restoration for Smart TransportationabstractAdverse weather conditions such as haze, rain, and snow significantly degrade the quality of images and videos, posing serious challenges to intelligent transportation systems that rely on visual input. These degradations affect critical applications including autonomous driving, traffic monitoring, and surveillance. This survey presents a comprehensive review of image and video restoration techniques developed to mitigate weather-induced visual impairments. We categorize existing approaches into traditional prior-based methods and modern data-driven models, including CNNs, transformers, diffusion models, and emerging vision-language models. Restoration strategies are further classified based on their scope: single-task models, multi-task/multi-weather systems, and all-in-one frameworks. In addition, we discuss day and night time restoration challenges, benchmark datasets, and evaluation protocols. The survey concludes by discussing current limitations and future directions, including unified restoration with downstream perception, real-time video restoration, and benchmarks for compound degradations under dynamic lighting. This work aims to serve as a valuable reference for advancing weather-resilient vision systems in smart transportation environments. Lastly, to keep pace with the rapid progress in this area, we will regularly update the latest relevant papers and their open-source implementations athttps://github.com/ChaudharyUPES/A-comprehensive-review-on-Multi-weather-restoration Vijay M. Galshetwar, Praful Hambarde, Prashant W. Patil, Akshay Dudhane, Sachin Chaudhary |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Unpaired recurrent learning for real-world video de-hazingabstractAutomated outdoor vision-based applications have become increasingly in demand for day-to-day life. Bad weather like haze, rain, snow, etc. may limit the reliability of these applications due to degradation in the overall video quality. So, there is a dire need to pre-process the weather-degraded videos before they are fed to downstream applications. Researchers generally adopt synthetically generated paired hazy frames for learning the task of video de-hazing. The models trained solely on synthetic data may have limited performance on different types of real-world hazy scenarios due to significant domain gap between synthetic and real-world hazy videos. One possible solution is to prove the generalization ability by training on unpaired data for video de-hazing. Some unpaired learning approaches are proposed for single image de-hazing. However, these unpaired single image de-hazing approaches compromise the performance in terms of temporal consistency, which is important for video de-hazing tasks. With this motivation, we have proposed a lightweight and temporally consistent architecture for video de-hazing tasks. To achieve this, diverse receptive and multi-scale features at various input resolutions are mixed and aggregated with multi-kernel attention to extract significant haze information. Furthermore, we propose a recurrent multi-attentive feature alignment concept to maintain temporal consistency with recurrent feedback of previously restored frames for temporal consistent video restoration . Comprehensive experiments are conducted on real-world and synthetic video databases (REVIDE and RSA100Haze). Both the qualitative and quantitative results show significant improvement of the proposed network with better temporal consistency over state-of-the-art methods for detailed video restoration in hazy weather. Source code is available at: https://github.com/pwp1208/UnpairedVideoDehazing . Prashant W. Patil, Santosh Nagnath Randive, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh, M. Subrahmanyam 0001 |
Pattern Recognit. | 1 |
| 2024 | Zero Reference based Low-light Enhancement with Wavelet OptimizationabstractImages captured in low light conditions usually suffer from poor visibility, a high amount of noise, and little information stored in the dark image, which has a negative impact on subsequent processing for outdoor computer vision applications. Presently, numerous deep learning based methods achieved superior performance with multi-exposure paired training data or additional information. However, obtaining multi-exposure data samples is a tedious task in real-time scenarios. To mitigate this challenge, we propose a zero reference based learnable wavelet approach without multi-exposure paired training data requirement for low-light image enhancement. Our proposed approach generates the low light image and learns to project an image into noise free similar looking image, then we enhance the image using retinex theory. Further, we have proposed learnable wavelet block to remove the hidden noise amplified while enhancement. We introduce Gaussian-based supervision to improve the smoothness of the image. Extensive experimental analysis on synthetic as well as real-world images, along with thorough ablation study demonstrate the effectiveness of our proposed method over the existing state-of-the-art methods for low-light image enhancement. The code is provided at https://github.com/vision-lab-sggsiet/Zero-Reference-based-Low-light-Enhancement-with-Wavelet-Optimization. Vivek Deshmukh, Adinath Madhavrao Dukre, Ashutosh Kulkarni, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Anil Balaji Gonde |
AVSS | 4 |
| 2024 | AeroDehazeNet: Exploiting Selective Multi-Scale Transformers for Aerial Image DehazingabstractRemote sensing is the task of analyzing and acquiring useful information from satellite images captured at a far distance from the earth’s surface. These images are vulnerable to degradation due to the presence of mist or haze. Existing methods either make use of prior information to estimate haze free images, or use CNN architectures based on generative adversarial networks (GANs) or Transformers. Though the state-of-the-art transformer-based architectures helped to dehaze the aerial images, they lacked the ability to capture multi-scale dependencies of the image. Identifying this shortcoming, we propose AeroDehazeNet based on a transformer that captures multi-scale dependencies along with global dependencies of the image. Our network comprises of three key components: (1) a multi-scale selective attention (MScA) network to attentively process the multi-scale information in an image, (2) residual attention network (RAN in feed forward network responsible for distilling non-degraded features passed from MScA, and (3) high frequency dominant skip connection (HFDS) block for passing diverse features (low frequency and high frequency) prominent with multi-scale edge features from encoder levels to adjacent decoder levels. The extensive quantitative and qualitative comparisons with existing methods on synthetic and realworld data plus exhaustive ablation study demonstrate the efficacy of our proposed network over transformer based state-of-the-art architectures with comparatively less number of parameters and FLOPs. Testing code is available at https://github.com/KartikGonde/AeroDehazeNet. Kartik Gonde, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Pramod Patil, Vinod V. Kimbahune |
AVSS | 2 |
| 2023 | Multi-weather Image Restoration via Domain TranslationabstractWeather degraded conditions such as rain, haze, snow, etc. may degrade the performance of most computer vision systems. Therefore, effective restoration of multi-weather degraded images is an essential prerequisite for successful functioning of such systems. The current multi-weather image restoration approaches utilize a model that is trained on a combined dataset consisting of individual images for rainy, snowy, and hazy weather degradations. These methods may face challenges when dealing with real-world situations where the images may have multiple, more intricate weather conditions. To address this issue, we propose a domain translation-based unified method for multi-weather image restoration. In this approach, the proposed network learns multiple weather degradations simultaneously, making it immune for real-world conditions. Specifically, we first propose an instance-level domain (weather) translation with multi-attentive feature learning approach to get different weather-degraded variants of the same scenario. Next, the original and translated images are used as input to the proposed novel multi-weather restoration network which utilizes a progressive multi-domain deformable alignment (PMDA) with cascaded multi-head attention (CMA). The proposed PMDA facilitates the restoration network to learn weather-invariant clues effectively. Further, PMDA and respective decoder features are merged via proposed CMA module for restoration. Extensive experimental results on synthetic and real-world hazy, rainy, and snowy image databases clearly demonstrate that our model outperforms the state-of-the-art multi-weather image restoration methods. Code is available at https://github.com/pwp1208/Domain_Translation_Multi-weather_Restoration. Prashant W. Patil, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh, M. Subrahmanyam 0001 |
ICCV | 1 |
| 2023 | Unified Multi-Weather Visibility RestorationabstractAutomated surveillance is widely opted for appli- cations such as traffic monitoring, vehicle identification, etc. But, various weather degradation factors such as rain and snow streaks, along with atmospheric veil severely affect the perceptual quality of an image, eventually affecting the performance of these applications. There exist weather specific (rain, haze, snow, etc.) methods focusing on respective restoration task. As image restoration is a preprocessing step for high level surveillance applications, it is practically inapplicable to have different architectures for different weather restoration. In this paper, we propose a lightweight unified network, having 1.1 M parameters (1/40th and 1/6th of the existing rain with veil removal, and snow with veil removal methods respectively) for removal of rain and snow along with the veiling effect present in the images. In this network, we propose two parallel streams to handle the degradations and restoration: First, degradation removal stream (DRS) focuses mainly on removing randomly repeating degradations i.e., rain and snow streaks, through the proposed adaptive multi-scale feature sharing block (AMFSB) and stage-wise subtractive block (SSB). Second, feature corrector stream (FCS) mainly focuses on refining the partial outputs of the first stream, reducing the veiling effect and acts supplementary to the first stream. Finally, we leverage contrastive regularization for better convergence of the proposed network. Substantial experiments on synthetic as well as real-world images, along with extensive ablation studies, demonstrate that the proposed method performs competitively with the existing methods for multi-weather image restoration. The code is available at:https://github.com/AshutoshKulkarni4998/UVRNet. Ashutosh Kulkarni, Prashant W. Patil, M. Subrahmanyam 0001, Sunil Gupta 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Deep Network for Extremely Low-Resolution Human Action RecognitionabstractDue to advancement in automated applications, privacy-preserving is an emerging concern. This concern is more significant in the case of human-centred surveillance application like human action recognition (HAR). Along with privacy concern, the computational complexity due to the huge size of video data is another major concern. To overcome these limitations, an attempt is made to examine the domain of human action recognition in low-resolution (LR) videos. The extremely LR video data ensures sufficient distortion in visual information to hide the identity of the person. Therefore, working with LR videos can resolve the above mentioned concerns of privacy preserving and computational complexity up to a certain extent. In this paper, a new generative adversarial network (GAN) based neural architecture is proposed for HAR in extremely low-resolution videos. The extensive results analysis with ablation study on the state-of-the-art datasets proves the effectiveness of the proposed method over the existing methods for LR-HAR. Sachin Chaudhary, Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001 |
AVSS | 2 |
| 2022 | Robust Unseen Video Understanding for Various Surveillance EnvironmentsabstractAutomated video-based applications are a highly demanding technique from a security perspective, where detection of moving objects i.e., moving object segmentation (MOS) is performed. Therefore, we have proposed an effective solution with a spatio-temporal squeeze excitation mechanism (SqEm) based multi-level feature sharing encoder-decoder network for MOS. Here, the SqEm module is proposed to get prominent foreground edge information using spatio-temporal features. Further, a multi-level feature sharing residual decoder module is proposed with respective SqEm features and previous output features for accurate and consistent foreground segmentation. To handle the foreground or background class imbalance issue, we propose a region of interest-based edge loss. The extensive experimental analysis on three databases is conducted. Result analysis and ablation study proved the robustness of the proposed network for unseen video understanding over SOTA methods. Prashant W. Patil, Jasdeep Singh, Praful Hambarde, Ashutosh Kulkarni, Sachin Chaudhary, M. Subrahmanyam 0001 |
AVSS | 1 |
| 2022 | Video Restoration Framework and Its Meta-adaptations to Data-Poor Conditions
Prashant W. Patil, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh |
ECCV (28) | 1 |
| 2022 | Multi-frame based adversarial learning approach for video surveillance
Prashant W. Patil, Akshay Dudhane, Sachin Chaudhary, M. Subrahmanyam 0001 |
Pattern Recognit. | 1 |
| 2022 | Dual-frame spatio-temporal feature modulation for video enhancement
Prashant W. Patil, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh |
Pattern Recognit. | 1 |
| 2022 | Progressive Subtractive Recurrent Lightweight Network for Video DerainingabstractPresence of rainy artifacts severely degrade the overall visual quality of a video and tend to overlap with the useful information present in the video frames. This degraded video affects the effectiveness of many automated applications like traffic monitoring, surveillance,etc.As video deraining is a pre-processing step for automated applications, it is highly demanded to have a lightweight deraining module. Therefore, in this paper, a“Progressive Subtractive Recurrent Lightweight Network”is proposed for video deraining. Initially, the Multi-Kernel feature Sharing Residual Block (MKSRB) is designed to learn different sizes of rain streaks which facilitates the complete removal of rain streaks through progressive subtractions. These MKSRB features are merged with previous frame output recurrently to maintain the temporal consistency. Further, multi-receptive feature subtraction is performed through Multi-scale Multi-Receptive Difference Block (MMRDB) to avoid loss of details and extract high-frequency information. Finally, progressively learned features through MKSRB and recurrent feature merging are aggregated with fused MMRDB features which outputs the rain-free frame. Substantial experiments on prevailing synthetic datasets and real-world videos verify the superior performance of the proposed method over the existing state-of-the-art methods for video deraining. Ashutosh Kulkarni, Prashant W. Patil, M. Subrahmanyam 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Multi-frame Recurrent Adversarial Network for Moving Object SegmentationabstractMoving object segmentation (MOS) in different practical scenarios like weather degraded, dynamic background, etc. videos is a challenging and high demanding task for various computer vision applications. Existing supervised approaches achieve remarkable performance with complicated training or extensive fine-tuning or inappropriate training-testing data distribution. Also, the generalized effect of existing works with completely unseen data is difficult to identify. In this work, the recurrent feature sharing based generative adversarial network is proposed with unseen video analysis. The proposed network comprises of dilated convolution to extract the spatial features at multiple scales. Along with the temporally sampled multiple frames, previous frame output is considered as input to the network. As the motion is very minute between the two consecutive frames, the previous frame decoder features are shared with encoder features recurrently for current frame foreground segmentation. This recurrent feature sharing of different layers helps the encoder network to learn the hierarchical interactions between the motion and appearance-based features. Also, the learning of the proposed network is concentrated in different ways, like disjoint and global training-testing for MOS. An extensive experimental analysis of the proposed network is carried out on two benchmark video datasets with seen and unseen MOS video. Qualitative and quantitative experimental study shows that the proposed network outperforms the existing methods. Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001 |
WACV | 1 |
| 2021 | Motion estimation in hazy videos
Sachin Chaudhary, Akshay Dudhane, Prashant W. Patil, M. Subrahmanyam 0001, Sanjay N. Talbar |
Pattern Recognit. Lett. | 3 |
| 2021 | Deep Adversarial Network for Scene Independent Moving Object SegmentationabstractThe current prevailing algorithms highly depend on additional pre-trained modules trained for other applications or complicated training procedures or neglect the inter-frame spatio-temporal structural dependencies. Also, the generalized effect of existing works with completely unseen data is difficult to identify. Specifically, the outdoor videos suffer from adverse atmospheric conditions like poor visibility, inclement weather, etc. In this letter, a novel end-to-end multi-scale temporal edge aggregation (MTPA) network is proposed with adversarial learning for scene dependent and independent object segmentation. The MTPA is proposed to extract the comprehensive spatio-temporal features from the current and reference frame. These MTPA features are used to guide the respective decoder through skip connections. To get authentic and consistent foreground object(s), the respective scale feedback of previous frame output is provided with respective MTPA features at each decoder input. The performance analysis of the proposed method is verified on CDnet-2014 and LASIESTA video datasets. The proposed method outperforms the existing state-of-the-art methods with scene dependent and independent analysis. Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001, Anil Balaji Gonde |
IEEE Signal Process. Lett. | 1 |
| 2021 | An Unified Recurrent Video Object Segmentation Framework for Various Surveillance EnvironmentsabstractMoving object segmentation (MOS) in videos received considerable attention because of its broad security-based applications like robotics, outdoor video surveillance, self-driving cars, etc. The current prevailing algorithms highly depend on additional trained modules for other applications or complicated training procedures or neglect the inter-frame spatio-temporal structural dependencies. To address these issues, a simple, robust, and effective unified recurrent edge aggregation approach is proposed for MOS, in which additional trained modules or fine-tuning on a test video frame(s) are not required. Here, a recurrent edge aggregation module (REAM) is proposed to extract effective foreground relevant features capturing spatio-temporal structural dependencies with encoder and respective decoder features connected recurrently from previous frame. These REAM features are then connected to a decoder through skip connections for comprehensive learning named as temporal information propagation. Further, the motion refinement block with multi-scale dense residual is proposed to combine the features from the optical flow encoder stream and the last REAM module for holistic feature learning. Finally, these holistic features and REAM features are given to the decoder block for segmentation. To guide the decoder block, previous frame output with respective scales is utilized. The different configurations of training-testing techniques are examined to evaluate the performance of the proposed method. Specifically, outdoor videos often suffer from constrained visibility due to different environmental conditions and other small particles in the air that scatter the light in the atmosphere. Thus, comprehensive result analysis is conducted on six benchmark video datasets with different surveillance environments. We demonstrate that the proposed method outperforms the state-of-the-art methods for MOS without any pre-trained module, fine-tuning on the test video frame(s) or complicated training. Prashant W. Patil, Akshay Dudhane, Ashutosh Kulkarni, M. Subrahmanyam 0001, Anil Balaji Gonde, Sunil Gupta 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Varicolored Image De-HazingabstractThe quality of images captured in bad weather is often affected by chromatic casts and low visibility due to the presence of atmospheric particles. Restoration of the color balance is often ignored in most of the existing image de-hazing methods. In this paper, we propose a varicolored end-to-end image de-hazing network which restores the color balance in a given varicolored hazy image and recovers the haze-free image. The proposed network comprises of 1) Haze color correction (HCC) module and 2) Visibility improvement (VI) module. The proposed HCC module provides required attention to each color channel and generates a color balanced hazy image. While the proposed VI module processes the color balanced hazy image through novel inception attention block to recover the haze-free image. We also propose a novel approach to generate a large-scale varicolored synthetic hazy image database. An ablation study has been carried out to demonstrate the effect of different factors on the performance of the proposed network for image de-hazing. Three benchmark synthetic datasets have been used for quantitative analysis of the proposed network. Visual results on a set of real-world hazy images captured in different weather conditions demonstrate the effectiveness of the proposed approach for varicolored image de-hazing. Akshay Dudhane, Kuldeep Marotirao Biradar, Prashant W. Patil, Praful Hambarde, M. Subrahmanyam 0001 |
CVPR | 3 |
| 2020 | An End-to-End Edge Aggregation Network for Moving Object SegmentationabstractMoving object segmentation in videos (MOS) is a highly demanding task for security-based applications like automated outdoor video surveillance. Most of the existing techniques proposed for MOS are highly depend on fine-tuning a model on the first frame(s) of test sequence or complicated training procedure, which leads to limited practical serviceability of the algorithm. In this paper, the inherent correlation learning-based edge extraction mechanism (EEM) and dense residual block (DRB) are proposed for the discriminative foreground representation. The multi-scale EEM module provides the efficient foreground edge related information (with the help of encoder) to the decoder through skip connection at subsequent scale. Further, the response of the optical flow encoder stream and the last EEM module are embedded in the bridge network. The bridge network comprises of multi-scale residual blocks with dense connections to learn the effective and efficient foreground relevant features. Finally, to generate accurate and consistent foreground object maps, a decoder block is proposed with skip connections from respective multi-scale EEM module feature maps and the subsequent down-sampled response of previous frame output. Specifically, the proposed network does not require any pre-trained models or fine-tuning of the parameters with the initial frame(s) of the test video. The performance of the proposed network is evaluated with different configurations like disjoint, cross-data, and global training-testing techniques. The ablation study is conducted to analyse each model of the proposed network. To demonstrate the effectiveness of the proposed framework, a comprehensive analysis on four benchmark video datasets is conducted. Experimental results show that the proposed approach outperforms the state-of-the-art methods for MOS Prashant W. Patil, Kuldeep Marotirao Biradar, Akshay Dudhane, M. Subrahmanyam 0001 |
CVPR | 1 |
| 2020 | Depth Estimation From Single Image And Semantic PriorabstractThe multi-modality sensor fusion technique is an active research area in scene understating. In this work, we explore the RGB image and semantic-map fusion methods for depth estimation. The LiDARs, Kinect, and TOF depth sensors are unable to predict the depth-map at illuminate and monotonous pattern surface. In this paper, we propose a semantic-to-depth generative adversarial network (S2D-GAN) for depth estimation from RGB image and its semantic-map. In the first stage, the proposed S2D-GAN estimates the coarse level depthmap using a semantic-to-coarse-depth generative adversarial network (S2CD-GAN) while the second stage estimates the fine-level depth-map using a cascaded multi-scale spatial pooling network. The experimental analysis of the proposed S2D-GAN performed on NYU-Depth-V2 dataset shows that the proposed S2D-GAN gives outstanding result over existing single image depth estimation and RGB with sparse samples methods. The proposed S2D-GAN also gives efficient results on the real-world indoor and outdoor image depth estimation. Praful Hambarde, Akshay Dudhane, Prashant W. Patil, M. Subrahmanyam 0001, Abhinav Dhall |
ICIP | 3 |
| 2020 | Deep Underwater Image Restoration and BeyondabstractUnderwater image restoration is a challenging problem due to the multiple distortions. Degradation in the information is mainly due to the 1) light scattering effect 2) wavelength dependent color attenuation and 3) object blurriness effect. In this letter, we propose a novel end-to-end deep network for underwater image restoration. The proposed network is divided into two parts viz. channel-wise color feature extraction module and dense-residual feature extraction module. A custom loss function is proposed, which preserves the structural details and generates the true edge information in the restored underwater scene. Also, to train the proposed network for underwater image enhancement, a new synthetic underwater image database is proposed. Existing synthetic underwater database images are characterized by light scattering and color attenuation distortions. However, object blurriness effect is ignored. We, on the other hand, introduced the blurring effect along with the light scattering and color attenuation distortions. The proposed network is validated for underwater image restoration task on real-world underwater images. Experimental analysis shows that the proposed network is superior than the existing state-of-the-art approaches for underwater image restoration. Akshay Dudhane, Praful Hambarde, Prashant W. Patil, M. Subrahmanyam 0001 |
IEEE Signal Process. Lett. | 3 |
| 2019 | Pose Guided Dynamic Image Network for Human Action Recognition in Person Centric VideosabstractThe most emerging concerns in computer vision are size of data to process and privacy preserving of the end user. Camera sensors are all around us these days, recording and analysing our day-to-day activities. In this scenario the privacy perseverance becomes a question of concern especially in case of devices working on the basis of human action recognition (HAR). Another important concern in computer vision is the size of data. The surveillance requires continues transfer of huge amount of data through the network. The processing time required to transfer the video to central server and analyses the video directly depends on the resolution of the video. The research in computer vision is exploring the possibility of working on different aspects of videos such as using only pose information or representing whole video using a single frame for the purpose of HAR. Here, an attempt is made to explore the concept of pose estimation and video representation using dynamic image to solve the dual purpose of privacy preserving and decreasing the load on network for transfer of videos over the network for analysis. In this paper, a new Pose Guided Dynamic Image (PDI) network is proposed for HAR which is capable of providing a summarized single frame for the person's activity in any given video. Unlike dynamic image network, this approach considers only the person's motion and discards the background motion. Therefore, PDI provides more specific information required for HAR as compared to the dynamic image. Also, by summarizing the video, the identity of the person remains masked. The proposed method is able to provide better result on both of the benchmark datasets used namely JHMDB and UCF-sports for the experimentation. Sachin Chaudhary, Akshay Dudhane, Prashant W. Patil, M. Subrahmanyam 0001 |
AVSS | 3 |
| 2019 | Image and Video Super Resolution using Recurrent Generative Adversarial NetworkabstractRecently, the convolutional neural network with residual learning models achieves high accuracy for single image super-resolution with different scale factors. With adversarial learning model, effective learning of transformation function for the low-resolution input image to a high-resolution target image can be achieved. In this paper, we propose a method for image and video super-resolution using the recurrent generative adversarial network named SR2GAN. In the proposed model (SR2GAN) we use recursive learning for video super-resolution to overcome the difficulty of learning transformation function for synthesizing realistic high-resolution images. This recursive approach helps to reduce the parameters with increasing depth of the model. An extensive evaluation is performed to examine the effectiveness of the proposed model, which shows that SR2GAN performs better in terms of peak signal to noise ratio (PSNR) and structural self-similarity index (SSIM) as compared to the state-of-the-art methods for super-resolution. For source code and supplementary material visit: https://github.com/OmkarThawakar/SR2GAN/. Omkar Thawakar, Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001, Uday Kulkarni |
AVSS | 2 |
| 2019 | Motion Saliency Based Generative Adversarial Network for Underwater Moving Object SegmentationabstractThe underwater moving object segmentation is a challenging task. The problems like absorbing, scattering and attenuation of light rays between the scene and the imaging platform degrades the visibility of image or video frames. Also, the back-scattering of light rays further increases the problem of underwater video analysis, because the light rays interact with underwater particles and scattered back to the sensor. In this paper, a novel Motion Saliency Based Generative Adversarial Network (GAN) for Underwater Moving Object Segmentation (MOS) is proposed. The proposed network comprises of both identity mapping and dense connections for underwater MOS. To the best of our knowledge, this is the first paper with the concept of GAN-based unpaired learning for MOS in underwater videos. Initially, current frame motion saliency is estimated using few initial video frames and current frame. Further, estimated motion saliency is given as input to the proposed network for foreground estimation. To examine the effectiveness of proposed network, the Fish4Knowledge [1] underwater video dataset and challenging video categories of ChangeDetection.net-2014 [2] datasets are considered. The segmentation accuracy of existing state-of-the-art methods are used for comparison with proposed approach in terms of average F-measure. From experimental results, it is evident that the proposed network shows significant improvement as compared to the existing state-of-the-art methods for MOS. Prashant W. Patil, Omkar Thawakar, Akshay Dudhane, M. Subrahmanyam 0001 |
ICIP | 1 |
| 2019 | FgGAN: A Cascaded Unpaired Learning for Background Estimation and Foreground SegmentationabstractThe moving object segmentation (MOS) in videos with bad weather, irregular motion of objects, camera jitter, shadow and dynamic background scenarios is still an open problem for computer vision applications. To address these issues, in this paper, we propose an approach named as Foreground Generative Adversarial Network (FgGAN) with the recent concepts of generative adversarial network (GAN) and unpaired training for background estimation and foreground segmentation. To the best of our knowledge, this is the first paper with the concept of GAN-based unpaired learning for MOS. Initially, video-wise background is estimated using GAN-based unpaired learning network (network-I). Then, to extract the motion information related to foreground, motion saliency is estimated using estimated background and current video frame. Further, estimated motion saliency is given as input to the GANbased unpaired learning network (network-II) for foreground segmentation. To examine the effectiveness of proposed FgGAN (cascaded networks I and II), the challenging video categories like dynamic background, bad weather, intermittent object motion and shadow are collected from ChangeDetection.net-2014 [26] database. The segmentation accuracy is observed qualitatively and quantitatively in terms of F-measure and percentage of wrong classification (PWC) and compared with the existing state-of-the-art methods. From experimental results, it is evident that the proposed FgGAN shows significant improvement in terms of F-measure and PWC as compared to the existing stateof-the-art methods for MOS. Prashant W. Patil, M. Subrahmanyam 0001 |
WACV | 1 |
| 2019 | MSFgNet: A Novel Compact End-to-End Deep Network for Moving Object DetectionabstractMoving object detection (MOD) in videos is a challenging task. Estimation of accurate background is the key to extracting the foreground from video frames. In this paper, we have proposed a novel compact end-to-end convolutional neural network architecture, motion saliency foreground network (MSFgNet), to estimate the background and to extract the foreground from video frames. Initially, the long streaming video is divided into a number of small video streams (SVS). The proposed network takes the SVS as an input and estimates the background frame for each SVS. Second, the saliency map is extracted using the current video frame and estimated background. Furthermore, a compact encoder-decoder network is proposed to extract the foreground from the estimated saliency maps. The performance of the proposed MSFgNet is tested on three benchmark datasets (CDnet-2014, LASIESTA, and PTIS) for MOD. The computational complexity (handling of number of parameters and execution time) and the performance of the proposed MSFgNet are compared with the existing state-of-the-art methods for MOD in terms of precision, recall, and F-measure. Performance analysis shows that the proposed network is very compact and outperforms the existing state-of-the-art methods for MOD in videos. Prashant W. Patil, M. Subrahmanyam 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2018 | MsEDNet: Multi-Scale Deep Saliency Learning for Moving Object DetectionabstractMoving object detection (foreground and background) is an important problem in computer vision. Most of the works in this problem are based on background subtraction. However, these approaches are not able to handle scenarios with infrequent motion of object, illumination changes, shadow, camouflage etc. To overcome these, here a two stage robust and compact method for moving object detection (MOD) is proposed. In first stage, to generate the saliency map, background image is estimated using a temporal histogram technique with the help of several input frames. In the second stage, multiscale encoder-decoder network is used to learn multiscale semantic feature of estimated saliency for foreground extraction. The encoder is used to extract multi-scale features from multi-scale saliency map. The decoder part is designed to learn the mapping of low resolution multi-scale features into high resolution output frame. To observe the efficacy of proposed MsEDNet, experiments are conducted on two benchmark datasets (change detection (CDnet-2014) [1] and Wallflower [2]) for MOD. The precision, recall and F-measure are used as performance parameter for comparison with the existing state-of-the-art methods. Experimental results show a significant improvement in detection accuracy and decrement in execution time as compared to the state-of-the-art methods for MOD. Prashant W. Patil, M. Subrahmanyam 0001, Abhinav Dhall, Sachin Chaudhary |
SMC | 1 |