EDBT 2026 Demo / reviewers in the wild / expert
Akshay Dudhane
dblp:213/7979 · also Akshay A. Dudhane
· DBLP profile ↗
25ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0002-0908-1930ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 11 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Clear Roads, Clear Vision: Advancements in Multi-Weather Restoration for Smart TransportationabstractAdverse weather conditions such as haze, rain, and snow significantly degrade the quality of images and videos, posing serious challenges to intelligent transportation systems that rely on visual input. These degradations affect critical applications including autonomous driving, traffic monitoring, and surveillance. This survey presents a comprehensive review of image and video restoration techniques developed to mitigate weather-induced visual impairments. We categorize existing approaches into traditional prior-based methods and modern data-driven models, including CNNs, transformers, diffusion models, and emerging vision-language models. Restoration strategies are further classified based on their scope: single-task models, multi-task/multi-weather systems, and all-in-one frameworks. In addition, we discuss day and night time restoration challenges, benchmark datasets, and evaluation protocols. The survey concludes by discussing current limitations and future directions, including unified restoration with downstream perception, real-time video restoration, and benchmarks for compound degradations under dynamic lighting. This work aims to serve as a valuable reference for advancing weather-resilient vision systems in smart transportation environments. Lastly, to keep pace with the rapid progress in this area, we will regularly update the latest relevant papers and their open-source implementations athttps://github.com/ChaudharyUPES/A-comprehensive-review-on-Multi-weather-restoration Vijay M. Galshetwar, Praful Hambarde, Prashant W. Patil, Akshay Dudhane, Sachin Chaudhary |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | EarthDial: Turning Multi-sensory Earth Observations to Interactive DialoguesabstractAutomated analysis of vast Earth observation data via interactive Vision-Language Models (VLMs) can unlock new opportunities for environmental monitoring, disaster response, and resource management. Existing generic VLMs do not perform well on Remote Sensing data, while the recent Geo-spatial VLMs remain restricted to a fixed resolution and few sensor modalities. In this paper, we introduce EarthDial, a conversational assistant specifically designed for Earth Observation (EO) data, transforming complex, multi-sensory Earth observations into interactive, natural language dialogues. EarthDial supports multi- spectral, multi-temporal, and multi-resolution imagery, enabling a wide range of remote sensing tasks, including classification, detection, captioning, question answering, visual reasoning, and visual grounding. To achieve this, we introduce an extensive instruction tuning dataset comprising over 11.11M instruction pairs covering RGB, Synthetic Aperture Radar (SAR), and multispectral modalities such as Near-Infrared (NIR) and infrared. Furthermore, EarthDial handles bi-temporal and multi-temporal sequence analysis for applications like change detection. Our extensive experimental results on 44 downstream datasets demonstrate that EarthDial outperforms existing generic and domain-specific models, achieving better generalization across various EO tasks. Our source codes and pre-trained models are at https://github.com/hiyamdebary/EarthDial. Sagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz, Muhammad Akhtar Munir, Muhammad Sohail Danish, Paolo Fraccaro, Campbell D. Watson, Levente J. Klein, Fahad Shahbaz Khan, Salman Khan 0001 |
CVPR | 2 |
| 2025 | Burst Image Restoration and EnhancementabstractBurst Image Restoration aims to reconstruct a high-quality image by efficiently combining complementary inter-frame information. However, it is quite challenging since individual burst images often have inter-frame misalignments that usually lead to ghosting and zipper artifacts. To mitigate this, we develop a novel approach for burst image processing named BIPNet that focuses solely on the information exchange between burst frames and filter-out the inherent degradations while preserving and enhancing the actual scene details. Our central idea is to generate a set of pseudo-burst features that combine complementary information from all the burst frames to exchange information seamlessly. However, due to inter-frame misalignment, the information cannot be effectively combined in pseudo-burst. Thus, we initially align the incoming burst features regarding the reference frame using the proposed edge-boosting feature alignment. Lastly, we progressively upscale the pseudo-burst features in multiple stages while adaptively combining the complementary information. Unlike the existing works, that usually deploy single-stage up-sampling with a late fusion scheme, we first deploy a pseudo-burst mechanism followed by the adaptive-progressive feature up-sampling. The proposed BIPNet significantly outperforms the existing methods on burst super-resolution, low-light image enhancement, low-light image super-resolution, and denoising tasks. Akshay Dudhane, Syed Waqas Zamir, Salman Khan 0001, Fahad Shahbaz Khan, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Burstormer: Burst Image Restoration and Enhancement TransformerabstractOn a shutter press, modern handheld cameras capture multiple images in rapid succession and merge them to gen-erate a single image. However, individual frames in a burst are misaligned due to inevitable motions and contain mul-tiple degradations. The challenge is to properly align the successive image shots and merge their complimentary in-formation to achieve high-quality outputs. Towards this direction, we propose Burstormer: a novel transformer-based architecture for burst image restoration and enhancement. In comparison to existing works, our approach exploits multi-scale local and non-local features to achieve improved alignment and feature fusion. Our key idea is to enable inter-frame communication in the burst neighborhoods for information aggregation and progressive fusion while modeling the burst-wide context. However, the input burst frames need to be properly aligned before fusing their information. Therefore, we propose an enhanced de-formable alignment module for aligning burst features with regards to the reference frame. Unlike existing methods, the proposed alignment module not only aligns burst features but also exchanges fea-ture information and maintains focused communication with the reference frame through the proposed reference-based feature enrichment mechanism, which facilitates handling complex motions. After multi-level alignment and enrichment, we re-emphasize on inter-frame communication within burst using a cyclic burst sampling module. Finally, the inter-frame information is aggre-gated using the proposed burst feature fusion module followed by progressive upsampling. Our Burstormer outperforms state-of-the-art methods on burst super-resolution, burst denoising and burst low-light enhance-ment. Our codes and pre-trained models are available at https://github.com/akshaydudhane16/Burstormer. Akshay Dudhane, Syed Waqas Zamir, Salman Khan 0001, Fahad Shahbaz Khan, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2023 | Gated Multi-Resolution Transfer Network for Burst Restoration and EnhancementabstractBurst image processing is becoming increasingly popular in recent years. However, it is a challenging task since individual burst images undergo multiple degradations and often have mutual misalignments resulting in ghosting and zipper artifacts. Existing burst restoration methods usually do not consider the mutual correlation and non-local contextual information among burst frames, which tends to limit these approaches in challenging cases. Another key challenge lies in the robust up-sampling of burst frames. The existing up-sampling methods cannot effectively utilize the advantages of single-stage and progressive up-sampling strategies with conventional and/or recent up-samplers at the same time. To address these challenges, we propose a novel Gated Multi-Resolution Transfer Network (GMTNet) to reconstruct a spatially precise high-quality image from a burst of low-quality raw images. GMT-Net consists of three modules optimized for burst processing tasks: Multi-scale Burst Feature Alignment (MBFA) for feature denoising and alignment, Transposed-Attention Feature Merging (TAFM) for multi-frame feature aggregation, and Resolution Transfer Feature Up-sampler (RTFU) to up-scale merged features and construct a high-quality output image. Detailed experimental analysis on five datasets validate our approach and sets a state-of-the-art for burst super-resolution, burst denoising, and low-light burst enhancement. Our codes and models are available at https://github.com/nanmehta/GMTNet. Nancy Mehta, Akshay Dudhane, M. Subrahmanyam 0001, Syed Waqas Zamir, Salman Khan 0001, Fahad Shahbaz Khan |
CVPR | 2 |
| 2022 | Deep Network for Extremely Low-Resolution Human Action RecognitionabstractDue to advancement in automated applications, privacy-preserving is an emerging concern. This concern is more significant in the case of human-centred surveillance application like human action recognition (HAR). Along with privacy concern, the computational complexity due to the huge size of video data is another major concern. To overcome these limitations, an attempt is made to examine the domain of human action recognition in low-resolution (LR) videos. The extremely LR video data ensures sufficient distortion in visual information to hide the identity of the person. Therefore, working with LR videos can resolve the above mentioned concerns of privacy preserving and computational complexity up to a certain extent. In this paper, a new generative adversarial network (GAN) based neural architecture is proposed for HAR in extremely low-resolution videos. The extensive results analysis with ablation study on the state-of-the-art datasets proves the effectiveness of the proposed method over the existing methods for LR-HAR. Sachin Chaudhary, Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001 |
AVSS | 3 |
| 2022 | AVisT: A Benchmark for Visual Object Tracking in Adverse Visibility
Mubashir Noman, Wafa Al Ghallabi, Daniya Kareem, Christoph Mayer 0007, Akshay Dudhane, Martin Danelljan, Hisham Cholakkal, Salman Khan 0001, Luc Van Gool, Fahad Shahbaz Khan |
BMVC | 5 |
| 2022 | Burst Image Restoration and EnhancementabstractModern handheld devices can acquire burst image sequence in a quick succession. However, the individual acquired frames suffer from multiple degradations and are misaligned due to camera shake and object motions. The goal of Burst Image Restoration is to effectively combine complimentary cues across multiple burst frames to generate high-quality outputs. Towards this goal, we develop a novel approach by solely focusing on the effective information exchange between burst frames, such that the degradations get filtered out while the actual scene details are preserved and enhanced. Our central idea is to create a set of pseudo-burst features that combine complimentary information from all the input burst frames to seamlessly exchange information. However, the pseudo-burst cannot be successfully created unless the individual burst frames are properly aligned to discount inter-frame movements. Therefore, our approach initially extracts pre-processed features from each burst frame and matches them using an edge-boosting burst alignment module. The pseudo-burst features are then created and enriched using multi-scale contextual information. Our final step is to adaptively aggregate information from the pseudo-burst features to progressively increase resolution in multiple stages while merging the pseudo-burst features. In comparison to existing works that usually follow a late fusion scheme with single-stage upsampling, our approach performs favorably, delivering state-of-the-art performance on burst super-resolution, burst low-light image enhancement and burst denoising tasks. The source code and pre-trained models are available at https://github.com/akshaydudhane16/BIPNet. Akshay Dudhane, Syed Waqas Zamir, Salman Khan 0001, Fahad Shahbaz Khan, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2022 | Multi-frame based adversarial learning approach for video surveillance
Prashant W. Patil, Akshay Dudhane, Sachin Chaudhary, M. Subrahmanyam 0001 |
Pattern Recognit. | 2 |
| 2021 | Multi-frame Recurrent Adversarial Network for Moving Object SegmentationabstractMoving object segmentation (MOS) in different practical scenarios like weather degraded, dynamic background, etc. videos is a challenging and high demanding task for various computer vision applications. Existing supervised approaches achieve remarkable performance with complicated training or extensive fine-tuning or inappropriate training-testing data distribution. Also, the generalized effect of existing works with completely unseen data is difficult to identify. In this work, the recurrent feature sharing based generative adversarial network is proposed with unseen video analysis. The proposed network comprises of dilated convolution to extract the spatial features at multiple scales. Along with the temporally sampled multiple frames, previous frame output is considered as input to the network. As the motion is very minute between the two consecutive frames, the previous frame decoder features are shared with encoder features recurrently for current frame foreground segmentation. This recurrent feature sharing of different layers helps the encoder network to learn the hierarchical interactions between the motion and appearance-based features. Also, the learning of the proposed network is concentrated in different ways, like disjoint and global training-testing for MOS. An extensive experimental analysis of the proposed network is carried out on two benchmark video datasets with seen and unseen MOS video. Qualitative and quantitative experimental study shows that the proposed network outperforms the existing methods. Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001 |
WACV | 2 |
| 2021 | Motion estimation in hazy videos
Sachin Chaudhary, Akshay Dudhane, Prashant W. Patil, M. Subrahmanyam 0001, Sanjay N. Talbar |
Pattern Recognit. Lett. | 2 |
| 2021 | Deep Adversarial Network for Scene Independent Moving Object SegmentationabstractThe current prevailing algorithms highly depend on additional pre-trained modules trained for other applications or complicated training procedures or neglect the inter-frame spatio-temporal structural dependencies. Also, the generalized effect of existing works with completely unseen data is difficult to identify. Specifically, the outdoor videos suffer from adverse atmospheric conditions like poor visibility, inclement weather, etc. In this letter, a novel end-to-end multi-scale temporal edge aggregation (MTPA) network is proposed with adversarial learning for scene dependent and independent object segmentation. The MTPA is proposed to extract the comprehensive spatio-temporal features from the current and reference frame. These MTPA features are used to guide the respective decoder through skip connections. To get authentic and consistent foreground object(s), the respective scale feedback of previous frame output is provided with respective MTPA features at each decoder input. The performance analysis of the proposed method is verified on CDnet-2014 and LASIESTA video datasets. The proposed method outperforms the existing state-of-the-art methods with scene dependent and independent analysis. Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001, Anil Balaji Gonde |
IEEE Signal Process. Lett. | 2 |
| 2021 | An Unified Recurrent Video Object Segmentation Framework for Various Surveillance EnvironmentsabstractMoving object segmentation (MOS) in videos received considerable attention because of its broad security-based applications like robotics, outdoor video surveillance, self-driving cars, etc. The current prevailing algorithms highly depend on additional trained modules for other applications or complicated training procedures or neglect the inter-frame spatio-temporal structural dependencies. To address these issues, a simple, robust, and effective unified recurrent edge aggregation approach is proposed for MOS, in which additional trained modules or fine-tuning on a test video frame(s) are not required. Here, a recurrent edge aggregation module (REAM) is proposed to extract effective foreground relevant features capturing spatio-temporal structural dependencies with encoder and respective decoder features connected recurrently from previous frame. These REAM features are then connected to a decoder through skip connections for comprehensive learning named as temporal information propagation. Further, the motion refinement block with multi-scale dense residual is proposed to combine the features from the optical flow encoder stream and the last REAM module for holistic feature learning. Finally, these holistic features and REAM features are given to the decoder block for segmentation. To guide the decoder block, previous frame output with respective scales is utilized. The different configurations of training-testing techniques are examined to evaluate the performance of the proposed method. Specifically, outdoor videos often suffer from constrained visibility due to different environmental conditions and other small particles in the air that scatter the light in the atmosphere. Thus, comprehensive result analysis is conducted on six benchmark video datasets with different surveillance environments. We demonstrate that the proposed method outperforms the state-of-the-art methods for MOS without any pre-trained module, fine-tuning on the test video frame(s) or complicated training. Prashant W. Patil, Akshay Dudhane, Ashutosh Kulkarni, M. Subrahmanyam 0001, Anil Balaji Gonde, Sunil Gupta 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Varicolored Image De-HazingabstractThe quality of images captured in bad weather is often affected by chromatic casts and low visibility due to the presence of atmospheric particles. Restoration of the color balance is often ignored in most of the existing image de-hazing methods. In this paper, we propose a varicolored end-to-end image de-hazing network which restores the color balance in a given varicolored hazy image and recovers the haze-free image. The proposed network comprises of 1) Haze color correction (HCC) module and 2) Visibility improvement (VI) module. The proposed HCC module provides required attention to each color channel and generates a color balanced hazy image. While the proposed VI module processes the color balanced hazy image through novel inception attention block to recover the haze-free image. We also propose a novel approach to generate a large-scale varicolored synthetic hazy image database. An ablation study has been carried out to demonstrate the effect of different factors on the performance of the proposed network for image de-hazing. Three benchmark synthetic datasets have been used for quantitative analysis of the proposed network. Visual results on a set of real-world hazy images captured in different weather conditions demonstrate the effectiveness of the proposed approach for varicolored image de-hazing. Akshay Dudhane, Kuldeep Marotirao Biradar, Prashant W. Patil, Praful Hambarde, M. Subrahmanyam 0001 |
CVPR | 1 |
| 2020 | An End-to-End Edge Aggregation Network for Moving Object SegmentationabstractMoving object segmentation in videos (MOS) is a highly demanding task for security-based applications like automated outdoor video surveillance. Most of the existing techniques proposed for MOS are highly depend on fine-tuning a model on the first frame(s) of test sequence or complicated training procedure, which leads to limited practical serviceability of the algorithm. In this paper, the inherent correlation learning-based edge extraction mechanism (EEM) and dense residual block (DRB) are proposed for the discriminative foreground representation. The multi-scale EEM module provides the efficient foreground edge related information (with the help of encoder) to the decoder through skip connection at subsequent scale. Further, the response of the optical flow encoder stream and the last EEM module are embedded in the bridge network. The bridge network comprises of multi-scale residual blocks with dense connections to learn the effective and efficient foreground relevant features. Finally, to generate accurate and consistent foreground object maps, a decoder block is proposed with skip connections from respective multi-scale EEM module feature maps and the subsequent down-sampled response of previous frame output. Specifically, the proposed network does not require any pre-trained models or fine-tuning of the parameters with the initial frame(s) of the test video. The performance of the proposed network is evaluated with different configurations like disjoint, cross-data, and global training-testing techniques. The ablation study is conducted to analyse each model of the proposed network. To demonstrate the effectiveness of the proposed framework, a comprehensive analysis on four benchmark video datasets is conducted. Experimental results show that the proposed approach outperforms the state-of-the-art methods for MOS Prashant W. Patil, Kuldeep Marotirao Biradar, Akshay Dudhane, M. Subrahmanyam 0001 |
CVPR | 3 |
| 2020 | Depth Estimation From Single Image And Semantic PriorabstractThe multi-modality sensor fusion technique is an active research area in scene understating. In this work, we explore the RGB image and semantic-map fusion methods for depth estimation. The LiDARs, Kinect, and TOF depth sensors are unable to predict the depth-map at illuminate and monotonous pattern surface. In this paper, we propose a semantic-to-depth generative adversarial network (S2D-GAN) for depth estimation from RGB image and its semantic-map. In the first stage, the proposed S2D-GAN estimates the coarse level depthmap using a semantic-to-coarse-depth generative adversarial network (S2CD-GAN) while the second stage estimates the fine-level depth-map using a cascaded multi-scale spatial pooling network. The experimental analysis of the proposed S2D-GAN performed on NYU-Depth-V2 dataset shows that the proposed S2D-GAN gives outstanding result over existing single image depth estimation and RGB with sparse samples methods. The proposed S2D-GAN also gives efficient results on the real-world indoor and outdoor image depth estimation. Praful Hambarde, Akshay Dudhane, Prashant W. Patil, M. Subrahmanyam 0001, Abhinav Dhall |
ICIP | 2 |
| 2020 | Deep Underwater Image Restoration and BeyondabstractUnderwater image restoration is a challenging problem due to the multiple distortions. Degradation in the information is mainly due to the 1) light scattering effect 2) wavelength dependent color attenuation and 3) object blurriness effect. In this letter, we propose a novel end-to-end deep network for underwater image restoration. The proposed network is divided into two parts viz. channel-wise color feature extraction module and dense-residual feature extraction module. A custom loss function is proposed, which preserves the structural details and generates the true edge information in the restored underwater scene. Also, to train the proposed network for underwater image enhancement, a new synthetic underwater image database is proposed. Existing synthetic underwater database images are characterized by light scattering and color attenuation distortions. However, object blurriness effect is ignored. We, on the other hand, introduced the blurring effect along with the light scattering and color attenuation distortions. The proposed network is validated for underwater image restoration task on real-world underwater images. Experimental analysis shows that the proposed network is superior than the existing state-of-the-art approaches for underwater image restoration. Akshay Dudhane, Praful Hambarde, Prashant W. Patil, M. Subrahmanyam 0001 |
IEEE Signal Process. Lett. | 1 |
| 2020 | RYF-Net: Deep Fusion Network for Single Image Haze RemovalabstractHaze removal from a single image is a challenging task. Estimation of accurate scene transmission map (TrMap) is the key to reconstruct the haze-free scene. In this paper, we propose a convolutional neural network based architecture to estimate the TrMap of the hazy scene. The proposed network takes the hazy image as an input and extracts the haze relevant features using proposed RNet and YNet through RGB and YCbCr color spaces respectively and generates two TrMaps. Further, we propose a novel TrMap fusion network (FNet) to integrate two TrMaPs and estimate robust TrMap for the hazy scene. To analyze the robustness of FNet, we tested it on combinations of TrMaps obtained from existing state-of-the-art methods. Performance evaluation of the proposed approach has been carried out using the structural similarity index, mean square error and peak signal to noise ratio. We conduct experiments on five datasets namely: D-HAZY ancuti2016d, Imagenet deng2009imagenet, Indoor SOTS li2017reside, HazeRD zhang2017hazerd and set of real-world hazy images. Performance analysis shows that the proposed approach outperforms the existing state-of-the-art methods for single image dehazing. Further, we extended our work to address high-level vision task such as object detection in hazy scenes. It is observed that there is a significant improvement in accurate object detection in hazy scenes using proposed approach. Akshay Dudhane, M. Subrahmanyam 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Pose Guided Dynamic Image Network for Human Action Recognition in Person Centric VideosabstractThe most emerging concerns in computer vision are size of data to process and privacy preserving of the end user. Camera sensors are all around us these days, recording and analysing our day-to-day activities. In this scenario the privacy perseverance becomes a question of concern especially in case of devices working on the basis of human action recognition (HAR). Another important concern in computer vision is the size of data. The surveillance requires continues transfer of huge amount of data through the network. The processing time required to transfer the video to central server and analyses the video directly depends on the resolution of the video. The research in computer vision is exploring the possibility of working on different aspects of videos such as using only pose information or representing whole video using a single frame for the purpose of HAR. Here, an attempt is made to explore the concept of pose estimation and video representation using dynamic image to solve the dual purpose of privacy preserving and decreasing the load on network for transfer of videos over the network for analysis. In this paper, a new Pose Guided Dynamic Image (PDI) network is proposed for HAR which is capable of providing a summarized single frame for the person's activity in any given video. Unlike dynamic image network, this approach considers only the person's motion and discards the background motion. Therefore, PDI provides more specific information required for HAR as compared to the dynamic image. Also, by summarizing the video, the identity of the person remains masked. The proposed method is able to provide better result on both of the benchmark datasets used namely JHMDB and UCF-sports for the experimentation. Sachin Chaudhary, Akshay Dudhane, Prashant W. Patil, M. Subrahmanyam 0001 |
AVSS | 2 |
| 2019 | Image and Video Super Resolution using Recurrent Generative Adversarial NetworkabstractRecently, the convolutional neural network with residual learning models achieves high accuracy for single image super-resolution with different scale factors. With adversarial learning model, effective learning of transformation function for the low-resolution input image to a high-resolution target image can be achieved. In this paper, we propose a method for image and video super-resolution using the recurrent generative adversarial network named SR2GAN. In the proposed model (SR2GAN) we use recursive learning for video super-resolution to overcome the difficulty of learning transformation function for synthesizing realistic high-resolution images. This recursive approach helps to reduce the parameters with increasing depth of the model. An extensive evaluation is performed to examine the effectiveness of the proposed model, which shows that SR2GAN performs better in terms of peak signal to noise ratio (PSNR) and structural self-similarity index (SSIM) as compared to the state-of-the-art methods for super-resolution. For source code and supplementary material visit: https://github.com/OmkarThawakar/SR2GAN/. Omkar Thawakar, Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001, Uday Kulkarni |
AVSS | 3 |
| 2019 | Single Image Depth Estimation Using Deep Adversarial TrainingabstractScene understanding is an active area of research in computer vision that encompasses several different problems. The LiDARs and stereo depth sensor have their own restrictions such as light sensitiveness, power consumption and short-range [1]. In this paper, we propose a two-stream deep adversarial network for single image depth estimation in RGB images. For stream I network, we propose a novel encoder-decoder architecture using residual concepts to extract course-level depth features. Stream II network purely processes the information through the residual architecture for fine-level depth estimation. Also, we designed a feature map sharing architecture to share the learned feature maps of the decoder module of stream I. Sharing feature maps strengthen the residual learning to estimate the scene depth and increase the robustness of the proposed network. A benchmark NYU RGB-D v2 [2] database is used to evaluate the proposed network for single image depth estimation. Both qualitative and quantitative analysis has been carried out to analyze the effectiveness of the proposed network for scene depth prediction. Performance analysis shows that the proposed method outperforms other existing methods for single image depth estimation. Praful Hambarde, Akshay Dudhane, M. Subrahmanyam 0001 |
ICIP | 2 |
| 2019 | Motion Saliency Based Generative Adversarial Network for Underwater Moving Object SegmentationabstractThe underwater moving object segmentation is a challenging task. The problems like absorbing, scattering and attenuation of light rays between the scene and the imaging platform degrades the visibility of image or video frames. Also, the back-scattering of light rays further increases the problem of underwater video analysis, because the light rays interact with underwater particles and scattered back to the sensor. In this paper, a novel Motion Saliency Based Generative Adversarial Network (GAN) for Underwater Moving Object Segmentation (MOS) is proposed. The proposed network comprises of both identity mapping and dense connections for underwater MOS. To the best of our knowledge, this is the first paper with the concept of GAN-based unpaired learning for MOS in underwater videos. Initially, current frame motion saliency is estimated using few initial video frames and current frame. Further, estimated motion saliency is given as input to the proposed network for foreground estimation. To examine the effectiveness of proposed network, the Fish4Knowledge [1] underwater video dataset and challenging video categories of ChangeDetection.net-2014 [2] datasets are considered. The segmentation accuracy of existing state-of-the-art methods are used for comparison with proposed approach in terms of average F-measure. From experimental results, it is evident that the proposed network shows significant improvement as compared to the existing state-of-the-art methods for MOS. Prashant W. Patil, Omkar Thawakar, Akshay Dudhane, M. Subrahmanyam 0001 |
ICIP | 3 |
| 2019 | CDNet: Single Image De-Hazing Using Unpaired Adversarial TrainingabstractOutdoor scene images generally undergo visibility degradation in presence of aerosol particles such as haze, fog and smoke. The reason behind this is, aerosol particles scatter the light rays reflected from the object surface and thus results in attenuation of light intensity. Effect of haze is inversely proportional to the transmission coefficient of the scene point. Thus, estimation of accurate transmission map (TrMap) is a key step to reconstruct the haze-free scene. Previous methods used various assumptions/priors to estimate the scene TrMap. Also, available end-to-end dehazing approaches make use of supervised training to anticipate the TrMap on synthetically generated paired hazy images. Despite the success of previous approaches, they fail in real-world extreme vague conditions due to unavailability of the real-world hazy image pairs for training the network. Thus, in this paper, Cycle-consistent generative adversarial network for single image De-hazing named as CDNet is proposed which is trained in an unpaired manner on real-world hazy image dataset. Generator network of CDNet comprises of encoder-decoder architecture which aims to estimate the object level TrMap followed by optical model to recover the haze-free scene. We conduct experiments on four datasets namely: D-HAZY [1], Imagenet [5], SOTS [20] and real-world images. Structural similarity index, peak signal to noise ratio and CIEDE2000 metric are used to evaluate the performance of the proposed CDNet. Experiments on benchmark datasets show that the proposed CDNet outperforms the existing state-of-the-art methods for single image haze removal. Akshay Dudhane, M. Subrahmanyam 0001 |
WACV | 1 |
| 2019 | Cardinal color fusion network for single image haze removal
Akshay Dudhane, M. Subrahmanyam 0001 |
Mach. Vis. Appl. | 1 |
| 2018 | C^2MSNet: A Novel Approach for Single Image Haze RemovalabstractDegradation of image quality due to the presence of haze is a very common phenomenon. Existing DehazeNet [3], MSCNN [11] tackled the drawbacks of hand crafted haze relevant features. However, these methods have the problem of color distortion in gloomy (poor illumination) environment. In this paper, a cardinal (red, green and blue) color fusion network for single image haze removal is proposed. In first stage, network fusses color information present in hazy images and generates multi-channel depth maps. The second stage estimates the scene transmission map from generated dark channels using multi channel multi scale convolutional neural network (McMs-CNN) to recover the original scene. To train the proposed network, we have used two standard datasets namely: ImageNet [5] and D-HAZY [1]. Performance evaluation of the proposed approach has been carried out using structural similarity index (SSIM), mean square error (MSE) and peak signal to noise ratio (PSNR). Performance analysis shows that the proposed approach outperforms the existing state-of-the-art methods for single image dehazing. Akshay Dudhane, M. Subrahmanyam 0001 |
WACV | 1 |