VLDB 2026 Research / reviewers in the wild / expert
M. Subrahmanyam 0001
dblp:61/10849 · also Subrahmanyam Murala
· DBLP profile ↗
83ranked-venue papers
8as first author
47since 2021 · last 2026
0000-0003-3384-4368ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 53 · 4 first-author · 33 since 2021Artificial intelligence and machine learning · 33 · 3 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QuEENet: Quantum-Enhanced Expressive Network for Image ClassificationabstractThis paper presents QuEENet, a hybrid quantum-classical architecture for image classification that incorporates parameterized quantum circuits within a convolution neural network. This study investigates how quantum circuit expressivity and entanglement strategies influence classification performance, with a focus on configurations involving a CNOT gate followed by a rotational gate on the target qubit. Non-Clifford gates, such as Rx/Ry/Rzsupports larger state-space coverage and expressivity in quantum models. The proposed QuEENet explored the aspect of non-Clifford gates in parameterized quantum circuits. While non-Clifford gates are theoretically critical for universal quantum computation, but their role in image classification task is unexplored. Experimental results across multiple benchmark datasets suggest that while increased expressivity via non-Clifford gates can be beneficial, it should be carefully balanced with circuit interpretability and trainability. QuEENet demonstrates that hybrid models can leverage quantum circuits not merely as architectural novelties, but as controllable modules for enhancing learning in classical pipelines. An extensive ablation study was conducted across multiple datasets to highlight the effects of Clifford and non-Clifford gate combinations and entanglement configurations. Shashank Bayal, Rushikesh Govind Dawane, Komal, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 5 |
| 2026 | DTMIR-Pro: Domain Translation with Prompt-based Latent-Space Generalization for Multi-Weather Image RestorationabstractMulti-weather image restoration seeks to recover scene visibility under rainy, snowy, and hazy conditions, thereby enhancing high-level vision tasks. Existing methods typically train on combined datasets with single-type weather degradations, limiting their generalization to real-world scenarios involving mixed degradations. Domain translation has emerged as a viable solution by generating diverse weather-degraded variants of the same scene. However, current approaches require separate models for each degradation type, resulting in increased system complexity. To address this, we propose DTMIR-Pro, a prompt-based domain translation framework with latent space generalization for multi-weather image restoration. A single trainable network performs multi-domain translation using domain-adaptive prompts and dynamic kernel selection via a proposed Dynamic Multi-Head Attention block to learn diverse degradation patterns. The restoration network takes translated outputs and employs a Multi-Weather Fusion Block with global-local feature streams to capture complex degradations. Furthermore, we introduce a Similarity-Based Encoder Routing mechanism to transfer domain-specific features from the translation encoder to the restoration stage. Extensive experiments on both synthetic and real-world weather-degraded datasets demonstrate the effectiveness and generalizability of the proposed method. The code is made available at https://github.com/AshutoshKulkarni4998/DTMIR-Pro. Ashutosh Kulkarni, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Balasubramanian Raman |
WACV | 4 |
| 2026 | Editorial to special issue on selected extended works from 9th international conference on computer vision & image processing (CVIP) 2024
Mohan Kankanhalli, Balasubramanian Raman, M. Subrahmanyam 0001, Jagadeesh Kakarla, Sambit Bakshi |
Image Vis. Comput. | 3 |
| 2026 | FedHC: Enhanced federated learning with Hessian and cosine correlation for proximal correlation
Kushall Singh, Monu Verma, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, G. Shankara Raju Kosuru, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb |
Knowl. Based Syst. | 6 |
| 2026 | FedHMed: Adaptive progressive loss and KL-divergence regularization for federated heterogeneous medical image classification tasks
Kushall Singh, Monu Verma, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb |
Knowl. Based Syst. | 5 |
| 2026 | ME-NAS: A Micro Expression Feature Adaptive Neural Architecture SearchabstractConvolution neural networks (CNN) have emerged as a prevailing paradigm for micro-expression recognition (MER) yet, it is inefficient and time-intensive to design optimal CNN-based MER models manually. In recent times, the neural architecture search (NAS) has garnered attention due to its automatic CNN architecture searching ability. However, the performance of NAS in MER is limited by challenges such as rapid duration, subtle intensity, and a mismatch between architecture and cell-level search. The existing search space, which stacks 12 cells with 3 transition paths (downsample, upsample, and same resolution), creates deep networks that may diminish minute spatiotemporal features due to progressive convolution and pooling. Therefore, motivated by these factors, in this article, we introduce a novel approach, the Micro-Expression Feature Adaptive NAS (ME-NAS), to analyze true human emotions through MER. While NAS has gained attention for its automatic CNN architecture search ability, its application in MER faces challenges due to ingrained challenges (rapid duration, subtle and low intensity) and the discrepancy between architecture and cell-level search. The existing NAS architecture search space is designed by stacking 12 cells with 3 transition paths (downsample, upsample, and same resolution), resulting in a deep network. Such deep networks may diminish minute spatiotemporal features due to the progressive convolution and pooling operations. Motivated by these factors, we designed a new NAS algorithm: ME-NAS. The ME-NAS comprises f (EXPERT) in architecture search, along with refined and complementary feature derivative (ReCODE) operations in cell-level search. The EXPERT aims to trace the optimal paths instead of covering all possible paths between cells. The ReCODE operations capture micro-level variations from spatial and temporal domains by introducing 24 3D convolution operations. The proposed ReCODE and EXPERT search space jointly lead to the search for a robust and shallow CNN architecture for micro-expressions (MEs). The proposed ME-NAS is evaluated on six datasets: CASME-I, CASME-II, CAS(ME) \({}^{2}\) , SAMM, SMIC, and MEGC-19 composite, with two evaluation strategies: LOSO and cross-domain, respectively. The experimental results manifest that the proposed ME-NAS outperformed the state-of-the-art approaches on both evaluation strategies. Monu Verma, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Mohamed Abdel-Mottaleb |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2025 | TRUST: Time-Domain Residual Unsupervised Stability Technique for Improved Heart Rate EstimationabstractCamera-based estimation of vital signs is a promising method for non-contact health monitoring, which analyzes minute changes in video data. However, the creation of accurate models for this task is challenging due to the scarcity of datasets that possess synchronized vital sign recordings. Our research enhances an existing non-contrastive unsupervised learning technique for extracting rPPG signals, which does not necessitate ground-truth signals during the training process. We have incorporated new time-domain loss functions and added a feature stabilization block to improve the model's stability and accuracy in detecting low-level features. Additionally, we have devised a metric to evaluate the feature instability in the model's final layer. Our experiments on four public datasets demonstrate that our method surpasses the performance of current state-of-the-art methods. These advancements make our approach a significant breakthrough in the development of scalable deep-learning models for camera-based heart-rate estimation. Shahzad Ahmad 0002, Sania Bano, Sukalpa Chanda, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 5 |
| 2025 | PULSE: Physiological Understanding with Liquid Signal ExtractionabstractThe non-contact estimation of vital signs, particularly heart rate, from video data is a promising method for remote health monitoring. 3D convolutional layers are widely used for this task due to their ability to capture both spatial and temporal features. However, traditional 3D convolutions, while effective in many cases, lack the capacity to adjust dy-namically to the temporal variability inherent in physiological signals such as remote photoplethysmography (rPPG), which are characterized by subtle frequency changes over time. To address this, we propose PULSE (Physiological Understanding with Liquid Signal Extraction), a frame-work that employs Liquid Time-Constant (LTC) models with 3D convolutional layers to enhance temporal sensitivity and improve the extraction of these fine-grained rPPG signals. In PULSE, traditional 3D-conv layers are deployed for ini-tial feature extraction, while LTC-based 3D-conv layers dy-namically adapt and guide the temporal processing, allowing the model to better track and interpret the subtle variations in heart rate signals under different conditions, such as motion artifacts and lighting changes. We evaluated the effectiveness of PULSE in an unsupervised training setting, demonstrating that our solution performs well even in the absence of labeled datasets a common challenge in rPPG signal extraction. Experimental evaluations on three public datasets confirm that PULSE achieves comparable or supe-rior results to existing methods, proving its robustness and efficacy for real-world, non-contact health monitoring applications. Shahzad Ahmad 0002, Sania Bano, Sachin Verma, Yogesh S. Rawat, Sukalpa Chanda, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 7 |
| 2025 | Phaseformer: Phase-Based Attention Mechanism for Underwater Image Restoration and BeyondabstractQuality degradation is observed in underwater images due to the effects of light refraction and absorption by water, leading to issues like color cast, haziness, and limited visibility. This degradation negatively affects the performance of autonomous underwater vehicles used in marine applications. To address these challenges, we propose a lightweight phase-based transformer network with 1.77M parameters for underwater image restoration (UIR). Our approachfocuses on effectively extracting non-contaminated features using a phase-based self-attention mechanism. We also introduce an optimized phase attention block to restore structural information by propagating prominent attentive features from the input. We evaluate our method on both synthetic (UIEB, UFO-120) and real-world (UIEB, U45, UCCS, SQUID) underwater image datasets. Additionally, we demonstrate its effectiveness for low-light image enhancement using the LOL dataset. Through extensive ab-lation studies and comparative analysis, it is clear that the proposed approach outperforms existing state-of-the-art (SOTA) methods. Code is available at Phaseformer. Md Raqib Khan, Anshul Negi, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 6 |
| 2025 | USWformer: Efficient Sparse Wavelet Transformer for Underwater Image EnhancementabstractTransformer-based methods have shown great promise in underwater image enhancement (UIE) tasks due to their capability to model long-range dependencies, which are vital for reconstructing clear images. While numerous effective attention mechanisms have been devised to handle the computational requirements of transformers, they frequently incorporate redundant information and noisy interactions from irrelevant regions. Additionally, the current methods focusing solely on the raw pixel space constrains the exploration of the underwater image frequency dynamics, thus hindering the models from fully leveraging their potential for producing high-quality images. To address these challenges, we propose USWformer, an efficient UIE Sparse Wavelet Transformer Network (1.19 M parameters) to eliminate the redundant features in both the spatial and frequency domains. The USWformer consists of two fundamental components: a Sparse Wavelet Self-Attention (SWSA) block and a Multi-scale Wavelet Feed-Forward Network (MWFN). The SWSA block selectively preserves essential attention scores from the keys corresponding to each query, adjusting the feature details. MWFN further diminishes the feature redundancy in the aggregated features thereby improving the enhancement of the underwater images. We assess the efficacy of our approach across benchmark datasets comprising synthetic and real-world under-water images, showcasing its superiority via thorough ablation studies and comparative analyses. Nancy Mehta, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 4 |
| 2025 | Hierarchical motion magnification
Jasdeep Singh, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, G. Sankara Raju Kosuru, Hasan Al-Marzouqi |
Neurocomputing | 3 |
| 2025 | Learnable directional scale space filters for video motion magnification
Jasdeep Singh, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, G. Sankara Raju Kosuru, Hasan Al-Marzouqi |
Knowl. Based Syst. | 3 |
| 2025 | Unpaired recurrent learning for real-world video de-hazingabstractAutomated outdoor vision-based applications have become increasingly in demand for day-to-day life. Bad weather like haze, rain, snow, etc. may limit the reliability of these applications due to degradation in the overall video quality. So, there is a dire need to pre-process the weather-degraded videos before they are fed to downstream applications. Researchers generally adopt synthetically generated paired hazy frames for learning the task of video de-hazing. The models trained solely on synthetic data may have limited performance on different types of real-world hazy scenarios due to significant domain gap between synthetic and real-world hazy videos. One possible solution is to prove the generalization ability by training on unpaired data for video de-hazing. Some unpaired learning approaches are proposed for single image de-hazing. However, these unpaired single image de-hazing approaches compromise the performance in terms of temporal consistency, which is important for video de-hazing tasks. With this motivation, we have proposed a lightweight and temporally consistent architecture for video de-hazing tasks. To achieve this, diverse receptive and multi-scale features at various input resolutions are mixed and aggregated with multi-kernel attention to extract significant haze information. Furthermore, we propose a recurrent multi-attentive feature alignment concept to maintain temporal consistency with recurrent feedback of previously restored frames for temporal consistent video restoration . Comprehensive experiments are conducted on real-world and synthetic video databases (REVIDE and RSA100Haze). Both the qualitative and quantitative results show significant improvement of the proposed network with better temporal consistency over state-of-the-art methods for detailed video restoration in hazy weather. Source code is available at: https://github.com/pwp1208/UnpairedVideoDehazing . Prashant W. Patil, Santosh Nagnath Randive, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh, M. Subrahmanyam 0001 |
Pattern Recognit. | 6 |
| 2025 | KL-DNAS: Knowledge Distillation-Based Latency Aware-Differentiable Architecture Search for Video Motion MagnificationabstractVideo motion magnification is the task of making subtle minute motions visible. Many times subtle motion occurs while being invisible to the naked eye, e.g., slight deformations in muscles of an athlete, small vibrations in the objects, microexpression, and chest movement while breathing. Magnification of such small motions has resulted in various applications like posture deformities detection, microexpression recognition, and studying the structural properties. State-of-the-art (SOTA) methods have fixed computational complexity, which makes them less suitable for applications requiring different time constraints, e.g., real-time respiratory rate measurement and microexpression classification. To solve this problem, we propose a knowledge distillation-based latency aware-differentiable architecture search (KL-DNAS) method for video motion magnification. To reduce memory requirements and to improve denoising characteristics, we use a teacher network to search the network by parts using knowledge distillation (KD). Furthermore, search among different receptive fields and multifeature connections are applied for individual layers. Also, a novel latency loss is proposed to jointly optimize the target latency constraint and output quality. We are able to find smaller model than the SOTA method and better motion magnification with lesser distortions. https://github.com/jasdeep-singh-007/KL-DNAS. Jasdeep Singh, M. Subrahmanyam 0001, G. Sankara Raju Kosuru |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Zero Reference based Low-light Enhancement with Wavelet OptimizationabstractImages captured in low light conditions usually suffer from poor visibility, a high amount of noise, and little information stored in the dark image, which has a negative impact on subsequent processing for outdoor computer vision applications. Presently, numerous deep learning based methods achieved superior performance with multi-exposure paired training data or additional information. However, obtaining multi-exposure data samples is a tedious task in real-time scenarios. To mitigate this challenge, we propose a zero reference based learnable wavelet approach without multi-exposure paired training data requirement for low-light image enhancement. Our proposed approach generates the low light image and learns to project an image into noise free similar looking image, then we enhance the image using retinex theory. Further, we have proposed learnable wavelet block to remove the hidden noise amplified while enhancement. We introduce Gaussian-based supervision to improve the smoothness of the image. Extensive experimental analysis on synthetic as well as real-world images, along with thorough ablation study demonstrate the effectiveness of our proposed method over the existing state-of-the-art methods for low-light image enhancement. The code is provided at https://github.com/vision-lab-sggsiet/Zero-Reference-based-Low-light-Enhancement-with-Wavelet-Optimization. Vivek Deshmukh, Adinath Madhavrao Dukre, Ashutosh Kulkarni, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Anil Balaji Gonde |
AVSS | 6 |
| 2024 | AeroDehazeNet: Exploiting Selective Multi-Scale Transformers for Aerial Image DehazingabstractRemote sensing is the task of analyzing and acquiring useful information from satellite images captured at a far distance from the earth’s surface. These images are vulnerable to degradation due to the presence of mist or haze. Existing methods either make use of prior information to estimate haze free images, or use CNN architectures based on generative adversarial networks (GANs) or Transformers. Though the state-of-the-art transformer-based architectures helped to dehaze the aerial images, they lacked the ability to capture multi-scale dependencies of the image. Identifying this shortcoming, we propose AeroDehazeNet based on a transformer that captures multi-scale dependencies along with global dependencies of the image. Our network comprises of three key components: (1) a multi-scale selective attention (MScA) network to attentively process the multi-scale information in an image, (2) residual attention network (RAN in feed forward network responsible for distilling non-degraded features passed from MScA, and (3) high frequency dominant skip connection (HFDS) block for passing diverse features (low frequency and high frequency) prominent with multi-scale edge features from encoder levels to adjacent decoder levels. The extensive quantitative and qualitative comparisons with existing methods on synthetic and realworld data plus exhaustive ablation study demonstrate the efficacy of our proposed network over transformer based state-of-the-art architectures with comparatively less number of parameters and FLOPs. Testing code is available at https://github.com/KartikGonde/AeroDehazeNet. Kartik Gonde, Prashant W. Patil, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Pramod Patil, Vinod V. Kimbahune |
AVSS | 4 |
| 2024 | RefMOS: A Robust Referred Moving Object Segmentation framework based on text queryabstractReferred Moving object segmentation is a very challenging task in automated video surveillance applications as it requires additional information to learn about object representation referred by natural language expression. In segmenting specific moving objects targeted by a text, suppressing other moving as well as stationary objects is a crucial task. A better context needs to be learned where linguistic, spatial, and temporal features need to be taken into account. In this work, we have proposed a robust referred moving object segmentation (RefMOS) framework to capture moving objects referred by text query. Most of the earlier state-of-the-art methods exploit a different type of supervision by treating video frames as images but lack temporal information during processing. In this work, we have proposed an inter-frame movement detector (IFCD) module, which extracts the movement information between the consecutive frames and helps integrate temporal information with spatial visual features. Language embedding is utilized to capture the information of referred moving objects in the text by extracting linguistic features from a pre-trained language model, i.e., BERT. Furthermore, the cross-entropy loss and SGD optimizer are used to train the network. Our RefMOS framework competes with the state-of-the-art approaches and achieves 48.6 mean IOU on the ref-DAVIS 17 dataset. Prafulla Saxena, Susim Mukul Roy, Dinesh Kumar Tyagi, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Balasubramanian Raman |
AVSS | 5 |
| 2024 | Luminate: Linguistic Understanding and Multi-Granularity Interaction for Video Object SegmentationabstractReferring Video Object Segmentation (R-VOS) is a challenging task that involves segmenting objects in a video based on linguistic descriptions. In this paper, we introduce a novel multi-granularity referring video Object segmentation framework, termed as LUMINATE. The LUMINATE framework introduces a streamlined approach to cross-modal fusion. The proposed LUMINATE enhanced interaction between visual and textual modalities begins with cross-attention between the vision encoder’s query and the text encoder’s key-value pairs, and vice versa. The results are then concatenated with the respective queries of the vision and text encoders, fostering a comprehensive understanding of semantic relationships. The combined features are fed into the Transformer Encoder for further refinement and integration into the segmentation pipeline. Extensive experiments on benchmark datasets, including Ref-DAVIS, demonstrate that our proposed LUMINATE approach achieves better results than state-of-the-art methods in terms of Jaccard and F-measure evaluation metrics. Furthermore, the efficiency of our multi-object R-VOS variant is highlighted, achieving a threefold speed improvement while maintaining satisfactory segmentation performance. The proposed approach contributes to advancing the capabilities of R-VOS models, paving the way for improved multimodal reasoning and real-world applications. Rahul Tekchandani, Ritik Maheshwari, Praful Hambarde, Satya Narayan Tazi, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
ICIP | 6 |
| 2024 | Frequency Modulated Deformable Transformer for Underwater Image Enhancement
Adinath Madhavrao Dukre, Vivek Deshmukh, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, Anil Balaji Gonde, M. Subrahmanyam 0001 |
ICPR (32) | 7 |
| 2024 | Probing Attention-Driven Normalizing Flow Network for Low-Light Image Enhancement
Nancy Mehta, K. N. Prakash, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
ICPR (32) | 5 |
| 2024 | Attentive Color Fusion Transformer Network (ACFTNet) for Underwater Image Enhancement
Mohd Ubaid Wani, Md Raqib Khan, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
ICPR (21) | 6 |
| 2024 | Spectroformer: Multi-Domain Query Cascaded Transformer Network For Underwater Image EnhancementabstractUnderwater images often suffer from color distortion, haze, and limited visibility due to light refraction and absorption in water. These challenges significantly impact autonomous underwater vehicle applications, necessitating efficient image enhancement techniques. To address these challenges, we propose a Multi-Domain Query Cascaded Transformer Network for underwater image enhancement. Our approach includes a novel Multi-Domain Query Cascaded Attention mechanism that integrates localized transmission features and global illumination features. To improve feature propagation from the encoder to the decoder, we propose a Spatio-Spectro Fusion-Based Attention Block. Additionally, we introduce a Hybrid Fourier-Spatial Up-sampling Block, which uniquely combines Fourier and spatial upsampling techniques to enhance feature resolution effectively. We evaluate our method on benchmark synthetic and real-world underwater image datasets, demonstrating its superiority through extensive ablation studies and comparative analysis. The testing code is available at: https://github.com/Mdraqibkhan/Spectroformer. Md Raqib Khan, Nancy Mehta, Shruti S. Phutke, Santosh Kumar Vipparthi, Sukumar Nandi, M. Subrahmanyam 0001 |
WACV | 7 |
| 2024 | C2AIR: Consolidated Compact Aerial Image Haze RemovalabstractAerial image haze removal deals with improving the visibility and quality of images captured from aerial platforms, such as drones and satellites. Aerial images are commonly used in various applications such as environmental monitoring, and disaster response. These applications usually require cleaner data for accurate functioning. However, atmospheric conditions such as haze or fog can significantly degrade the quality of these images, reducing their contrast, color saturation, and sharpness, making it difficult to extract meaningful information from them. Existing methods rely on computationally heavy and haze density (light, moderate, dense) specific architectures for aerial image dehazing. In light of these limitations, we propose a novel lightweight and consolidated approach for aerial image dehazing. In this approach, we propose Density Aware Query Modulated Block for learning weather degradations in input features and guiding the restoration process. Further, we propose Cross Collaborative Feed-Forward Block for learning to restore varying sizes of the structures in the input images. Finally, we propose Gated Adaptive Feature Fusion block to achieve inter-scale and intra-feature attentive fusion, effective for aerial image restoration. Extensive analysis on benchmark aerial image dehazing datasets and real-world images, along with detailed ablation studies validate the effectiveness of the proposed approach. Further, we have analysed our method for other restoration task such as underwater image enhancement to experiment its wide applicability. The code is available at https://github.com/AshutoshKulkarni4998/C2AIR. Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 4 |
| 2024 | Image Inpainting via Correlated Multi-Resolution Feature ProjectionabstractWith the advancement in image editing applications, image inpainting is gaining more attention due to its ability to recover corrupted images efficiently. Also, the existing methods for image inpainting either use two-stage coarse-to-fine architectures or single-stage architectures with a deeper network. On the other hand, shallow network architectures lack the quality of results and the methods with remarkable inpainting quality have high complexity in terms of number of parameters or average run time. Despite the improvement in the inpainting quality, these methods still lack the correlated local and global information. In this work, we propose a single-stage multi-resolution generator architecture for image inpainting with moderate complexity and superior outcomes. Here, a multi-kernel non-local (MKNL) attention block is proposed to merge the feature maps from all the resolutions. Further, a feature projection block is proposed to project features of MKNL to respective decoder for effective reconstruction of image. Also, a valid feature fusion block is proposed to merge encoder skip connection features at valid region and respective decoder features at hole region. This ensures that there will not be any redundant feature merging while reconstruction of image. Effectiveness of the proposed architecture is verified on CelebA-HQ Liu, et al. 2015, Karras et al. 2017, and Places2 Zhou et al. 2018 datasets corrupted with publicly available NVIDIA mask dataset Liu et al. 2018. The detailed ablation study, extensive result analysis, and application of object removal prove the robustness of the proposed method over existing state-of-the-art methods for image inpainting. Shruti S. Phutke, M. Subrahmanyam 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Gated Multi-Resolution Transfer Network for Burst Restoration and EnhancementabstractBurst image processing is becoming increasingly popular in recent years. However, it is a challenging task since individual burst images undergo multiple degradations and often have mutual misalignments resulting in ghosting and zipper artifacts. Existing burst restoration methods usually do not consider the mutual correlation and non-local contextual information among burst frames, which tends to limit these approaches in challenging cases. Another key challenge lies in the robust up-sampling of burst frames. The existing up-sampling methods cannot effectively utilize the advantages of single-stage and progressive up-sampling strategies with conventional and/or recent up-samplers at the same time. To address these challenges, we propose a novel Gated Multi-Resolution Transfer Network (GMTNet) to reconstruct a spatially precise high-quality image from a burst of low-quality raw images. GMT-Net consists of three modules optimized for burst processing tasks: Multi-scale Burst Feature Alignment (MBFA) for feature denoising and alignment, Transposed-Attention Feature Merging (TAFM) for multi-frame feature aggregation, and Resolution Transfer Feature Up-sampler (RTFU) to up-scale merged features and construct a high-quality output image. Detailed experimental analysis on five datasets validate our approach and sets a state-of-the-art for burst super-resolution, burst denoising, and low-light burst enhancement. Our codes and models are available at https://github.com/nanmehta/GMTNet. Nancy Mehta, Akshay Dudhane, M. Subrahmanyam 0001, Syed Waqas Zamir, Salman Khan 0001, Fahad Shahbaz Khan |
CVPR | 3 |
| 2023 | Multi Domain Learning for Motion MagnificationabstractVideo motion magnification makes subtle invisible motions visible, such as small chest movements while breathing, subtle vibrations in the moving objects etc. But small motions are prone to noise, illumination changes, large motions, etc. making the task difficult. Most state-of-the-art methods use hand-crafted concepts which result in small magnification, ringing artifacts etc. The deep learning-based approach has higher magnification but is prone to severe artifacts in some scenarios. We propose a new phase-based deep network for video motion magnification that operates in both domains (frequency and spatial) to address this issue. It generates motion magnification from frequency domain phase fluctuations and then improves its quality in the spatial domain. The proposed models are lightweight networks with fewer parameters (∼0.11M and ∼0. 05M). Further, the proposed networks performance is compared to the SOTA approaches and evaluated on real-world and synthetic videos. Finally, an ablation study is also conducted to show the impact of different parts of the network. Jasdeep Singh, M. Subrahmanyam 0001, G. Sankara Raju Kosuru |
CVPR | 2 |
| 2023 | Multi-weather Image Restoration via Domain TranslationabstractWeather degraded conditions such as rain, haze, snow, etc. may degrade the performance of most computer vision systems. Therefore, effective restoration of multi-weather degraded images is an essential prerequisite for successful functioning of such systems. The current multi-weather image restoration approaches utilize a model that is trained on a combined dataset consisting of individual images for rainy, snowy, and hazy weather degradations. These methods may face challenges when dealing with real-world situations where the images may have multiple, more intricate weather conditions. To address this issue, we propose a domain translation-based unified method for multi-weather image restoration. In this approach, the proposed network learns multiple weather degradations simultaneously, making it immune for real-world conditions. Specifically, we first propose an instance-level domain (weather) translation with multi-attentive feature learning approach to get different weather-degraded variants of the same scenario. Next, the original and translated images are used as input to the proposed novel multi-weather restoration network which utilizes a progressive multi-domain deformable alignment (PMDA) with cascaded multi-head attention (CMA). The proposed PMDA facilitates the restoration network to learn weather-invariant clues effectively. Further, PMDA and respective decoder features are merged via proposed CMA module for restoration. Extensive experimental results on synthetic and real-world hazy, rainy, and snowy image databases clearly demonstrate that our model outperforms the state-of-the-art multi-weather image restoration methods. Code is available at https://github.com/pwp1208/Domain_Translation_Multi-weather_Restoration. Prashant W. Patil, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh, M. Subrahmanyam 0001 |
ICCV | 5 |
| 2023 | Underwater Image Enhancement with Phase Transfer and AttentionabstractUnderwater pictures typically suffer from substantial deterioration due to the refraction and absorption of light by water, including color cast, hazy blur, and limited visibility. Such degradation in visibility eventually reduces the effectiveness of marine applications installed on autonomous underwater vehicles. Hence, an efficient pre-processing step is required for the significant performance of these applications. As a solution, underwater image enhancement (UIE) mainly focuses on enhancing the visibility of degraded images along with restoring crucial details. Existing methods generally utilize (a) complex cascaded architectures, (b) different degradation-prone color spaces, and (c) direct skip connections that pass irrelevant content. In light of this, we propose a lightweight transformer network with 1.7M parameters (1/6thof the existing state-of-the-art method) consisting of the proposed gray-scale attention and phase transformer block for UIE. A gray-scale attention block is proposed for the effective extraction of non-contaminated features. Further, a phase transfer block is proposed for effectively restoring the structural information in the outputs by propagating most relevant and undegraded features from the inputs. A comprehensive evaluation of the proposed method on synthetic (EUVP, UIEB) and real-world (UIEB, UCCS) image datasets as well as extensive ablation studies confirm its effectiveness over existing state-of-the-art approaches. The source code is provided at: https://github.com/Mdraqibkhan/UIEPTA. Md Raqib Khan, Ashutosh Kulkarni, Shruti S. Phutke, M. Subrahmanyam 0001 |
IJCNN | 4 |
| 2023 | Aerial Image Dehazing with Attentive Deformable TransformersabstractAerial imagery is widely utilized in visual data dependent applications such as military surveillance, earthquake assessment, etc. For these applications, minute texture in the aerial image are essential as any disturbance can cause inaccurate prediction. However, atmospheric haze severely reduces the visibility of the scene to be analysed, and hence takes a toll on accuracy of higher level applications. Existing methods either utilize additional prior while training, or produce sub-optimal outputs on different densities of haze degradation, due to absence of local and global dependencies in the extracted features. Therefore, it is essential to have a texture preserving algorithm for aerial image dehazing. In light of this, we propose a work that introduces a novel deformable multi-head attention with spatially attentive offset extraction based solution for aerial image dehazing. Here, the deformable multi-head attention is introduced to reconstruct fine level texture in the restored image. We also introduce spatially attentive offset extractor in the deformable convolution for focusing on relevant contextual information. Further, edge boosting skip connections are proposed for effectively passing edge features from shallow layers to deeper layers of the network. Thorough experimentation on synthetic as well as real-world data, along with extensive ablation study, demonstrate that the proposed method outperforms the prevailing works on aerial image dehazing. The code is provided at https://github.com/AshutoshKulkarni4998/AIDTransformer. Ashutosh Kulkarni, M. Subrahmanyam 0001 |
WACV | 2 |
| 2023 | Nested Deformable Multi-head Attention for Facial Image InpaintingabstractExtracting adequate contextual information is an important aspect of any image inpainting method. To achieve this, ample image inpainting methods are available that aim to focus on large receptive fields. Recent advancements in the deep learning field with the introduction of transformers for image inpainting paved the way toward plausible results. Stacking multiple transformer blocks in a single layer causes the architecture to become computationally complex. In this context, we propose a novel lightweight architecture with a nested deformable attention-based transformer layer for feature fusion. The nested attention helps the network to focus on long-term dependencies from encoder and decoder features. Also, multi-head attention consisting of a deformable convolution is proposed to delve into the diverse receptive fields. With the advantage of nested and deformable attention, we propose a lightweight architecture for facial image inpainting. The results comparison on Celeb HQ [25] dataset using known (NVIDIA) and unknown (QD-IMD) masks and Places2 [57] dataset with NVIDIA masks along with extensive ablation study prove the superiority of the proposed approach for image inpainting tasks. The code is available at: https://github.com/shrutiphutke/NDMA_Facial_Inpainting. Shruti S. Phutke, M. Subrahmanyam 0001 |
WACV | 2 |
| 2023 | Lightweight Network For Video Motion MagnificationabstractVideo motion magnification provides information to understand the subtle changes present in objects for applications like industrial, healthcare, sports, etc. Most state-of- the-art (SOTA) methods use hand-crafted bandpass filters, which require prior information for the motion magnification, produces ringing artifacts, and small magnification etc. While others use deep-learning based techniques for higher magnification, but their output suffers from artificially induced motion, distortions, blurriness, etc. Further, SOTA methods are computationally complex, which makes them less suitable for real-time applications. To address these problems, we proposed deep learning based simple yet effective solution for motion magnification. The proposed method uses a feature sharing and appearance encoder for better motion magnification with fewer distortions, artifacts etc. Additionally, for reducing magnification of noise and other unwanted changes, proxy-model based training is proposed. A computationally lightweight model (~0.12 M parameters) is proposed along with the base model. The performance of the proposed models is tested qualitatively and quantitatively, with the SOTA methods. Results demonstrate the effectiveness of the proposed lightweight and base model over the existing SOTA methods. Jasdeep Singh, M. Subrahmanyam 0001, G. Sankara Raju Kosuru |
WACV | 2 |
| 2023 | Image inpainting via spatial projections
Shruti S. Phutke, M. Subrahmanyam 0001 |
Pattern Recognit. | 2 |
| 2023 | Unified Multi-Weather Visibility RestorationabstractAutomated surveillance is widely opted for appli- cations such as traffic monitoring, vehicle identification, etc. But, various weather degradation factors such as rain and snow streaks, along with atmospheric veil severely affect the perceptual quality of an image, eventually affecting the performance of these applications. There exist weather specific (rain, haze, snow, etc.) methods focusing on respective restoration task. As image restoration is a preprocessing step for high level surveillance applications, it is practically inapplicable to have different architectures for different weather restoration. In this paper, we propose a lightweight unified network, having 1.1 M parameters (1/40th and 1/6th of the existing rain with veil removal, and snow with veil removal methods respectively) for removal of rain and snow along with the veiling effect present in the images. In this network, we propose two parallel streams to handle the degradations and restoration: First, degradation removal stream (DRS) focuses mainly on removing randomly repeating degradations i.e., rain and snow streaks, through the proposed adaptive multi-scale feature sharing block (AMFSB) and stage-wise subtractive block (SSB). Second, feature corrector stream (FCS) mainly focuses on refining the partial outputs of the first stream, reducing the veiling effect and acts supplementary to the first stream. Finally, we leverage contrastive regularization for better convergence of the proposed network. Substantial experiments on synthetic as well as real-world images, along with extensive ablation studies, demonstrate that the proposed method performs competitively with the existing methods for multi-weather image restoration. The code is available at:https://github.com/AshutoshKulkarni4998/UVRNet. Ashutosh Kulkarni, Prashant W. Patil, M. Subrahmanyam 0001, Sunil Gupta 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Deep Network for Extremely Low-Resolution Human Action RecognitionabstractDue to advancement in automated applications, privacy-preserving is an emerging concern. This concern is more significant in the case of human-centred surveillance application like human action recognition (HAR). Along with privacy concern, the computational complexity due to the huge size of video data is another major concern. To overcome these limitations, an attempt is made to examine the domain of human action recognition in low-resolution (LR) videos. The extremely LR video data ensures sufficient distortion in visual information to hide the identity of the person. Therefore, working with LR videos can resolve the above mentioned concerns of privacy preserving and computational complexity up to a certain extent. In this paper, a new generative adversarial network (GAN) based neural architecture is proposed for HAR in extremely low-resolution videos. The extensive results analysis with ablation study on the state-of-the-art datasets proves the effectiveness of the proposed method over the existing methods for LR-HAR. Sachin Chaudhary, Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001 |
AVSS | 4 |
| 2022 | Robust Unseen Video Understanding for Various Surveillance EnvironmentsabstractAutomated video-based applications are a highly demanding technique from a security perspective, where detection of moving objects i.e., moving object segmentation (MOS) is performed. Therefore, we have proposed an effective solution with a spatio-temporal squeeze excitation mechanism (SqEm) based multi-level feature sharing encoder-decoder network for MOS. Here, the SqEm module is proposed to get prominent foreground edge information using spatio-temporal features. Further, a multi-level feature sharing residual decoder module is proposed with respective SqEm features and previous output features for accurate and consistent foreground segmentation. To handle the foreground or background class imbalance issue, we propose a region of interest-based edge loss. The extensive experimental analysis on three databases is conducted. Result analysis and ablation study proved the robustness of the proposed network for unseen video understanding over SOTA methods. Prashant W. Patil, Jasdeep Singh, Praful Hambarde, Ashutosh Kulkarni, Sachin Chaudhary, M. Subrahmanyam 0001 |
AVSS | 6 |
| 2022 | Multi-frame based adversarial learning approach for video surveillance
Prashant W. Patil, Akshay Dudhane, Sachin Chaudhary, M. Subrahmanyam 0001 |
Pattern Recognit. | 4 |
| 2022 | Progressive Subtractive Recurrent Lightweight Network for Video DerainingabstractPresence of rainy artifacts severely degrade the overall visual quality of a video and tend to overlap with the useful information present in the video frames. This degraded video affects the effectiveness of many automated applications like traffic monitoring, surveillance,etc.As video deraining is a pre-processing step for automated applications, it is highly demanded to have a lightweight deraining module. Therefore, in this paper, a“Progressive Subtractive Recurrent Lightweight Network”is proposed for video deraining. Initially, the Multi-Kernel feature Sharing Residual Block (MKSRB) is designed to learn different sizes of rain streaks which facilitates the complete removal of rain streaks through progressive subtractions. These MKSRB features are merged with previous frame output recurrently to maintain the temporal consistency. Further, multi-receptive feature subtraction is performed through Multi-scale Multi-Receptive Difference Block (MMRDB) to avoid loss of details and extract high-frequency information. Finally, progressively learned features through MKSRB and recurrent feature merging are aggregated with fused MMRDB features which outputs the rain-free frame. Substantial experiments on prevailing synthetic datasets and real-world videos verify the superior performance of the proposed method over the existing state-of-the-art methods for video deraining. Ashutosh Kulkarni, Prashant W. Patil, M. Subrahmanyam 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | FASNet: Feature Aggregation and Sharing Network for Image InpaintingabstractImage inpainting is a reconstruction method, where a corrupted image consisting of holes is filled with the most relevant contents from the valid region of an image. To inpaint an image, we have proposed a lightweight cascaded architecture with2.5M parametersconsisting of encoder feature aggregation block (FAB) with decoder feature sharing (DFS) inpainting network followed by a refinement network. Initially, the FAB with DFS (inpainting) generator network is proposed which comprises of multi-level feature aggregation mechanism and feature sharing decoder. The FAB makes use of multi-scale spatial channel-wise attention to fuse weighted features from all the encoder levels. The DFS reconstructs the inpainted image with multi-scale and multi-receptive feature sharing in order to inpaint the image with smaller to larger hole regions effectively. Further, the refinement generator network is proposed for refining the inpainted image from the inpainting generator network. The effectiveness of proposed architecture is verified on CelebA-HQ [1], [2], Paris Street View (PARIS_SV) [3] and Places2 [4] datasets corrupted using publicly available NVIDIA mask dataset [5]. Extensive result analysis with detailed ablation study prove the robustness of the proposed architecture over state-of-the-art methods for image inpainting. Shruti S. Phutke, M. Subrahmanyam 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Pseudo Decoder Guided Light-Weight Architecture for Image InpaintingabstractImage inpainting is one of the most important and widely used approaches where input image is synthesized at the missing regions. This has various applications like undesired object removal, virtual garment shopping, etc. The methods used for image inpainting may use the knowledge of hole locations to effectively regenerate contents in an image. Existing image inpainting methods give astonishing results with coarse-to-fine architectures or with use of guided information like edges, structures, etc. The coarse-to-fine architectures require umpteen resources leading to high computation cost of the architecture. Other methods with edge or structural information depend on the available models to generate guiding information for inpainting. In this context, we have proposed computationally efficient, light-weight network for image inpainting with very less number of parameters (0.97M) and without any guided information. The proposed architecture consists of the multi-encoder level feature fusion module, pseudo decoder and regeneration decoder. The encoder multi level feature fusion module extracts relevant information from each of the encoder levels to merge structural and textural information from various receptive fields. This information is then processed with pseudo decoder followed by space depth correlation module to assist regeneration decoder for inpainting task. The experiments are performed with different types of masks and compared with the state-of-the-art methods on three benchmark datasets i.e., Paris Street View (PARIS_SV), Places2 and CelebA_HQ. Along with this, the proposed network is tested on high resolution images ( 1024×1024 and 2048 ×2048 ) and compared with the existing methods. The extensive comparison with state-of-the-art methods, computational complexity analysis, and ablation study prove the effectiveness of the proposed framework for image inpainting. Shruti S. Phutke, M. Subrahmanyam 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | WiperNet: A Lightweight Multi-Weather Restoration Network for Enhanced SurveillanceabstractAdherent raindrops, rainstreaks and snow severely degrade the perceptual quality of an image, eventually affecting the performance of several computer vision based applications which are applied in outdoor scenarios, e.g., traffic monitoring, autonomous driving, etc. Due to the complex appearance properties, removal of such degradations from an image is a challenging task. Working towards mitigating this problem, in this paper, a lightweight network named as WiperNet is proposed which tackles the problem of raindrops, rain streaks and snow removal present in an image. The WiperNet makes use of the proposed Dual Restoration (DR) mechanism, where the input features are processed twice through the network. In the network, Multi-scale Context Aware Residual Block (MCARB) is proposed for integrating contextual information from various scales. Also, Adaptive Varying Receptive Fusion Block (AVRFB) is proposed for adaptively fusing the information acquired through different dilation rates. Finally, we propose a Feature Refinement Stream which makes use of multiple kernel sizes of convolution filters and spatio-channel attention blocks for focusing on relevant information for effective removal of the degradations while using the coarse outputs of the features from the initial layers of the network. Substantial experiments and ablation study scrutinize that the proposed lightweight WiperNet outperforms the existing state-of-the-art methods for raindrop, rain streak and snow removal. The code is provided athttps://github.com/AshutoshKulkarni4998/WiperNet. Ashutosh Kulkarni, M. Subrahmanyam 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Multi-frame Recurrent Adversarial Network for Moving Object SegmentationabstractMoving object segmentation (MOS) in different practical scenarios like weather degraded, dynamic background, etc. videos is a challenging and high demanding task for various computer vision applications. Existing supervised approaches achieve remarkable performance with complicated training or extensive fine-tuning or inappropriate training-testing data distribution. Also, the generalized effect of existing works with completely unseen data is difficult to identify. In this work, the recurrent feature sharing based generative adversarial network is proposed with unseen video analysis. The proposed network comprises of dilated convolution to extract the spatial features at multiple scales. Along with the temporally sampled multiple frames, previous frame output is considered as input to the network. As the motion is very minute between the two consecutive frames, the previous frame decoder features are shared with encoder features recurrently for current frame foreground segmentation. This recurrent feature sharing of different layers helps the encoder network to learn the hierarchical interactions between the motion and appearance-based features. Also, the learning of the proposed network is concentrated in different ways, like disjoint and global training-testing for MOS. An extensive experimental analysis of the proposed network is carried out on two benchmark video datasets with seen and unseen MOS video. Qualitative and quantitative experimental study shows that the proposed network outperforms the existing methods. Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001 |
WACV | 3 |
| 2021 | Hyperrealistic Image Inpainting with HypergraphsabstractImage inpainting is a non-trivial task in computer vision due to multiple possibilities for filling the missing data, which may be dependent on the global information of the image. Most of the existing approaches use the attention mechanism to learn the global context of the image. This attention mechanism produces semantically plausible but blurry results because of incapability to capture the global context. In this paper, we introduce hypergraph convolution on spatial features to learn the complex relationship among the data. We introduce a trainable mechanism to connect nodes using hyperedges for hypergraph convolution. To the best of our knowledge, hypergraph convolution have never been used on spatial features for any image-to-image tasks in computer vision. Further, we introduce gated convolution in the discriminator to enforce local consistency in the predicted image. The experiments on Places2, CelebA-HQ, Paris Street View, and Facades datasets, show that our approach achieves state-of-the-art results. Gourav Wadhwa, Abhinav Dhall, M. Subrahmanyam 0001, Usman Tariq |
WACV | 3 |
| 2021 | Motion estimation in hazy videos
Sachin Chaudhary, Akshay Dudhane, Prashant W. Patil, M. Subrahmanyam 0001, Sanjay N. Talbar |
Pattern Recognit. Lett. | 4 |
| 2021 | MSAR-Net: Multi-scale attention based light-weight image super-resolution
Nancy Mehta, M. Subrahmanyam 0001 |
Pattern Recognit. Lett. | 2 |
| 2021 | Deep Adversarial Network for Scene Independent Moving Object SegmentationabstractThe current prevailing algorithms highly depend on additional pre-trained modules trained for other applications or complicated training procedures or neglect the inter-frame spatio-temporal structural dependencies. Also, the generalized effect of existing works with completely unseen data is difficult to identify. Specifically, the outdoor videos suffer from adverse atmospheric conditions like poor visibility, inclement weather, etc. In this letter, a novel end-to-end multi-scale temporal edge aggregation (MTPA) network is proposed with adversarial learning for scene dependent and independent object segmentation. The MTPA is proposed to extract the comprehensive spatio-temporal features from the current and reference frame. These MTPA features are used to guide the respective decoder through skip connections. To get authentic and consistent foreground object(s), the respective scale feedback of previous frame output is provided with respective MTPA features at each decoder input. The performance analysis of the proposed method is verified on CDnet-2014 and LASIESTA video datasets. The proposed method outperforms the existing state-of-the-art methods with scene dependent and independent analysis. Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001, Anil Balaji Gonde |
IEEE Signal Process. Lett. | 3 |
| 2021 | Diverse Receptive Field Based Adversarial Concurrent Encoder Network for Image InpaintingabstractImage inpainting is nowadays demanding because of its wide applications such as removing the unwanted objects from the image or recovering the old corrupted photo. Existing approaches achieved superior performance with coarse-to-fine or progressive or recurrent architectures for image inpainting regardless of computational complexity. In these types, the disturbance at the first instance or first iteration may lead to semantically unambiguous results. Also, to inpaint the image with varying hole sizes it is desirable to focus on the diverse receptive fields without deeper network i.e, network with less number of parameters. Therefore, we have proposed a lightweight adversarial concurrent encoder architecture with a diverse receptive field for image inpainting. Here, the concurrent encoder is integrated with diverse receptive fields to benefit with lower computational complexity. The proposed method is compared with state-of-the-art (SOTA) methods on Places2 and Paris Street View dataset in terms of peak signal-to-noise ratio and structural similarity index. Along with the extensive results analysis and ablation study, the proposed method proves the effectiveness in terms of less computational complexity compared to existing SOTA methods. Shruti S. Phutke, M. Subrahmanyam 0001 |
IEEE Signal Process. Lett. | 2 |
| 2021 | An Unified Recurrent Video Object Segmentation Framework for Various Surveillance EnvironmentsabstractMoving object segmentation (MOS) in videos received considerable attention because of its broad security-based applications like robotics, outdoor video surveillance, self-driving cars, etc. The current prevailing algorithms highly depend on additional trained modules for other applications or complicated training procedures or neglect the inter-frame spatio-temporal structural dependencies. To address these issues, a simple, robust, and effective unified recurrent edge aggregation approach is proposed for MOS, in which additional trained modules or fine-tuning on a test video frame(s) are not required. Here, a recurrent edge aggregation module (REAM) is proposed to extract effective foreground relevant features capturing spatio-temporal structural dependencies with encoder and respective decoder features connected recurrently from previous frame. These REAM features are then connected to a decoder through skip connections for comprehensive learning named as temporal information propagation. Further, the motion refinement block with multi-scale dense residual is proposed to combine the features from the optical flow encoder stream and the last REAM module for holistic feature learning. Finally, these holistic features and REAM features are given to the decoder block for segmentation. To guide the decoder block, previous frame output with respective scales is utilized. The different configurations of training-testing techniques are examined to evaluate the performance of the proposed method. Specifically, outdoor videos often suffer from constrained visibility due to different environmental conditions and other small particles in the air that scatter the light in the atmosphere. Thus, comprehensive result analysis is conducted on six benchmark video datasets with different surveillance environments. We demonstrate that the proposed method outperforms the state-of-the-art methods for MOS without any pre-trained module, fine-tuning on the test video frame(s) or complicated training. Prashant W. Patil, Akshay Dudhane, Ashutosh Kulkarni, M. Subrahmanyam 0001, Anil Balaji Gonde, Sunil Gupta 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Varicolored Image De-HazingabstractThe quality of images captured in bad weather is often affected by chromatic casts and low visibility due to the presence of atmospheric particles. Restoration of the color balance is often ignored in most of the existing image de-hazing methods. In this paper, we propose a varicolored end-to-end image de-hazing network which restores the color balance in a given varicolored hazy image and recovers the haze-free image. The proposed network comprises of 1) Haze color correction (HCC) module and 2) Visibility improvement (VI) module. The proposed HCC module provides required attention to each color channel and generates a color balanced hazy image. While the proposed VI module processes the color balanced hazy image through novel inception attention block to recover the haze-free image. We also propose a novel approach to generate a large-scale varicolored synthetic hazy image database. An ablation study has been carried out to demonstrate the effect of different factors on the performance of the proposed network for image de-hazing. Three benchmark synthetic datasets have been used for quantitative analysis of the proposed network. Visual results on a set of real-world hazy images captured in different weather conditions demonstrate the effectiveness of the proposed approach for varicolored image de-hazing. Akshay Dudhane, Kuldeep Marotirao Biradar, Prashant W. Patil, Praful Hambarde, M. Subrahmanyam 0001 |
CVPR | 5 |
| 2020 | An End-to-End Edge Aggregation Network for Moving Object SegmentationabstractMoving object segmentation in videos (MOS) is a highly demanding task for security-based applications like automated outdoor video surveillance. Most of the existing techniques proposed for MOS are highly depend on fine-tuning a model on the first frame(s) of test sequence or complicated training procedure, which leads to limited practical serviceability of the algorithm. In this paper, the inherent correlation learning-based edge extraction mechanism (EEM) and dense residual block (DRB) are proposed for the discriminative foreground representation. The multi-scale EEM module provides the efficient foreground edge related information (with the help of encoder) to the decoder through skip connection at subsequent scale. Further, the response of the optical flow encoder stream and the last EEM module are embedded in the bridge network. The bridge network comprises of multi-scale residual blocks with dense connections to learn the effective and efficient foreground relevant features. Finally, to generate accurate and consistent foreground object maps, a decoder block is proposed with skip connections from respective multi-scale EEM module feature maps and the subsequent down-sampled response of previous frame output. Specifically, the proposed network does not require any pre-trained models or fine-tuning of the parameters with the initial frame(s) of the test video. The performance of the proposed network is evaluated with different configurations like disjoint, cross-data, and global training-testing techniques. The ablation study is conducted to analyse each model of the proposed network. To demonstrate the effectiveness of the proposed framework, a comprehensive analysis on four benchmark video datasets is conducted. Experimental results show that the proposed approach outperforms the state-of-the-art methods for MOS Prashant W. Patil, Kuldeep Marotirao Biradar, Akshay Dudhane, M. Subrahmanyam 0001 |
CVPR | 4 |
| 2020 | Depth Estimation From Single Image And Semantic PriorabstractThe multi-modality sensor fusion technique is an active research area in scene understating. In this work, we explore the RGB image and semantic-map fusion methods for depth estimation. The LiDARs, Kinect, and TOF depth sensors are unable to predict the depth-map at illuminate and monotonous pattern surface. In this paper, we propose a semantic-to-depth generative adversarial network (S2D-GAN) for depth estimation from RGB image and its semantic-map. In the first stage, the proposed S2D-GAN estimates the coarse level depthmap using a semantic-to-coarse-depth generative adversarial network (S2CD-GAN) while the second stage estimates the fine-level depth-map using a cascaded multi-scale spatial pooling network. The experimental analysis of the proposed S2D-GAN performed on NYU-Depth-V2 dataset shows that the proposed S2D-GAN gives outstanding result over existing single image depth estimation and RGB with sparse samples methods. The proposed S2D-GAN also gives efficient results on the real-world indoor and outdoor image depth estimation. Praful Hambarde, Akshay Dudhane, Prashant W. Patil, M. Subrahmanyam 0001, Abhinav Dhall |
ICIP | 4 |
| 2020 | A novel feature descriptor for image retrieval by combining modified color histogram and diagonally symmetric co-occurrence texture pattern
Ayan Kumar Bhunia, Avirup Bhattacharyya, Prithaj Banerjee, Partha Pratim Roy 0001, M. Subrahmanyam 0001 |
Pattern Anal. Appl. | 5 |
| 2020 | Deep Underwater Image Restoration and BeyondabstractUnderwater image restoration is a challenging problem due to the multiple distortions. Degradation in the information is mainly due to the 1) light scattering effect 2) wavelength dependent color attenuation and 3) object blurriness effect. In this letter, we propose a novel end-to-end deep network for underwater image restoration. The proposed network is divided into two parts viz. channel-wise color feature extraction module and dense-residual feature extraction module. A custom loss function is proposed, which preserves the structural details and generates the true edge information in the restored underwater scene. Also, to train the proposed network for underwater image enhancement, a new synthetic underwater image database is proposed. Existing synthetic underwater database images are characterized by light scattering and color attenuation distortions. However, object blurriness effect is ignored. We, on the other hand, introduced the blurring effect along with the light scattering and color attenuation distortions. The proposed network is validated for underwater image restoration task on real-world underwater images. Experimental analysis shows that the proposed network is superior than the existing state-of-the-art approaches for underwater image restoration. Akshay Dudhane, Praful Hambarde, Prashant W. Patil, M. Subrahmanyam 0001 |
IEEE Signal Process. Lett. | 4 |
| 2020 | RYF-Net: Deep Fusion Network for Single Image Haze RemovalabstractHaze removal from a single image is a challenging task. Estimation of accurate scene transmission map (TrMap) is the key to reconstruct the haze-free scene. In this paper, we propose a convolutional neural network based architecture to estimate the TrMap of the hazy scene. The proposed network takes the hazy image as an input and extracts the haze relevant features using proposed RNet and YNet through RGB and YCbCr color spaces respectively and generates two TrMaps. Further, we propose a novel TrMap fusion network (FNet) to integrate two TrMaPs and estimate robust TrMap for the hazy scene. To analyze the robustness of FNet, we tested it on combinations of TrMaps obtained from existing state-of-the-art methods. Performance evaluation of the proposed approach has been carried out using the structural similarity index, mean square error and peak signal to noise ratio. We conduct experiments on five datasets namely: D-HAZY ancuti2016d, Imagenet deng2009imagenet, Indoor SOTS li2017reside, HazeRD zhang2017hazerd and set of real-world hazy images. Performance analysis shows that the proposed approach outperforms the existing state-of-the-art methods for single image dehazing. Further, we extended our work to address high-level vision task such as object detection in hazy scenes. It is observed that there is a significant improvement in accurate object detection in hazy scenes using proposed approach. Akshay Dudhane, M. Subrahmanyam 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | LEARNet: Dynamic Imaging Network for Micro Expression RecognitionabstractUnlike prevalent facial expressions, micro expressions have subtle, involuntary muscle movements which are short-lived in nature. These minute muscle movements reflect true emotions of a person. Due to the short duration and low intensity, these micro-expressions are very difficult to perceive and interpret correctly. In this paper, we propose the dynamic representation of micro-expressions to preserve facial movement information of a video in a single frame. We also propose a Lateral Accretive Hybrid Network (LEARNet) to capture micro-level features of an expression in the facial region. The LEARNet refines the salient expression features in accretive manner by incorporating accretion layers (AL) in the network. The response of the AL holds the hybrid feature maps generated by prior laterally connected convolution layers. Moreover, LEARNet architecture incorporates the cross decoupled relationship between convolution layers which helps in preserving the tiny but influential facial muscle change information. The visual responses of the proposed LEARNet depict the effectiveness of the system by preserving both high- and micro-level edge features of facial expression. The effectiveness of the proposed LEARNet is evaluated on four benchmark datasets: CASME-I, CASME-II, CAS(ME)'2 and SMIC. The experimental results after investigation show a significant improvement of 4.03%, 1.90%, 1.79% and 2.82% as compared with ResNet on CASME-I, CASME-II, CAS(ME)'2 and SMIC datasets respectively. Monu Verma, Santosh Kumar Vipparthi, Girdhari Singh, M. Subrahmanyam 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Pose Guided Dynamic Image Network for Human Action Recognition in Person Centric VideosabstractThe most emerging concerns in computer vision are size of data to process and privacy preserving of the end user. Camera sensors are all around us these days, recording and analysing our day-to-day activities. In this scenario the privacy perseverance becomes a question of concern especially in case of devices working on the basis of human action recognition (HAR). Another important concern in computer vision is the size of data. The surveillance requires continues transfer of huge amount of data through the network. The processing time required to transfer the video to central server and analyses the video directly depends on the resolution of the video. The research in computer vision is exploring the possibility of working on different aspects of videos such as using only pose information or representing whole video using a single frame for the purpose of HAR. Here, an attempt is made to explore the concept of pose estimation and video representation using dynamic image to solve the dual purpose of privacy preserving and decreasing the load on network for transfer of videos over the network for analysis. In this paper, a new Pose Guided Dynamic Image (PDI) network is proposed for HAR which is capable of providing a summarized single frame for the person's activity in any given video. Unlike dynamic image network, this approach considers only the person's motion and discards the background motion. Therefore, PDI provides more specific information required for HAR as compared to the dynamic image. Also, by summarizing the video, the identity of the person remains masked. The proposed method is able to provide better result on both of the benchmark datasets used namely JHMDB and UCF-sports for the experimentation. Sachin Chaudhary, Akshay Dudhane, Prashant W. Patil, M. Subrahmanyam 0001 |
AVSS | 4 |
| 2019 | Image and Video Super Resolution using Recurrent Generative Adversarial NetworkabstractRecently, the convolutional neural network with residual learning models achieves high accuracy for single image super-resolution with different scale factors. With adversarial learning model, effective learning of transformation function for the low-resolution input image to a high-resolution target image can be achieved. In this paper, we propose a method for image and video super-resolution using the recurrent generative adversarial network named SR2GAN. In the proposed model (SR2GAN) we use recursive learning for video super-resolution to overcome the difficulty of learning transformation function for synthesizing realistic high-resolution images. This recursive approach helps to reduce the parameters with increasing depth of the model. An extensive evaluation is performed to examine the effectiveness of the proposed model, which shows that SR2GAN performs better in terms of peak signal to noise ratio (PSNR) and structural self-similarity index (SSIM) as compared to the state-of-the-art methods for super-resolution. For source code and supplementary material visit: https://github.com/OmkarThawakar/SR2GAN/. Omkar Thawakar, Prashant W. Patil, Akshay Dudhane, M. Subrahmanyam 0001, Uday Kulkarni |
AVSS | 4 |
| 2019 | Single Image Depth Estimation Using Deep Adversarial TrainingabstractScene understanding is an active area of research in computer vision that encompasses several different problems. The LiDARs and stereo depth sensor have their own restrictions such as light sensitiveness, power consumption and short-range [1]. In this paper, we propose a two-stream deep adversarial network for single image depth estimation in RGB images. For stream I network, we propose a novel encoder-decoder architecture using residual concepts to extract course-level depth features. Stream II network purely processes the information through the residual architecture for fine-level depth estimation. Also, we designed a feature map sharing architecture to share the learned feature maps of the decoder module of stream I. Sharing feature maps strengthen the residual learning to estimate the scene depth and increase the robustness of the proposed network. A benchmark NYU RGB-D v2 [2] database is used to evaluate the proposed network for single image depth estimation. Both qualitative and quantitative analysis has been carried out to analyze the effectiveness of the proposed network for scene depth prediction. Performance analysis shows that the proposed method outperforms other existing methods for single image depth estimation. Praful Hambarde, Akshay Dudhane, M. Subrahmanyam 0001 |
ICIP | 3 |
| 2019 | Motion Saliency Based Generative Adversarial Network for Underwater Moving Object SegmentationabstractThe underwater moving object segmentation is a challenging task. The problems like absorbing, scattering and attenuation of light rays between the scene and the imaging platform degrades the visibility of image or video frames. Also, the back-scattering of light rays further increases the problem of underwater video analysis, because the light rays interact with underwater particles and scattered back to the sensor. In this paper, a novel Motion Saliency Based Generative Adversarial Network (GAN) for Underwater Moving Object Segmentation (MOS) is proposed. The proposed network comprises of both identity mapping and dense connections for underwater MOS. To the best of our knowledge, this is the first paper with the concept of GAN-based unpaired learning for MOS in underwater videos. Initially, current frame motion saliency is estimated using few initial video frames and current frame. Further, estimated motion saliency is given as input to the proposed network for foreground estimation. To examine the effectiveness of proposed network, the Fish4Knowledge [1] underwater video dataset and challenging video categories of ChangeDetection.net-2014 [2] datasets are considered. The segmentation accuracy of existing state-of-the-art methods are used for comparison with proposed approach in terms of average F-measure. From experimental results, it is evident that the proposed network shows significant improvement as compared to the existing state-of-the-art methods for MOS. Prashant W. Patil, Omkar Thawakar, Akshay Dudhane, M. Subrahmanyam 0001 |
ICIP | 4 |
| 2019 | CDNet: Single Image De-Hazing Using Unpaired Adversarial TrainingabstractOutdoor scene images generally undergo visibility degradation in presence of aerosol particles such as haze, fog and smoke. The reason behind this is, aerosol particles scatter the light rays reflected from the object surface and thus results in attenuation of light intensity. Effect of haze is inversely proportional to the transmission coefficient of the scene point. Thus, estimation of accurate transmission map (TrMap) is a key step to reconstruct the haze-free scene. Previous methods used various assumptions/priors to estimate the scene TrMap. Also, available end-to-end dehazing approaches make use of supervised training to anticipate the TrMap on synthetically generated paired hazy images. Despite the success of previous approaches, they fail in real-world extreme vague conditions due to unavailability of the real-world hazy image pairs for training the network. Thus, in this paper, Cycle-consistent generative adversarial network for single image De-hazing named as CDNet is proposed which is trained in an unpaired manner on real-world hazy image dataset. Generator network of CDNet comprises of encoder-decoder architecture which aims to estimate the object level TrMap followed by optical model to recover the haze-free scene. We conduct experiments on four datasets namely: D-HAZY [1], Imagenet [5], SOTS [20] and real-world images. Structural similarity index, peak signal to noise ratio and CIEDE2000 metric are used to evaluate the performance of the proposed CDNet. Experiments on benchmark datasets show that the proposed CDNet outperforms the existing state-of-the-art methods for single image haze removal. Akshay Dudhane, M. Subrahmanyam 0001 |
WACV | 2 |
| 2019 | FgGAN: A Cascaded Unpaired Learning for Background Estimation and Foreground SegmentationabstractThe moving object segmentation (MOS) in videos with bad weather, irregular motion of objects, camera jitter, shadow and dynamic background scenarios is still an open problem for computer vision applications. To address these issues, in this paper, we propose an approach named as Foreground Generative Adversarial Network (FgGAN) with the recent concepts of generative adversarial network (GAN) and unpaired training for background estimation and foreground segmentation. To the best of our knowledge, this is the first paper with the concept of GAN-based unpaired learning for MOS. Initially, video-wise background is estimated using GAN-based unpaired learning network (network-I). Then, to extract the motion information related to foreground, motion saliency is estimated using estimated background and current video frame. Further, estimated motion saliency is given as input to the GANbased unpaired learning network (network-II) for foreground segmentation. To examine the effectiveness of proposed FgGAN (cascaded networks I and II), the challenging video categories like dynamic background, bad weather, intermittent object motion and shadow are collected from ChangeDetection.net-2014 [26] database. The segmentation accuracy is observed qualitatively and quantitatively in terms of F-measure and percentage of wrong classification (PWC) and compared with the existing state-of-the-art methods. From experimental results, it is evident that the proposed FgGAN shows significant improvement in terms of F-measure and PWC as compared to the existing stateof-the-art methods for MOS. Prashant W. Patil, M. Subrahmanyam 0001 |
WACV | 2 |
| 2019 | Depth-based end-to-end deep network for human action recognitionabstractRecognition of human actions from videos can be improved if depth information is available. Depth information certainly helps in segregating foreground motion from the background. Single image depth estimation (SIDE) is a commonly used method for the analysis of weather degraded images. In this study, the idea of SIDE is extended to human action recognition (HAR) on datasets where depth information is not available. Several depth‐based HAR algorithms are available but all of them are using the depth information given with the dataset. Some other methods are using depth motion map which refers to the depth of motion in a temporal direction. Here, a new depth‐based end‐to‐end deep network is proposed for HAR in which the frame‐wise depth is estimated and this estimated depth is used for processing instead of RGB frame. As colour information is not required for estimating motion, a single channel depth map is used for estimating motion in the video. It makes the system computationally efficient. The proposed method is tested and verified on three benchmark datasets namely JHMDB, HMDB51 and UCF101. The proposed method outperforms the existing state‐of‐the‐art methods for HAR on all the three tested datasets. Sachin Chaudhary, M. Subrahmanyam 0001 |
IET Comput. Vis. | 2 |
| 2019 | ANTIC: antithetic isomeric cluster patterns for medical image retrieval and change detectionabstractIn this study, new feature descriptors are designed for medical image retrieval and change detection applications, respectively. Inspired by isomerism, the authors propose a novel feature descriptor named antithetic isomeric cluster pattern (ANTIC). The ANTIC is defined by the two properties: cluster patterns and antithetic isomerism (ANTI). The cluster pattern corresponds to successive pixel intensity differences at antithetical orientations. Furthermore, the ANTI is characterised by two aspects: first, the clusters are oppositely oriented (antithetical) to each other and second, both adhere to a defined isomeric property. The ANTIC identifies the lines and corner point information in the local neighbourhood across various directions. To attain enhanced robustness, they further proposed multiresolution ANTIC by integrating the multiresolution Gaussian filter. Moreover, to reduce the feature dimensionality, they extended their work to rotation invariant features. The proposed method outperforms other widely used feature descriptors in biomedical and retinopathy image retrieval applications. In addition, they extracted spatiotemporal features by designing intra‐ANTIC and inter‐ANTIC to detect motion changes in video sequences. They validated the effectiveness of these features by conducting experiments on CDNet 2014 dataset. The proposed descriptor achieves better performance in various challenging conditions for change detection as compared to other state‐of‐the‐art techniques. Murari Mandal, Mallika Chaudhary, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Anil Balaji Gonde, Shyam Krishna Nagar |
IET Comput. Vis. | 4 |
| 2019 | Regional adaptive affinitive patterns (RADAP) with logical operators for facial expression recognitionabstractAutomated facial expression recognition plays a significant role in the study of human behaviour analysis. In this study, the authors propose a robust feature descriptor named regional adaptive affinitive patterns (RADAP) for facial expression recognition. The RADAP computes positional adaptive thresholds in the local neighbourhood and encodes multi‐distance magnitude features which are robust to intra‐class variations and irregular illumination variation in an image. Furthermore, they established cross‐distance co‐occurrence relations in RADAP by using logical operators. They proposed XRADAP, ARADAP, and DRADAP using xor, adder and decoder, respectively. The XRADAP engrains the quality of robustness to intra‐class variations in RADAP features using pairwise co‐occurrence. Similarly, ARADAP and DRADAP extract more stable and illumination invariant features and capture the minute expression features which are usually missed by regular descriptors. The performance of the proposed methods is evaluated by conducting experiments on nine benchmark datasets Cohn–Kanade+ (CK+), Japanese female facial expression (JAFFE), Multimedia Understanding Group (MUG), MMI, OULU‐CASIA, Indian spontaneous expression database, DISFA, AFEW and Combined (CK+, JAFFE, MUG, MMI & GEMEP‐FERA) database in both person dependent and person independent setup. The experimental results demonstrate the effectiveness of the proposed method over state‐of‐the‐art approaches. Murari Mandal, Monu Verma, Sonakshi Mathur, Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Kranthi Kumar Deveerasetty |
IET Image Process. | 5 |
| 2019 | Deep network for human action recognition using Weber motion
Sachin Chaudhary, M. Subrahmanyam 0001 |
Neurocomputing | 2 |
| 2019 | Local energy oriented pattern for image indexing and retrieval
Gajanan M. Galshetwar, Laxman M. Waghmare, Anil Balaji Gonde, M. Subrahmanyam 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Cardinal color fusion network for single image haze removal
Akshay Dudhane, M. Subrahmanyam 0001 |
Mach. Vis. Appl. | 2 |
| 2019 | MSFgNet: A Novel Compact End-to-End Deep Network for Moving Object DetectionabstractMoving object detection (MOD) in videos is a challenging task. Estimation of accurate background is the key to extracting the foreground from video frames. In this paper, we have proposed a novel compact end-to-end convolutional neural network architecture, motion saliency foreground network (MSFgNet), to estimate the background and to extract the foreground from video frames. Initially, the long streaming video is divided into a number of small video streams (SVS). The proposed network takes the SVS as an input and estimates the background frame for each SVS. Second, the saliency map is extracted using the current video frame and estimated background. Furthermore, a compact encoder-decoder network is proposed to extract the foreground from the estimated saliency maps. The performance of the proposed MSFgNet is tested on three benchmark datasets (CDnet-2014, LASIESTA, and PTIS) for MOD. The computational complexity (handling of number of parameters and execution time) and the performance of the proposed MSFgNet are compared with the existing state-of-the-art methods for MOD in terms of precision, recall, and F-measure. Performance analysis shows that the proposed network is very compact and outperforms the existing state-of-the-art methods for MOD in videos. Prashant W. Patil, M. Subrahmanyam 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | CANDID: Robust Change Dynamics and Deterministic Update Policy for Dynamic Background SubtractionabstractBackground subtraction in video provides the preliminary information essential for many computer vision applications. In this paper, we propose a sequence of approaches named CANDID to solve the change detection problem in challenging video scenarios. The CANDID adaptively initializes the pixel-level distance threshold and update rate. These parameters are updated by computing the change dynamics at a location. Further, the background model is maintained by formulating a deterministic update policy. The performance of the proposed method is evaluated over various challenging scenarios such as dynamic background and extreme weather conditions. The qualitative and quantitative measures of the proposed method outperform the existing state-of-the-art approaches. Murari Mandal, Prafulla Saxena, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
ICPR | 4 |
| 2018 | TSNet: Deep Network for Human Action Recognition in Hazy VideosabstractThe all-weather intelligent surveillance system is the prime challenge for computer vision researchers. The surveillance is mostly done to analyze the human activity in a particular region. Several extreme weather conditions like rain, snow, haze, fog etc. halts the surveillance process and thus decreases the reliability of the surveillance system. Here, an attempt is made to tackle one of these weather situation i.e. haze in case of surveillance. Haze distorts the quality of images and videos captured by camera. Due to poor quality, it is difficult to analyze the haze degraded video for the human activities using the existing state-of-the-art methods for human action recognition (HAR). Therefore, in this paper, a new two level saliency based end-to-end network (TSNet) for HAR in hazy videos is proposed. De-hazing approaches given in [1]-[9] have certain limitations and therefore we fine-tuned the de-hazing network given in [10] for HAR. The concept of rank pooling given in [11] is further utilized to efficiently represent the temporal saliency of the video. The transmission map information is utilized here to fix the spatial saliency in each frame. As currently, there is no dataset available for HAR in hazy video, here a new dataset of hazy video is generated from two benchmark datasets namely HMDB51 [12] and UCF101 [13] by adding synthetic haze. The existing methods for HAR proposed in [11], [14], [15] are applied and compared with the proposed method on proposed hazy-HMDB51 and hazy-UCF101. The proposed method clearly outperforms the above mentioned methods in terms of average recognition rate (ARR). Sachin Chaudhary, M. Subrahmanyam 0001 |
SMC | 2 |
| 2018 | MsEDNet: Multi-Scale Deep Saliency Learning for Moving Object DetectionabstractMoving object detection (foreground and background) is an important problem in computer vision. Most of the works in this problem are based on background subtraction. However, these approaches are not able to handle scenarios with infrequent motion of object, illumination changes, shadow, camouflage etc. To overcome these, here a two stage robust and compact method for moving object detection (MOD) is proposed. In first stage, to generate the saliency map, background image is estimated using a temporal histogram technique with the help of several input frames. In the second stage, multiscale encoder-decoder network is used to learn multiscale semantic feature of estimated saliency for foreground extraction. The encoder is used to extract multi-scale features from multi-scale saliency map. The decoder part is designed to learn the mapping of low resolution multi-scale features into high resolution output frame. To observe the efficacy of proposed MsEDNet, experiments are conducted on two benchmark datasets (change detection (CDnet-2014) [1] and Wallflower [2]) for MOD. The precision, recall and F-measure are used as performance parameter for comparison with the existing state-of-the-art methods. Experimental results show a significant improvement in detection accuracy and decrement in execution time as compared to the state-of-the-art methods for MOD. Prashant W. Patil, M. Subrahmanyam 0001, Abhinav Dhall, Sachin Chaudhary |
SMC | 2 |
| 2018 | C^2MSNet: A Novel Approach for Single Image Haze RemovalabstractDegradation of image quality due to the presence of haze is a very common phenomenon. Existing DehazeNet [3], MSCNN [11] tackled the drawbacks of hand crafted haze relevant features. However, these methods have the problem of color distortion in gloomy (poor illumination) environment. In this paper, a cardinal (red, green and blue) color fusion network for single image haze removal is proposed. In first stage, network fusses color information present in hazy images and generates multi-channel depth maps. The second stage estimates the scene transmission map from generated dark channels using multi channel multi scale convolutional neural network (McMs-CNN) to recover the original scene. To train the proposed network, we have used two standard datasets namely: ImageNet [5] and D-HAZY [1]. Performance evaluation of the proposed approach has been carried out using structural similarity index (SSIM), mean square error (MSE) and peak signal to noise ratio (PSNR). Performance analysis shows that the proposed approach outperforms the existing state-of-the-art methods for single image dehazing. Akshay Dudhane, M. Subrahmanyam 0001 |
WACV | 2 |
| 2018 | Local Neighborhood Intensity Pattern-A new texture feature descriptor for image retrieval
Prithaj Banerjee, Ayan Kumar Bhunia, Avirup Bhattacharyya, Partha Pratim Roy 0001, M. Subrahmanyam 0001 |
Expert Syst. Appl. | 5 |
| 2016 | Local directional mask maximum edge patterns for image retrieval and face recognitionabstractThis study proposes a new feature descriptor, local directional mask maximum edge pattern, for image retrieval and face recognition applications. Local binary pattern (LBP) and LBP variants collect the relationship between the centre pixel and its surrounding neighbours in an image. Thus, LBP based features are very sensitive to the noise variations in an image. Whereas the proposed method collects the maximum edge patterns (MEP) and maximum edge position patterns (MEPP) from the magnitude directional edges of face/image. These directional edges are computed with the aid of directional masks. Once the directional edges (DE) are computed, the MEP and MEPP are coded based on the magnitude of DE and position of maximum DE. Further, the robustness of the proposed method is increased by integrating it with the multiresolution Gaussian filters. The performance of the proposed method is tested by conducting four experiments onopen access series of imaging studies‐magnetic resonance imaging, Brodatz, MIT VisTex and Extended Yale B databases for biomedical image retrieval, texture retrieval and face recognition applications. The results after being investigated the proposed method shows a significant improvement as compared with LBP and LBP variant features in terms of their evaluation measures on respective databases. Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Anil Balaji Gonde, Q. M. Jonathan Wu |
IET Comput. Vis. | 2 |
| 2015 | Spherical symmetric 3D local ternary patterns for natural, texture and biomedical image indexing and retrieval
M. Subrahmanyam 0001, Q. M. Jonathan Wu |
Neurocomputing | 1 |
| 2015 | Local extrema co-occurrence pattern for color and texture image retrieval
Manisha Verma, Balasubramanian Raman, M. Subrahmanyam 0001 |
Neurocomputing | 3 |
| 2015 | Local Gabor maximum edge position octal patterns for image retrieval
Santosh Kumar Vipparthi, M. Subrahmanyam 0001, Shyam Krishna Nagar, Anil Balaji Gonde |
Neurocomputing | 2 |
| 2014 | Expert content-based image retrieval system using robust local patterns
M. Subrahmanyam 0001, Q. M. Jonathan Wu |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | MRI and CT image indexing and retrieval using local mesh peak valley edge patterns
M. Subrahmanyam 0001, Q. M. Jonathan Wu |
Signal Process. Image Commun. | 1 |
| 2014 | Local Mesh Patterns Versus Local Binary Patterns: Biomedical Image Indexing and RetrievalabstractIn this paper, a new image indexing and retrieval algorithm using local mesh patterns are proposed for biomedical image retrieval application. The standard local binary pattern encodes the relationship between the referenced pixel and its surrounding neighbors, whereas the proposed method encodes the relationship among the surrounding neighbors for a given referenced pixel in an image. The possible relationships among the surrounding neighbors are depending on the number of neighbors, P. In addition, the effectiveness of our algorithm is confirmed by combining it with the Gabor transform. To prove the effectiveness of our algorithm, three experiments have been carried out on three different biomedical image databases. Out of which two are meant for computer tomography (CT) and one for magnetic resonance (MR) image retrieval. It is further mentioned that the database considered for three experiments are OASIS-MRI database, NEMA-CT database, and VIA/I-ELCAP database which includes region of interest CT images. The results after being investigated show a significant improvement in terms of their evaluation measures as compared to LBP, LBP with Gabor transform, and other spatial and transform domain methods. M. Subrahmanyam 0001, Q. M. Jonathan Wu |
IEEE J. Biomed. Health Informatics | 1 |
| 2013 | Local ternary co-occurrence patterns: A new feature descriptor for MRI and CT image retrieval
M. Subrahmanyam 0001, Q. M. Jonathan Wu |
Neurocomputing | 1 |
| 2012 | Expert system design using wavelet and color vocabulary trees for image retrieval
M. Subrahmanyam 0001, R. P. Maheshwari 0001, Balasubramanian Raman |
Expert Syst. Appl. | 1 |
| 2012 | Local maximum edge binary patterns: A new descriptor for image retrieval and object tracking
M. Subrahmanyam 0001, R. P. Maheshwari 0001, Balasubramanian Raman |
Signal Process. | 1 |
| 2012 | Local Tetra Patterns: A New Feature Descriptor for Content-Based Image RetrievalabstractIn this paper, we propose a novel image indexing and retrieval algorithm using local tetra patterns (LTrPs) for content-based image retrieval (CBIR). The standard local binary pattern (LBP) and local ternary pattern (LTP) encode the relationship between the referenced pixel and its surrounding neighbors by computing gray-level difference. The proposed method encodes the relationship between the referenced pixel and its neighbors, based on the directions that are calculated using the first-order derivatives in vertical and horizontal directions. In addition, we propose a generic strategy to compute nth-order LTrP using (n - 1)th-order horizontal and vertical derivatives for efficient CBIR and analyze the effectiveness of our proposed algorithm by combining it with the Gabor transform. The performance of the proposed method is compared with the LBP, the local derivative patterns, and the LTP based on the results obtained using benchmark image databases viz., Corel 1000 database (DB1), Brodatz texture database (DB2), and MIT VisTex database (DB3). Performance analysis shows that the proposed method improves the retrieval result from 70.34%/44.9% to 75.9%/48.7% in terms of average precision/average recall on database DB1, and from 79.97% to 85.30% and 82.23% to 90.02% in terms of average retrieval rate on databases DB2 and DB3, respectively, as compared with the standard LBP. M. Subrahmanyam 0001, R. P. Maheshwari 0001, Balasubramanian Raman |
IEEE Trans. Image Process. | 1 |