VLDB 2026 Research / reviewers in the wild / expert
Shruti S. Phutke
dblp:303/0114
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-7627-5930ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TransWaveNet: Multi-scale Transformer-Wavelet Fusion for Colorectal Polyp Segmentation
Amit Shakya, Akanksha Yadav, Shruti S. Phutke, Lalit Sharma |
ICPR (6) | 3 |
| 2025 | Phaseformer: Phase-Based Attention Mechanism for Underwater Image Restoration and BeyondabstractQuality degradation is observed in underwater images due to the effects of light refraction and absorption by water, leading to issues like color cast, haziness, and limited visibility. This degradation negatively affects the performance of autonomous underwater vehicles used in marine applications. To address these challenges, we propose a lightweight phase-based transformer network with 1.77M parameters for underwater image restoration (UIR). Our approachfocuses on effectively extracting non-contaminated features using a phase-based self-attention mechanism. We also introduce an optimized phase attention block to restore structural information by propagating prominent attentive features from the input. We evaluate our method on both synthetic (UIEB, UFO-120) and real-world (UIEB, U45, UCCS, SQUID) underwater image datasets. Additionally, we demonstrate its effectiveness for low-light image enhancement using the LOL dataset. Through extensive ab-lation studies and comparative analysis, it is clear that the proposed approach outperforms existing state-of-the-art (SOTA) methods. Code is available at Phaseformer. Md Raqib Khan, Anshul Negi, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 4 |
| 2024 | Frequency Modulated Deformable Transformer for Underwater Image Enhancement
Adinath Madhavrao Dukre, Vivek Deshmukh, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, Anil Balaji Gonde, M. Subrahmanyam 0001 |
ICPR (32) | 4 |
| 2024 | Attentive Color Fusion Transformer Network (ACFTNet) for Underwater Image Enhancement
Mohd Ubaid Wani, Md Raqib Khan, Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
ICPR (21) | 4 |
| 2024 | Spectroformer: Multi-Domain Query Cascaded Transformer Network For Underwater Image EnhancementabstractUnderwater images often suffer from color distortion, haze, and limited visibility due to light refraction and absorption in water. These challenges significantly impact autonomous underwater vehicle applications, necessitating efficient image enhancement techniques. To address these challenges, we propose a Multi-Domain Query Cascaded Transformer Network for underwater image enhancement. Our approach includes a novel Multi-Domain Query Cascaded Attention mechanism that integrates localized transmission features and global illumination features. To improve feature propagation from the encoder to the decoder, we propose a Spatio-Spectro Fusion-Based Attention Block. Additionally, we introduce a Hybrid Fourier-Spatial Up-sampling Block, which uniquely combines Fourier and spatial upsampling techniques to enhance feature resolution effectively. We evaluate our method on benchmark synthetic and real-world underwater image datasets, demonstrating its superiority through extensive ablation studies and comparative analysis. The testing code is available at: https://github.com/Mdraqibkhan/Spectroformer. Md Raqib Khan, Nancy Mehta, Shruti S. Phutke, Santosh Kumar Vipparthi, Sukumar Nandi, M. Subrahmanyam 0001 |
WACV | 4 |
| 2024 | C2AIR: Consolidated Compact Aerial Image Haze RemovalabstractAerial image haze removal deals with improving the visibility and quality of images captured from aerial platforms, such as drones and satellites. Aerial images are commonly used in various applications such as environmental monitoring, and disaster response. These applications usually require cleaner data for accurate functioning. However, atmospheric conditions such as haze or fog can significantly degrade the quality of these images, reducing their contrast, color saturation, and sharpness, making it difficult to extract meaningful information from them. Existing methods rely on computationally heavy and haze density (light, moderate, dense) specific architectures for aerial image dehazing. In light of these limitations, we propose a novel lightweight and consolidated approach for aerial image dehazing. In this approach, we propose Density Aware Query Modulated Block for learning weather degradations in input features and guiding the restoration process. Further, we propose Cross Collaborative Feed-Forward Block for learning to restore varying sizes of the structures in the input images. Finally, we propose Gated Adaptive Feature Fusion block to achieve inter-scale and intra-feature attentive fusion, effective for aerial image restoration. Extensive analysis on benchmark aerial image dehazing datasets and real-world images, along with detailed ablation studies validate the effectiveness of the proposed approach. Further, we have analysed our method for other restoration task such as underwater image enhancement to experiment its wide applicability. The code is available at https://github.com/AshutoshKulkarni4998/C2AIR. Ashutosh Kulkarni, Shruti S. Phutke, Santosh Kumar Vipparthi, M. Subrahmanyam 0001 |
WACV | 2 |
| 2024 | Image Inpainting via Correlated Multi-Resolution Feature ProjectionabstractWith the advancement in image editing applications, image inpainting is gaining more attention due to its ability to recover corrupted images efficiently. Also, the existing methods for image inpainting either use two-stage coarse-to-fine architectures or single-stage architectures with a deeper network. On the other hand, shallow network architectures lack the quality of results and the methods with remarkable inpainting quality have high complexity in terms of number of parameters or average run time. Despite the improvement in the inpainting quality, these methods still lack the correlated local and global information. In this work, we propose a single-stage multi-resolution generator architecture for image inpainting with moderate complexity and superior outcomes. Here, a multi-kernel non-local (MKNL) attention block is proposed to merge the feature maps from all the resolutions. Further, a feature projection block is proposed to project features of MKNL to respective decoder for effective reconstruction of image. Also, a valid feature fusion block is proposed to merge encoder skip connection features at valid region and respective decoder features at hole region. This ensures that there will not be any redundant feature merging while reconstruction of image. Effectiveness of the proposed architecture is verified on CelebA-HQ Liu, et al. 2015, Karras et al. 2017, and Places2 Zhou et al. 2018 datasets corrupted with publicly available NVIDIA mask dataset Liu et al. 2018. The detailed ablation study, extensive result analysis, and application of object removal prove the robustness of the proposed method over existing state-of-the-art methods for image inpainting. Shruti S. Phutke, M. Subrahmanyam 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Underwater Image Enhancement with Phase Transfer and AttentionabstractUnderwater pictures typically suffer from substantial deterioration due to the refraction and absorption of light by water, including color cast, hazy blur, and limited visibility. Such degradation in visibility eventually reduces the effectiveness of marine applications installed on autonomous underwater vehicles. Hence, an efficient pre-processing step is required for the significant performance of these applications. As a solution, underwater image enhancement (UIE) mainly focuses on enhancing the visibility of degraded images along with restoring crucial details. Existing methods generally utilize (a) complex cascaded architectures, (b) different degradation-prone color spaces, and (c) direct skip connections that pass irrelevant content. In light of this, we propose a lightweight transformer network with 1.7M parameters (1/6thof the existing state-of-the-art method) consisting of the proposed gray-scale attention and phase transformer block for UIE. A gray-scale attention block is proposed for the effective extraction of non-contaminated features. Further, a phase transfer block is proposed for effectively restoring the structural information in the outputs by propagating most relevant and undegraded features from the inputs. A comprehensive evaluation of the proposed method on synthetic (EUVP, UIEB) and real-world (UIEB, UCCS) image datasets as well as extensive ablation studies confirm its effectiveness over existing state-of-the-art approaches. The source code is provided at: https://github.com/Mdraqibkhan/UIEPTA. Md Raqib Khan, Ashutosh Kulkarni, Shruti S. Phutke, M. Subrahmanyam 0001 |
IJCNN | 3 |
| 2023 | Nested Deformable Multi-head Attention for Facial Image InpaintingabstractExtracting adequate contextual information is an important aspect of any image inpainting method. To achieve this, ample image inpainting methods are available that aim to focus on large receptive fields. Recent advancements in the deep learning field with the introduction of transformers for image inpainting paved the way toward plausible results. Stacking multiple transformer blocks in a single layer causes the architecture to become computationally complex. In this context, we propose a novel lightweight architecture with a nested deformable attention-based transformer layer for feature fusion. The nested attention helps the network to focus on long-term dependencies from encoder and decoder features. Also, multi-head attention consisting of a deformable convolution is proposed to delve into the diverse receptive fields. With the advantage of nested and deformable attention, we propose a lightweight architecture for facial image inpainting. The results comparison on Celeb HQ [25] dataset using known (NVIDIA) and unknown (QD-IMD) masks and Places2 [57] dataset with NVIDIA masks along with extensive ablation study prove the superiority of the proposed approach for image inpainting tasks. The code is available at: https://github.com/shrutiphutke/NDMA_Facial_Inpainting. Shruti S. Phutke, M. Subrahmanyam 0001 |
WACV | 1 |
| 2023 | Image inpainting via spatial projections
Shruti S. Phutke, M. Subrahmanyam 0001 |
Pattern Recognit. | 1 |
| 2022 | FASNet: Feature Aggregation and Sharing Network for Image InpaintingabstractImage inpainting is a reconstruction method, where a corrupted image consisting of holes is filled with the most relevant contents from the valid region of an image. To inpaint an image, we have proposed a lightweight cascaded architecture with2.5M parametersconsisting of encoder feature aggregation block (FAB) with decoder feature sharing (DFS) inpainting network followed by a refinement network. Initially, the FAB with DFS (inpainting) generator network is proposed which comprises of multi-level feature aggregation mechanism and feature sharing decoder. The FAB makes use of multi-scale spatial channel-wise attention to fuse weighted features from all the encoder levels. The DFS reconstructs the inpainted image with multi-scale and multi-receptive feature sharing in order to inpaint the image with smaller to larger hole regions effectively. Further, the refinement generator network is proposed for refining the inpainted image from the inpainting generator network. The effectiveness of proposed architecture is verified on CelebA-HQ [1], [2], Paris Street View (PARIS_SV) [3] and Places2 [4] datasets corrupted using publicly available NVIDIA mask dataset [5]. Extensive result analysis with detailed ablation study prove the robustness of the proposed architecture over state-of-the-art methods for image inpainting. Shruti S. Phutke, M. Subrahmanyam 0001 |
IEEE Signal Process. Lett. | 1 |
| 2022 | Pseudo Decoder Guided Light-Weight Architecture for Image InpaintingabstractImage inpainting is one of the most important and widely used approaches where input image is synthesized at the missing regions. This has various applications like undesired object removal, virtual garment shopping, etc. The methods used for image inpainting may use the knowledge of hole locations to effectively regenerate contents in an image. Existing image inpainting methods give astonishing results with coarse-to-fine architectures or with use of guided information like edges, structures, etc. The coarse-to-fine architectures require umpteen resources leading to high computation cost of the architecture. Other methods with edge or structural information depend on the available models to generate guiding information for inpainting. In this context, we have proposed computationally efficient, light-weight network for image inpainting with very less number of parameters (0.97M) and without any guided information. The proposed architecture consists of the multi-encoder level feature fusion module, pseudo decoder and regeneration decoder. The encoder multi level feature fusion module extracts relevant information from each of the encoder levels to merge structural and textural information from various receptive fields. This information is then processed with pseudo decoder followed by space depth correlation module to assist regeneration decoder for inpainting task. The experiments are performed with different types of masks and compared with the state-of-the-art methods on three benchmark datasets i.e., Paris Street View (PARIS_SV), Places2 and CelebA_HQ. Along with this, the proposed network is tested on high resolution images ( 1024×1024 and 2048 ×2048 ) and compared with the existing methods. The extensive comparison with state-of-the-art methods, computational complexity analysis, and ablation study prove the effectiveness of the proposed framework for image inpainting. Shruti S. Phutke, M. Subrahmanyam 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Diverse Receptive Field Based Adversarial Concurrent Encoder Network for Image InpaintingabstractImage inpainting is nowadays demanding because of its wide applications such as removing the unwanted objects from the image or recovering the old corrupted photo. Existing approaches achieved superior performance with coarse-to-fine or progressive or recurrent architectures for image inpainting regardless of computational complexity. In these types, the disturbance at the first instance or first iteration may lead to semantically unambiguous results. Also, to inpaint the image with varying hole sizes it is desirable to focus on the diverse receptive fields without deeper network i.e, network with less number of parameters. Therefore, we have proposed a lightweight adversarial concurrent encoder architecture with a diverse receptive field for image inpainting. Here, the concurrent encoder is integrated with diverse receptive fields to benefit with lower computational complexity. The proposed method is compared with state-of-the-art (SOTA) methods on Places2 and Paris Street View dataset in terms of peak signal-to-noise ratio and structural similarity index. Along with the extensive results analysis and ablation study, the proposed method proves the effectiveness in terms of less computational complexity compared to existing SOTA methods. Shruti S. Phutke, M. Subrahmanyam 0001 |
IEEE Signal Process. Lett. | 1 |