EDBT 2026 Demo / reviewers in the wild / expert
Alik Pramanick
dblp:367/5486
· DBLP profile ↗
10ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0002-5993-3185ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D2Mamba: Dual Domain Guided Informed Search in State Space Model for Underwater Image EnhancementabstractUnderwater images suffer from color distortion, haziness, and low-contrast due to light absorption and scattering. Despite deep learning advances in enhancement, challenges persist in efficiency, global context modeling, spatial-spectral consistency, and perceptually accurate detail recovery. To address these challenges, we design a novel underwater image enhancement framework, D2Mamba, adopting a dual-domain information (spatial and frequency) with state space models (SSMs), enabling efficient global context modeling while preserving local details. Unlike conventional SSMs that rely on raster, bidirectional, cross or diagonal scans, D2Mamba uses an A* search guided by physics-based Geodesic Information-Field Heuristic (GIFH) scan for feature traversal based on input degradation characteristics. GIFH combines feature gradients, high-frequency heterogeneity, and low-frequency semantic distance to compute adaptive costs, enabling the capture of both spatial and spectral dependencies. Further, a Spectral Wasserstein Attenuation Loss (SWAL) is introduced to enforce distributional alignment in the spectral domain, enabling perceptually consistent and physically consistent color restoration in enhanced underwater images. Extensive experiments on benchmark datasets demonstrate that D2Mamba achieves state-of-the-art performance with only 788K parameters and 7.06 GFLOPs. The code is available at https://github.com/Alik033/D2Mamba. Alik Pramanick, Soumajit Roy, Arijit Sur |
WACV | 1 |
| 2025 | MedCAM-OsteoCls: Medical Context Aware Multimodal Classification of Knee OsteoarthritisabstractKnee Osteoarthritis (KOA) is a degenerative musculoskeletal joint disorder that significantly impacts middle-aged and elderly individuals. Although X-rays and MRIs are clinically used to identify such disorders, combining these imaging modalities is challenging due to the distinct nature of the data, as MRI volume provides detailed views of cartilage and soft tissues, while X-rays offer a global perspective of bone anatomy and positioning in single scan. In this context, the proposed work aims to address two key challenges: (1) how to automatically select the most informative slices from large and dis-organized MRI volume in a clinically relevant way and (2) how to effectively capture the discriminative multimodal interactions between X-ray and MRI modalities. In order to solve these issues, the Context-Guided Slice Selection and Prioritization (CG-SSP) module and Xray-MRI Cross-Attention (XMRCA) module are proposed. The CG-SSP module focuses on rejecting non-efficient slices and prioritizing the efficient slices in a multi-stage approach, first by segmenting the Femoral Cartilage (FC) then extracting the sequential patches of FC located in each MRI slice followed by assigning the attention weight to each of these slices based on learned structural complexity, while the XMRCA module focuses on learning the global interactions between X-ray and MRI feature space in a computationally efficient manner. Overall, the proposed model achieved an accuracy improvement of 2.66% over the state-of-the-art in unimodal and multimodal baselines on the OAI dataset. The proposed model identifies key diagnostic slices from MRI and can provide valuable insights into interactions between X-ray and MRI data for radiological KOA diagnosis in an end-to-end manner. Code will be made available at: https://github.com/adaydar/MedCAM-OsteoCls Akshay Daydar, Alik Pramanick, Arijit Sur, Subramani Kanagaraj |
ICASSP | 2 |
| 2025 | Efficient-USR: Prompt Guided Dual-Domain Feature Information for Efficient Underwater Image Super-ResolutionabstractRecent advances in deep learning have significantly improved underwater image super-resolution (UISR) performance. However, their large and complex architectures result in huge computational complexity, making them unsuitable for low-power devices such as autonomous underwater vehicles (AUVs) and remotely operated vehicles (ROVs). In addition, current research emphasizes designing deep models to enhance performance but overlooks the potential benefits of integrating frequency domain information. To overcome these challenges, we propose Efficient-USR, an effective dual-domain information-based lightweight framework for UISR. In Efficient-USR, we analyze the image in different scales by incorporating two core components: a) spatial-frequency interaction block (SFIB) to capture both the global and local context by operating spatial-channel cross-attention on spatial-frequency features and b) prompt-guided cross-attention block (PCAB) to efficiently encode degradation information in prompt component and perform feature fusion of multiple scales via cross-attention. Extensive experiments on two UISR benchmarks demonstrate the effectiveness of our Efficient-USR (e.g., ∼ 1.35% and ∼ 1.40% SSIM improvement on UFO-120 and USR-248 4 datasets, respectively, with only 0.18 M parameters). The code× is available at: https: //github.com/Alik033/Efficient-USR. Alik Pramanick, Utsav Bheda, Arijit Sur |
ICASSP | 1 |
| 2025 | River-GEM: Generating and Enhancing Muddy Water ImagesabstractUnderwater image enhancement is crucial for marine engineering and aquatic robotics. However, most recent methods have focused on ocean environments, where they trained and tested on oceanic images. As a result, these methods are less effective in river water, where relatively blurry images are produced due to extensive muddy environments. In river water-based image restoration, we have observed two main limitations: (1) the lack of datasets that accurately represent the muddy water conditions found in river environments and (2) the limited effectiveness of current enhancement models in dealing with the unique challenges of muddy water images. To address these limitations, we introduce the "River-GEM" framework, which includes (1) a dataset of highly degraded muddy water images generated using a principle component analysis-based fusion technique and (2) a novel enhancement model utilizing the depth-map guidance to learn complex contextual details of muddy water images. The proposed model captures the intricate details of the objects in degraded images by analyzing feature correlation at multiple scales. Our model shows improvements of 3.42% in SSIM and 5.04% in PSNR over existing methods on the muddy-UIEB dataset. This dataset and enhancement network marks a significant step forward in this area of research. The dataset and code will be available at: https://github.com/Alik033/River-GEM. Alik Pramanick, Chivukula Sairam Satwik, Sonal Kumar, Akshay Daydar, Arijit Sur |
ICASSP | 1 |
| 2025 | Diffusion Based Shape-Aware Learning with Multi-Scale Context for Segmentation of Tibiofemoral Knee Joint Tissues: an End-to-End ApproachabstractKnee Osteoarthritis (KOA) is commonly evaluated through an automated segmentation of the femur (FB), tibia (TB), and tibiofemoral cartilage (FC and TC) in Magnetic Resonance Imaging. But, in recent works, such automated segmentation (1) is possible only from the multistage framework, thus creating data handling challenges and requiring continuous manual intervention, and (2) lacks poor learning capability in multiclass segmentation problem. To solve these issues, Multi-Scale Attentive-Unet (MiSA-Unet) model is proposed that includes (1) Scale-aware Attentive Feature Enhancement (SAFE) module to learn the multi-contextual information by attending to varied receptive fields and (2) Diffusion-based Multiple Tissue Shape Reconstruction (DMulTiSR) loss to progressively refine the tissue details by contemplating the feature difference and utilizing discrete diffusion process. Unlike existing methods, the proposed framework is end-to-end and single-stage model that improved average DSC by 2.33% and segmentation speed by 82.6%, while with post-processing, it out-performed by 9.19% in FC, 33.82% in TC, and 13.31% in FB for Hausdorff distance. The code is available at: https://github.com/adaydar/MiSA-Unet. Akshay Daydar, Alik Pramanick, Arijit Sur, Subramani Kanagaraj |
ICIP | 2 |
| 2025 | TURBIT: Generating Turbid Underwater Images with Diffusion and Differential TransformersabstractUnderwater vision in muddy and turbid conditions poses significant challenges for underwater exploration and aquatic research. Existing benchmark datasets focus primarily on oceanic water scenes, limiting their applicability to highly turbid environments. This paper proposes a two-stage framework that combines a diffusion transformer (M1) and a differential attention model (M2) to generate synthetic turbid underwater datasets. M1, trained on the ImageNet Underwater Subset, acts as a foundational content generator, while M2, trained using a few-shot approach on a curated dataset, models turbidity-specific context. A conditional fusion module integrates both models, ensuring realism while preserving structural integrity. The plug-and-play design enables domain-specific fine-tuning of M2, enhancing adaptability across diverse conditions. We benchmarked the generated turbid underwater datasets against state-of-the-art underwater image enhancement (UIE) methods, with extensive experiments on fusion strategies and depth preservation validating their effectiveness. This framework establishes a scalable data generation pipeline that advances UIE research in challenging underwater scenarios. Codes and datasets are available at https://github.com/Alik033/TURBIT. Utkarsh Srivastava, Soumajit Roy, Alik Pramanick, Arijit Sur |
ICIP | 3 |
| 2025 | Harnessing multi-resolution and multi-scale attention for underwater image restoration
Alik Pramanick, Arijit Sur, V. Vijaya Saradhi |
Vis. Comput. | 1 |
| 2024 | Attention-Based Spatial-Frequency Information Network for Underwater Single Image Super-ResolutionabstractUnderwater single image super-resolution (UISR) is a challenging task as these images frequently suffer from poor visibility. The best-published UISR works continue to suffer from color degradation, poor texture representation, and loss of finer (high-frequency) details. We propose a novel deep learning-based (DL) UISR model that incorporates spatial information as well as the transformed (wavelet) coefficient of degraded low-resolution (LR) underwater images by intelligent feature management. To ensure the visual quality of the super-resolved image, color channel-specific L1 loss, perceptual loss, and difference of Gaussian (DoG) loss are used in tandem with SSIM loss. We employ publicly available datasets, namely UFO-120 and USR-248, to evaluate the proposed model. The results of our experiments show that our model outperforms existing state-of-the-art methods (e.g., $\sim 9.45\% $/$\sim 1.77\% $ in SSIM and $\sim 0.91\% $/ $\sim 1.44\% $ in PSNR on UFO-120/USR-248 ×4, respectively), as demonstrated through quantitative measurements and visual quality assessments. Alik Pramanick, Dhruvil Megha, Arijit Sur |
ICASSP | 1 |
| 2024 | X-CAUNET: Cross-Color Channel Attention with Underwater Image-Enhancing TransformerabstractUnderwater image enhancement is essential to mitigate the environment-centric noise in images, such as haziness, color degradation, etc. With most existing works focused on processing an RGB image as a whole, the explicit context that can be mined from each color channel separately goes unaccounted for, ignoring the effects produced by the wavelength of light in underwater conditions. In this work, we propose a framework called X-CAUNET that addresses this research gap by using cross-attention transformers. The input image is split into three channels (R-G-B), local context is captured using convolutional layers with different receptive field sizes, and a message-passing mechanism allows for context correlation between them. To maintain consistency, another transformer is used on the original image to aggregate global context, and a weighted combination of all the outputs enhances the input degraded image. Extensive experiments demonstrate we achieve state-of-the-art PSNR and SSIM with 2.66% and 2.11% relative gains. Code is available at: https://github.com/Alik033/X-CAUNET. Alik Pramanick, Sandipan Sarma, Arijit Sur |
ICASSP | 1 |
| 2024 | ML-CrAIST: Multi-scale Low-High Frequency Information-Based Cross Attention with Image Super-Resolving Transformer
Alik Pramanick, Utsav Bheda, Arijit Sur |
ICPR (21) | 1 |