Arijit Sur

dblp:07/5501 · also Jain Arijit Sur · DBLP profile ↗
← Back
71ranked-venue papers
5as first author
38since 2021 · last 2026
0000-0002-9038-8138ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 58 · 4 first-author · 31 since 2021Artificial intelligence and machine learning · 13 · 13 since 2021Computer networks · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 MMG-SLAM: Multimodal Visual SLAM with MambaVision Loops and Gaussian Splatting
Ashok Bandyopadhyay, Adarsh Gupta, Arijit Sur, U. P. Rajeev
ICPR (11)3
2026 DGSSM: Diffusion Guided State-Space Models for Multimodal Salient Object Detection
Suklav Ghosh, Arijit Sur, Pinaki Mitra
ICPR (14)2
2026 A Fully Unsupervised Framework for Object Mask Labeling with the Self-supervised Vision Transformer
Sonal Kumar, Kathrotiya Sanket Jitendrabhai, Akshay Daydar, Arijit Sur, Rashmi Dutta Baruah
ICPR (12)4
2026 Universal Adversarial Suffixes Using Calibrated Gumbel-Softmax Relaxation
Sampriti Soor, Suklav Ghosh, Arijit Sur
ICPR (15)3
2026 D2Mamba: Dual Domain Guided Informed Search in State Space Model for Underwater Image Enhancement
abstract
Underwater images suffer from color distortion, haziness, and low-contrast due to light absorption and scattering. Despite deep learning advances in enhancement, challenges persist in efficiency, global context modeling, spatial-spectral consistency, and perceptually accurate detail recovery. To address these challenges, we design a novel underwater image enhancement framework, D2Mamba, adopting a dual-domain information (spatial and frequency) with state space models (SSMs), enabling efficient global context modeling while preserving local details. Unlike conventional SSMs that rely on raster, bidirectional, cross or diagonal scans, D2Mamba uses an A* search guided by physics-based Geodesic Information-Field Heuristic (GIFH) scan for feature traversal based on input degradation characteristics. GIFH combines feature gradients, high-frequency heterogeneity, and low-frequency semantic distance to compute adaptive costs, enabling the capture of both spatial and spectral dependencies. Further, a Spectral Wasserstein Attenuation Loss (SWAL) is introduced to enforce distributional alignment in the spectral domain, enabling perceptually consistent and physically consistent color restoration in enhanced underwater images. Extensive experiments on benchmark datasets demonstrate that D2Mamba achieves state-of-the-art performance with only 788K parameters and 7.06 GFLOPs. The code is available at https://github.com/Alik033/D2Mamba.
Alik Pramanick, Soumajit Roy, Arijit Sur
WACV3
2026 Funnel-HOI: top-down perception for zero-shot HOI detection
Sandipan Sarma, Agney Talwarr, Arijit Sur
Mach. Vis. Appl.3
2025 MedCAM-OsteoCls: Medical Context Aware Multimodal Classification of Knee Osteoarthritis
abstract
Knee Osteoarthritis (KOA) is a degenerative musculoskeletal joint disorder that significantly impacts middle-aged and elderly individuals. Although X-rays and MRIs are clinically used to identify such disorders, combining these imaging modalities is challenging due to the distinct nature of the data, as MRI volume provides detailed views of cartilage and soft tissues, while X-rays offer a global perspective of bone anatomy and positioning in single scan. In this context, the proposed work aims to address two key challenges: (1) how to automatically select the most informative slices from large and dis-organized MRI volume in a clinically relevant way and (2) how to effectively capture the discriminative multimodal interactions between X-ray and MRI modalities. In order to solve these issues, the Context-Guided Slice Selection and Prioritization (CG-SSP) module and Xray-MRI Cross-Attention (XMRCA) module are proposed. The CG-SSP module focuses on rejecting non-efficient slices and prioritizing the efficient slices in a multi-stage approach, first by segmenting the Femoral Cartilage (FC) then extracting the sequential patches of FC located in each MRI slice followed by assigning the attention weight to each of these slices based on learned structural complexity, while the XMRCA module focuses on learning the global interactions between X-ray and MRI feature space in a computationally efficient manner. Overall, the proposed model achieved an accuracy improvement of 2.66% over the state-of-the-art in unimodal and multimodal baselines on the OAI dataset. The proposed model identifies key diagnostic slices from MRI and can provide valuable insights into interactions between X-ray and MRI data for radiological KOA diagnosis in an end-to-end manner. Code will be made available at: https://github.com/adaydar/MedCAM-OsteoCls
Akshay Daydar, Alik Pramanick, Arijit Sur, Subramani Kanagaraj
ICASSP3
2025 Efficient-USR: Prompt Guided Dual-Domain Feature Information for Efficient Underwater Image Super-Resolution
abstract
Recent advances in deep learning have significantly improved underwater image super-resolution (UISR) performance. However, their large and complex architectures result in huge computational complexity, making them unsuitable for low-power devices such as autonomous underwater vehicles (AUVs) and remotely operated vehicles (ROVs). In addition, current research emphasizes designing deep models to enhance performance but overlooks the potential benefits of integrating frequency domain information. To overcome these challenges, we propose Efficient-USR, an effective dual-domain information-based lightweight framework for UISR. In Efficient-USR, we analyze the image in different scales by incorporating two core components: a) spatial-frequency interaction block (SFIB) to capture both the global and local context by operating spatial-channel cross-attention on spatial-frequency features and b) prompt-guided cross-attention block (PCAB) to efficiently encode degradation information in prompt component and perform feature fusion of multiple scales via cross-attention. Extensive experiments on two UISR benchmarks demonstrate the effectiveness of our Efficient-USR (e.g., ∼ 1.35% and ∼ 1.40% SSIM improvement on UFO-120 and USR-248 4 datasets, respectively, with only 0.18 M parameters). The code× is available at: https: //github.com/Alik033/Efficient-USR.
Alik Pramanick, Utsav Bheda, Arijit Sur
ICASSP3
2025 River-GEM: Generating and Enhancing Muddy Water Images
abstract
Underwater image enhancement is crucial for marine engineering and aquatic robotics. However, most recent methods have focused on ocean environments, where they trained and tested on oceanic images. As a result, these methods are less effective in river water, where relatively blurry images are produced due to extensive muddy environments. In river water-based image restoration, we have observed two main limitations: (1) the lack of datasets that accurately represent the muddy water conditions found in river environments and (2) the limited effectiveness of current enhancement models in dealing with the unique challenges of muddy water images. To address these limitations, we introduce the "River-GEM" framework, which includes (1) a dataset of highly degraded muddy water images generated using a principle component analysis-based fusion technique and (2) a novel enhancement model utilizing the depth-map guidance to learn complex contextual details of muddy water images. The proposed model captures the intricate details of the objects in degraded images by analyzing feature correlation at multiple scales. Our model shows improvements of 3.42% in SSIM and 5.04% in PSNR over existing methods on the muddy-UIEB dataset. This dataset and enhancement network marks a significant step forward in this area of research. The dataset and code will be available at: https://github.com/Alik033/River-GEM.
Alik Pramanick, Chivukula Sairam Satwik, Sonal Kumar, Akshay Daydar, Arijit Sur
ICASSP5
2025 Diffusion Based Shape-Aware Learning with Multi-Scale Context for Segmentation of Tibiofemoral Knee Joint Tissues: an End-to-End Approach
abstract
Knee Osteoarthritis (KOA) is commonly evaluated through an automated segmentation of the femur (FB), tibia (TB), and tibiofemoral cartilage (FC and TC) in Magnetic Resonance Imaging. But, in recent works, such automated segmentation (1) is possible only from the multistage framework, thus creating data handling challenges and requiring continuous manual intervention, and (2) lacks poor learning capability in multiclass segmentation problem. To solve these issues, Multi-Scale Attentive-Unet (MiSA-Unet) model is proposed that includes (1) Scale-aware Attentive Feature Enhancement (SAFE) module to learn the multi-contextual information by attending to varied receptive fields and (2) Diffusion-based Multiple Tissue Shape Reconstruction (DMulTiSR) loss to progressively refine the tissue details by contemplating the feature difference and utilizing discrete diffusion process. Unlike existing methods, the proposed framework is end-to-end and single-stage model that improved average DSC by 2.33% and segmentation speed by 82.6%, while with post-processing, it out-performed by 9.19% in FC, 33.82% in TC, and 13.31% in FB for Hausdorff distance. The code is available at: https://github.com/adaydar/MiSA-Unet.
Akshay Daydar, Alik Pramanick, Arijit Sur, Subramani Kanagaraj
ICIP3
2025 TURBIT: Generating Turbid Underwater Images with Diffusion and Differential Transformers
abstract
Underwater vision in muddy and turbid conditions poses significant challenges for underwater exploration and aquatic research. Existing benchmark datasets focus primarily on oceanic water scenes, limiting their applicability to highly turbid environments. This paper proposes a two-stage framework that combines a diffusion transformer (M1) and a differential attention model (M2) to generate synthetic turbid underwater datasets. M1, trained on the ImageNet Underwater Subset, acts as a foundational content generator, while M2, trained using a few-shot approach on a curated dataset, models turbidity-specific context. A conditional fusion module integrates both models, ensuring realism while preserving structural integrity. The plug-and-play design enables domain-specific fine-tuning of M2, enhancing adaptability across diverse conditions. We benchmarked the generated turbid underwater datasets against state-of-the-art underwater image enhancement (UIE) methods, with extensive experiments on fusion strategies and depth preservation validating their effectiveness. This framework establishes a scalable data generation pipeline that advances UIE research in challenging underwater scenarios. Codes and datasets are available at https://github.com/Alik033/TURBIT.
Utkarsh Srivastava, Soumajit Roy, Alik Pramanick, Arijit Sur
ICIP4
2025 MDAAF: masked domain adversarial adaptation framework for unsupervised domain adaptive semantic segmentation
Avinash Chouhan, Arijit Sur, Dibyajyoti Chutia, Shiv Prasad Aggarwal
Pattern Anal. Appl.2
2025 Harnessing multi-resolution and multi-scale attention for underwater image restoration
Alik Pramanick, Arijit Sur, V. Vijaya Saradhi
Vis. Comput.2
2024 PV-SLAM: Panoptic Visual SLAM with Loop Closure and Online Bundle Adjustment
Ashok Bandyopadhyay, Pranjal Baranwal, Arijit Sur, U. P. Rajeev
BMVC3
2024 IPCL: Iterative Pseudo-Supervised Contrastive Learning to Improve Self-Supervised Feature Representation
abstract
Self-supervised learning with a contrastive batch approach has become a powerful tool for representation learning in computer vision. The performance of downstream tasks is proportional to the quality of visual features learned while self-supervised pre-training. The existing contrastive batch approaches heavily depend on data augmentation to learn latent information from unlabelled datasets. We argue that introducing the dataset’s intra-class variation in a contrastive batch approach improves visual representation quality further. In this paper, we propose a novel self-supervised learning approach named Iterative Pseudo-supervised Contrastive Learning (IPCL), which utilizes a balanced combination of image augmentations and pseudo-class information to improve the visual representation iteratively. Experimental results illustrate that our proposed method surpasses the baseline self-supervised method with the batch contrastive approach. It improves the visual representation quality over multiple datasets, leading to better performance on the downstream unsupervised image classification task. Code is available at https://github.com/SonalKumar95/IPCL.
Sonal Kumar, Anirudh Phukan, Arijit Sur
ICASSP3
2024 Attention-Based Spatial-Frequency Information Network for Underwater Single Image Super-Resolution
abstract
Underwater single image super-resolution (UISR) is a challenging task as these images frequently suffer from poor visibility. The best-published UISR works continue to suffer from color degradation, poor texture representation, and loss of finer (high-frequency) details. We propose a novel deep learning-based (DL) UISR model that incorporates spatial information as well as the transformed (wavelet) coefficient of degraded low-resolution (LR) underwater images by intelligent feature management. To ensure the visual quality of the super-resolved image, color channel-specific L1 loss, perceptual loss, and difference of Gaussian (DoG) loss are used in tandem with SSIM loss. We employ publicly available datasets, namely UFO-120 and USR-248, to evaluate the proposed model. The results of our experiments show that our model outperforms existing state-of-the-art methods (e.g., $\sim 9.45\% $/$\sim 1.77\% $ in SSIM and $\sim 0.91\% $/ $\sim 1.44\% $ in PSNR on UFO-120/USR-248 ×4, respectively), as demonstrated through quantitative measurements and visual quality assessments.
Alik Pramanick, Dhruvil Megha, Arijit Sur
ICASSP3
2024 X-CAUNET: Cross-Color Channel Attention with Underwater Image-Enhancing Transformer
abstract
Underwater image enhancement is essential to mitigate the environment-centric noise in images, such as haziness, color degradation, etc. With most existing works focused on processing an RGB image as a whole, the explicit context that can be mined from each color channel separately goes unaccounted for, ignoring the effects produced by the wavelength of light in underwater conditions. In this work, we propose a framework called X-CAUNET that addresses this research gap by using cross-attention transformers. The input image is split into three channels (R-G-B), local context is captured using convolutional layers with different receptive field sizes, and a message-passing mechanism allows for context correlation between them. To maintain consistency, another transformer is used on the original image to aggregate global context, and a weighted combination of all the outputs enhances the input degraded image. Extensive experiments demonstrate we achieve state-of-the-art PSNR and SSIM with 2.66% and 2.11% relative gains. Code is available at: https://github.com/Alik033/X-CAUNET.
Alik Pramanick, Sandipan Sarma, Arijit Sur
ICASSP3
2024 Boosting Zero-Shot Human-Object Interaction Detection with Vision-Language Transfer
abstract
Human-Object Interaction (HOI) detection is a crucial task that involves localizing interactive human-object pairs and identifying the actions being performed. Most existing HOI detectors are supervised in nature and lack the ability of zero-shot discovery of unseen interactions. Recently, transformer-based methods have superseded the traditional CNN detectors by aggregating image-wide context but still suffer from the long-tail distribution problem in HOI. In this work, our primary focus is improving HOI detection in images, particularly in zero-shot scenarios. We use an end-to-end transformer-based object detector to localize human-object pairs and yield visual features of actions and objects. Moreover, we adopt the text encoder from a popular visual-language model called CLIP with a novel prompting mechanism to extract semantic information for unseen actions and objects. Finally, we learn a strong visual-semantic alignment and achieve state-of-the-art performance on the challenging HICO-DET dataset across five zero-shot settings, with up to 70.88% relative gains. Code is available at https://github.com/sandipan211/ZSHOI-VLT.
Sandipan Sarma, Pradnesh Kalkar, Arijit Sur
ICASSP3
2024 RADA: Reconstruction Assisted Domain Adaptation for Nighttime Aerial Tracking
Avinash Chouhan, Mayank Chandak, Arijit Sur, Dibyajyoti Chutia, Shiv Prasad Aggarwal
ICPR (10)3
2024 GraPix: Exploring Graph Modularity Optimization for Unsupervised Pixel Clustering
Sonal Kumar, Arijit Sur, Rashmi Dutta Baruah
ICPR (10)2
2024 ML-CrAIST: Multi-scale Low-High Frequency Information-Based Cross Attention with Image Super-Resolving Transformer
Alik Pramanick, Utsav Bheda, Arijit Sur
ICPR (21)3
2024 Zero-Shot Underwater Gesture Recognition
Sandipan Sarma, Gundameedi Sai Ram Mohan, Hariansh Sehgal, Arijit Sur
ICPR (7)4
2024 TMLNet: Triad Multitask Learning Network for multiobjective based change detection
Avinash Chouhan, Arijit Sur, Dibyajyoti Chutia, Shiv Prasad Aggarwal
Neurocomputing2
2024 Knee osteoarthritis severity prediction using an attentive multi-scale deep convolutional neural network
Rohit Kumar Jain, Prasen Kumar Sharma, Sibaji Gaj, Arijit Sur, Palash Ghosh
Multim. Tools Appl.4
2024 Multi-contextual design of convolutional neural network for steganalysis
Brijesh Singh, Arijit Sur, Pinaki Mitra
Multim. Tools Appl.2
2024 Collaborative Video Caching in the Edge Network using Deep Reinforcement Learning
abstract
With the enormous growth in mobile data traffic over the 5G environment, Adaptive BitRate (ABR) video streaming has become a challenging problem. Recent advances in Mobile Edge Computing (MEC) technology make it feasible to use Base Stations (BSs) intelligently by network caching, popularity-based video streaming, and more. Additional computing resources on the edge node offer an opportunity to reduce network traffic on the backhaul links during peak traffic hours. More recently, it has been found in the literature that collaborative caching strategies between neighbouring BSs (i.e., MEC servers) make it more efficient to reduce backhaul traffic and network congestion and thus improve the viewer experience substantially. In this work, we propose a Reinforcement Learning (RL)–based collaborative caching mechanism in which the edge servers cooperate to serve the requested content from the end-users. Specifically, this research aims to improve the overall cache hit rate at the MEC, where the edge servers are clustered based on their geographic locations. This task is modelled as a multi-objective optimization problem and solved using an RL framework. In addition, a novel cache admission and eviction policy is defined by calculating the priority score of video segments in the clustered MEC mesh network.
Anirban Lekharu, Arijit Sur, Moumita Patra
ACM Trans. Internet Things3
2024 Reinforcement Learning-Based Adaptive Bitrate Caching at MEC Server
abstract
Mobile Edge Computing (MEC) has become an important concept in modern video communication and broadcasting scenarios to address varied user expectations in an ever-evolving network environment. It has been observed that by preventing redundant access to the Origin Server through backhaul links, caching popular video content in MEC servers minimizes network congestion. Designing an efficient caching mechanism in the MEC server is challenging. To maintain a decent QoE (Quality of Experience) for end-users, we must consider diverse parameters like content popularity, network conditions, etc. In this work, we have proposed a QoE-aware Adaptive BitRate (ABR) caching mechanism at the MEC server using Reinforcement Learning (RL). The proposed model predicts the content popularity of the video and the most preferred video quality for the end-users of a Base Station. In this work, an efficient caching mechanism is devised in the MEC server to provide a decent QoE among the end-users. The primary goal of our RL-based framework is to increase the cache hit rate and reduce the backhaul load while maintaining a satisfactory QoE. Experimental results demonstrate that the proposed model, which emphasizes on video quality, video quality switching, and cache hit rate, outperforms state-of-the-art caching algorithms in terms of the overall QoE reward.
Anirban Lekharu, Annanya Pratap Singh Chauhan, Arijit Sur, Moumita Patra
IEEE Trans. Netw. Serv. Manag.3
2024 DiRaC-I: Identifying Diverse and Rare Training Classes for Zero-Shot Learning
abstract
Zero-Shot Learning (ZSL) is an extreme form of transfer learning that aims at learning from a few “seen classes” to have an understanding about the “unseen classes” in the wild. Given a dataset in ZSL research, most existing works use a predetermined, disjoint set of seen-unseen classes to evaluate their methods. These seen (training) classes might be sub-optimal for ZSL methods to appreciate the diversity and rarity of an object domain. Inspired by strategies like active learning, it is intuitive that intelligently selecting the training classes can improve ZSL performance. In this work, we propose a framework called Diverse and Rare Class Identifier (DiRaC-I) which, given an attribute-based dataset, can intelligently yield the most suitable “seen classes” for training ZSL models. DiRaC-I has two main goals – constructing a diversified set of seed classes, and using them to initialize a visual-semantic mining algorithm for acquiring the classes capturing both diversity and rarity in the object domain adequately. These classes can then be used as “seen classes” to train ZSL models for image classification. We simulate a real-world scenario where visual samples of novel object classes in the wild are available to neither DiRaC-I nor the ZSL models during training and conducted extensive experiments on two benchmark data sets for zero-shot image classification — CUB and SUN. Our results demonstrate DiRaC-I helps ZSL models to achieve significant classification accuracy improvements – specifically, up to 8% for CUB and up to 5% for SUN dataset. Additionally, while recognizing classes exhibiting rare attributes we also observe a performance boost for ZSL models, which is up to 10% and 7% for CUB and SUN datasets, respectively.
Sandipan Sarma, Arijit Sur
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Wavelength-based Attributed Deep Neural Network for Underwater Image Restoration
abstract
Background: Underwater images, in general, suffer from low contrast and high color distortions due to the non-uniform attenuation of the light as it propagates through the water. In addition, the degree of attenuation varies with the wavelength, resulting in the asymmetric traversing of colors. Despite the prolific works for underwater image restoration (UIR) using deep learning, the above asymmetricity has not been addressed in the respective network engineering. Contributions: As the first novelty, this article shows that attributing the right receptive field size ( context ) based on the traversing range of the color channel may lead to a substantial performance gain for the task of UIR. Further, it is important to suppress the irrelevant multi-contextual features and increase the representational power of the model. Therefore, as a second novelty, we have incorporated an attentive skip mechanism to adaptively refine the learned multi-contextual features. The proposed framework, called Deep WaveNet , is optimized using the traditional pixel-wise and feature-based cost functions. An extensive set of experiments have been carried out to show the efficacy of the proposed scheme over existing best-published literature on benchmark datasets. More importantly, we have demonstrated a comprehensive validation of enhanced images across various high-level vision tasks, e.g., underwater image semantic segmentation and diver’s 2D pose estimation. A sample video to exhibit our real-world performance is available at https://tinyurl.com/yzcrup9n . Also, we have open-sourced our framework at https://github.com/pksvision/Deep-WaveNet-Underwater-Image-Restoration .
Prasen Kumar Sharma, Ira Bisht, Arijit Sur
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Resolving Semantic Confusions for Improved Zero-Shot Detection
Sandipan Sarma, Arijit Sur
BMVC3
2022 Aggregated Context Network For Semantic Segmentation Of Aerial Images
abstract
With the considerable advancement of remote sensing technology and computer vision, automatic scene understanding for very high-resolution aerial (VHR) imagery became a necessary research topic. Semantic segmentation of VHR imagery is an important task where context information plays a crucial role. Adequate feature delineation is difficult due to high-class imbalance in remotely sensed data. In this work, we proposed a variant of encoder-decoder-based architecture where residual attentive skip connections are incorporated. We added a multi-context block in each of the encoder units to capture multi-scale and multi-context features and used dense connections for effective feature extraction. A comprehensive set of experiments reveal that the proposed scheme outperformed recently published work by 3% in overall accuracy and F1 score for ISPRS Vaihingen and ISPRS Potsdam benchmark datasets.
Avinash Chouhan, Arijit Sur, Dibyajyoti Chutia
ICIP2
2022 StegGAN: hiding image within image using conditional generative adversarial networks
Brijesh Singh, Prasen Kumar Sharma, Shashank Anil Huddedar, Arijit Sur, Pinaki Mitra
Multim. Tools Appl.4
2022 Deep Learning Model for Content Aware Caching at MEC Servers
abstract
In recent years, mobile data traffic has increased enormously with an increase of mobile and smart devices. The global mobile data traffic is set to increase manifold in the coming years. With the rise in mobile data traffic and heterogeneous mobile devices, substantial improvement has been achieved in wireless media technology in providing a varied range of multimedia services. These multimedia services are often resource-hungry and require high-speed data and low latency transmissions. High-speed networks like the fifth-generation (5G) network helps in faster data delivery resulting in less congestion at the backhaul links and higher transmission capacity. Integrating Mobile Edge Computing (MEC) capabilities into the cellular architecture provides advantages like intelligent and efficient context-aware caching and video adaptations for content delivery. The primary objective of this work is to reduce the overall backhaul congestion and access delay by increasing the cache hit rate at the MEC server. This work proposes a deep learning-based model for caching at the MEC servers based on the content popularity at different time slots of a day. Experimental results reveal that the proposed model outperforms the state-of-the-art standard caching approaches, improves the cache hit probability by almost 21%, and decreases backhaul usage and access delay by approximately 18%.
Anirban Lekharu, Mitansh Jain, Arijit Sur, Arnab Sarkar 0001
IEEE Trans. Netw. Serv. Manag.3
2021 High-resolution image de-raining using conditional GAN with sub-pixel upscaling
Prasen Kumar Sharma, Sathisha Basavaraju, Arijit Sur
Multim. Tools Appl.3
2021 Steganalysis using learned denoising kernels
Brijesh Singh, Mohit Chhajed, Arijit Sur, Pinaki Mitra
Multim. Tools Appl.3
2021 Steganalysis of Digital Images Using Deep Fractal Network
abstract
In the recent literature on steganalysis, it has been observed that a deeper network is, in general, preferred for detecting low tone embedding noise, e.g., SRNet. However, very recently, a deep model, called FractalNet, became popular, which is based on self-similarity and grows deeper and wider by maintaining a balance between depth and width using a recurrent adaptation of a fundamental building block. In this work, the concept of the FractalNet model has been exploited for steganalytic detection, where the embedded image has been used as input. In a practical scenario, it has been observed that steganalytic detection for test images is increasing if the width of the network can be increased with a certain proportion to the depth. The proposed deep network is designed by repeating a basic fractal block in such a way that a balance between the depth and width of the overall network can be maintained. A comprehensive set of experiments reveals that the proposed model outperformed the state-of-the-art results. An ablation study is also included to justify the proposed architecture in favor of its performance.
Brijesh Singh, Arijit Sur, Pinaki Mitra
IEEE Trans. Comput. Soc. Syst.2
2021 High-quality Frame Recurrent Video De-raining with Multi-contextual Adversarial Network
abstract
In this article, we address the problem of rain-streak removal in the videos. Unlike the image, challenges in video restoration comprise temporal consistency besides spatial enhancement. The researchers across the world have proposed several effective methods for estimating the de-noised videos with outstanding temporal consistency. However, such methods also amplify the computational cost due to their larger size. By way of analysis, incorporating separate modules for spatial and temporal enhancement may require more computational resources. It motivates us to propose a unified architecture that directly estimates the de-rained frame with maximal visual quality and minimal computational cost. To this end, we present a deep learning-based Frame-recurrent Multi-contextual Adversarial Network for rain-streak removal in videos. The proposed model is built upon a Conditional Generative Adversarial Network (CGAN)-based framework where the generator model directly estimates the de-rained frame from the previously estimated one with the help of its multi-contextual adversary. To optimize the proposed model, we have incorporated the Perceptual loss function in addition to the conventional Euclidean distance. Also, instead of traditional entropy loss from the adversary, we propose to use the Euclidean distance between the features of de-rained and clean frames, extracted from the discriminator model as a cost function for video de-raining. Various experimental observations across 11 test sets, with over 10 state-of-the-art methods, using 14 image-quality metrics, prove the efficacy of the proposed work, both visually and computationally.
Prasen Kumar Sharma, Sujoy Ghosh, Arijit Sur
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Deep learning-based image de-raining using discrete Fourier transformation
Prasen Kumar Sharma, Sathisha Basavaraju, Arijit Sur
Vis. Comput.3
2020 Scale-aware Conditional Generative Adversarial Network for Image Dehazing
abstract
Outdoor images are often deteriorated due to the presence of haze in the atmosphere. Conventionally, the single image dehazing problem aims to restore the haze-free image. Previous successful approaches have utilized various hand-crafted features/priors. However, such images suffer from color degradation and halo artifacts. By way of analysis, these artifacts, in general, prevail around the regions with high-intensity variation, such as edgy structures. This finding inspires us to consider the Laplacians of Gaussian (LoG) of the images which exceptionally retains this information, to solve the problem of single image haze removal. In this line of thought, we present an end-to-end model that learns to remove the haze based on the per-pixel difference between LoGs of the dehazed and original haze-free images. The optimization of the proposed network is further enhanced by using the adversarial training and perceptual loss function. The proposed method has been appraised on Synthetic Objective Testing Set (SOTS) and benchmark real-world hazy images using 16 image quality measures. Based on the Color Difference (CIEDE 2000), an improvement of ~ 15.89% has been observed over the state-of-the-art method, Yang et al. [50]. An ablation study has been presented at the end to illustrate the improvements achieved by various modules of the proposed network.
Prasen Kumar, Sharma Priyankar, Arijit Sur
WACV3
2020 Prediction mode based H.265/HEVC video watermarking resisting re-compression attack
Sibaji Gaj, Arijit Sur, Prabin Kumar Bora
Multim. Tools Appl.2
2020 Motion vector based video steganography using homogeneous block selection
Shuvendu Rana, Rohit Kamra, Arijit Sur
Multim. Tools Appl.3
2020 Steganalysis for clustering modification directions steganography
Ritvik Rawat, Brijesh Singh, Arijit Sur, Pinaki Mitra
Multim. Tools Appl.3
2020 Image Memorability Prediction Using Depth and Motion Cues
abstract
Memorability is one of the intrinsic image properties, which enables one to quantify to what extent images are memorable to the human cognitive system. Many works have shed light on various visual factors that influence the image memorability. Recently, one such study showed the influence of image depth and motion information on image memorability [3]. Based on these findings, this article proposes a deep learning-based prediction model, which utilizes depth and motion cues to predict the image memorability scores. The proposed model contains three deep CNN networks. Each of these three networks is individually trained to utilize one of the three visual factors: 1) visual depth information; 2) optical flow information; and 3) fine-tuned scene- and object-related features. In the end, all three networks are ensembled to predict the final memorability scores for the given image. An extensive set of experiments are conducted on large scale image memorability data set LaMem. From the experimental results, it is observed that the proposed model performs better than the current state-of-the-art model [3] by 2.44%.
Sathisha Basavaraju, Arijit Sur
IEEE Trans. Comput. Soc. Syst.2
2019 Dual-Domain Single Image De-Raining Using Conditional Generative Adversarial Network
abstract
This paper presents a novel method for a single image rain streak removal problem which exploits the spatial as well as wavelet transformed coefficients of the rainy images. The proposed method adopts the Conditional Generative Adversarial Network [1] framework and consists of two following networks: Generator and Discriminator. The generator model receives the input from both spatial, frequency domain of the rainy image and yields five de-rained image candidates. A Deep Residual Network [2] has been used to merge these derained candidates and predict a single de-rained image. To ensure the visual quality of the de-rained image, Perceptual loss function [3] in addition to adversarial training has been incorporated. Extensive experiments on the synthetic and realworld rainy images dataset reveal an improvement over the existing state-of-the-art methods [4], [5] by ~ 1.08%, 2.57% in Structural Similarity Index [6] and ~ 7.39%, 9.95% in Peak signal-to-noise ratio respectively.
Prasen Kumar Sharma, Priyankar Jain, Arijit Sur
ICIP3
2019 Memorability based image to image translation
abstract
This paper presents a memorability based image-to-image translation technique to make an image more memorable while retaining its high-level contents. Conventionally, the image-to-image translation task aims to learn the mapping between images of two different domains using a set of aligned image pairs. However, dataset having such one-to-one mapping is not available for memorability based image-to-image translation. Therefore, the aim of the proposed task is defined to learn the mapping F: I → I' between two image domains I and I'. Here, I corresponds to input image domain and I' is the unknown image domain containing the modified version of the input images. Also, every image in I' is more memorable than its corresponding image in I. Therefore, the proposed task is achieved by developing a deep learning based method to learn the mapping F: I→ I' using mean-squared error and memorability loss between I and F(I). The experimental results showed that the proposed approach increases the memorability of the given image better than the state-of-the-art image-to-image translation techniques.
Sathisha Basavaraju, Prasen Kumar Sharma, Arijit Sur
ICMV3
2019 Object Memorability Prediction using Deep Learning: Location and Size Bias
Sathisha Basavaraju, Sibaji Gaj, Arijit Sur
J. Vis. Commun. Image Represent.3
2019 Multiple instance learning based deep CNN for image memorability prediction
Sathisha Basavaraju, Arijit Sur
Multim. Tools Appl.2
2019 View invariant DIBR-3D image watermarking using DT-CWT
abstract
In 3D image compression, depth image based rendering (DIBR) is one of the latest techniques where the center image (say the main view, is used to synthesise the left and the right view image) and the depth image are communicated to the receiver side. It has been observed in the literature that most of the existing 3D image watermarking schemes are not resilient to the view synthesis process used in the DIBR technique. In this paper, a 3D image watermarking scheme is proposed which is invariant to the DIBR view synthesis process. In this proposed scheme, 2D-dual-tree complex wavelet transform (2D-DT-CWT) coefficients of centre view are used for watermark embedding such that shift invariance and directional property of the DT-CWT can be exploited to make the scheme robust against view synthesis process. A comprehensive set of experiments has been carried out to justify the robustness of the proposed scheme over the related existing schemes with respect to the JPEG compression and synthesis view attack.
Shuvendu Rana, Arijit Sur
Multim. Tools Appl.2
2019 Client-Side QoE Management for SVC Video Streaming: An FSM Supported Design Approach
abstract
HTTP adaptive streaming (HAS) which provides the flexibility of video bit-rate adjustments, is becoming the de facto framework for live and on-demand video streaming services. The flexibility of HAS is further empowered by the scalable video coding (SVC) technique which allows low overhead quality upgradations of transmitted video segments. In this paper, we present a client-side finite-state machine (FSM) supported quality of experience (QoE) centric video bit-rate adaptation and management mechanism. The proposed strategy simultaneously manages three important QoE verticals: 1) providing stutter-free video viewing experience; 2) minimizing flickers in video outputs by controlling and smoothing the rates of encoding quality switches over time; and 3) maximizing aggregate video quality over a video playout session. Based on careful consideration of these QoE verticals, the management policy dynamically decides whether to download a new segment or upgrade the quality of an already downloaded segment in the playout buffer. We have implemented and evaluated the performance of our proposed framework in a multi-client SVC video steaming test-bed. Conducted experiments using real-world network traces reveal that the proposed strategy is able to outperform other state-of-the-art adaptive streaming techniques and deliver satisfactory QoE even in the face of highly varying channel conditions.
Rajesh Devaraj, Arnab Sarkar 0001, Arijit Sur
IEEE Trans. Netw. Serv. Manag.4
2018 Image Memorability: The Role of Depth and Motion
abstract
Understanding what makes an image memorable or forgettable is one of the interesting problems in the domain of computer vision and cognitive science. The initial studies have demonstrated that memorability is an inherent characteristic of an image and have found many high-level image properties (such as emotions, saliency, object statistics, popularity, aesthetics, etc.) which influence image memorability. This paper sheds light on the relationship between image memorability and two image features: motion and depth which, to the best of our knowledge, is still unexplored. Experimental analysis reveals that motion and depth cues have a positive influence in determining image memorability. Further, the paper presents a novel deep learning model, FOD-MemNet, to exploit motion and depth cues along with object features to predict image memorability. Experimental results demonstrate that the proposed FOD-MemNet model outperforms the current state-of-the-art model by achieving a rank correlation of 0.655 which is near to human consistency$(\rho=0.68)$.
Sathisha Basavaraju, Paritosh Mittal, Arijit Sur
ICIP3
2017 A resource allocation framework for adaptive video streaming over LTE
Arnab Sarkar 0001, Arijit Sur
J. Netw. Comput. Appl.3
2017 SIFT based video watermarking resistant to temporal scaling
Nilkanta Sahu, Arijit Sur
J. Vis. Commun. Image Represent.2
2017 A robust watermarking scheme against frame blending and projection attacks
Sibaji Gaj, Anoop Kumar Rathore, Arijit Sur, Prabin Kumar Bora
Multim. Tools Appl.3
2016 A drift compensated reversible watermarking scheme for H.265/HEVC
abstract
In this paper, a compressed domain drift compensated reversible watermarking scheme is proposed with a high embedding capacity and the least amount of visual quality degradation for H.265/HEVC videos. Using compressed domain syntax elements, such as motion vector and transformed residual, a set of 4 × 4 Transform Blocks (TB) of similar texture are chosen from consecutive I Frames for watermark embedding. Due to texture similarity of these selected TBs, the differences between the transformed coefficients are equal or close to zero. Utilizing this difference statistics, a multilevel watermarking is inserted in the compressed video by altering near zeros values in the difference transformed coefficients. A comprehensive set of experiments have been carried out to justify the efficacy of the proposed scheme over existing literature.
Sibaji Gaj, Shuvendu Rana, Arijit Sur, Prabin Kumar Bora
MMSP3
2016 Segmentation based 3D depth watermarking using SIFT
abstract
In this paper, a 3D image watermarking scheme is proposed to embed the watermark with the depth of the 3D image for depth image based rendering (DIBR) 3D image representation. To make the scheme invariant to view synthesis process, watermark is inserted with the scale invariant feature transform (SIFT) feature point locations obtained from the original image. Moreover, embedding zone for watermarking has been selected in such a way that no watermark can be inserted in the foreground object to avoid perceptible artefacts. Also, a novel watermark embedding policy is used to insert the watermark with the depth of the 3D image to resist the image processing attacks. A comprehensive set of experiments are carried out to justify the robustness of the proposed scheme.
Shuvendu Rana, Sibaji Gaj, Arijit Sur, Prabin Kumar Bora
MMSP3
2016 Detection of fake 3D video using CNN
abstract
In this paper, a novel automatic fake and the real 3D video recognition scheme is proposed to distinguish the 3D video converted from the 2D video using 2D to 3D conversion process (say fake 3D) from the 3D video captured using direct capturing of the 3D camera (say real 3D). To identify the real and fake 3D, pre-filtration is done using the dual tree complex wavelet transform to emerge the edge and vertical and horizontal parallax characteristics of real and fake 3D videos. Convolution neural network (CNN) is used to train the 3D characteristics to distinguish the fake 3D videos from the real ones. A comprehensive set of experiments has been carried out to justify the efficacy of the proposed scheme over the existing literature.
Shuvendu Rana, Sibaji Gaj, Arijit Sur, Prabin Kumar Bora
MMSP3
2016 A three level adaptive video streaming framework over LTE
abstract
The ever increasing demand for high bandwidth, low latency multimedia applications on mobile devices is set to pose a considerable challenge on the bandwidth allocation and multiplexing mechanisms in LTE and future wireless networks. This paper proposes a low overhead Scalable Video Coding (SVC) based dynamic adaptive streaming framework (called TLS-AV) which attempt to maintain a minimum satisfactory Quality of Experience (QoE) for all end users even during transient network overloads. Two fast and efficient radio resource allocation heuristics namely, the TLS-AV Water-Filling Heuristic (TWH) and TLS-AV Balanced Water-Filling Heuristic (TBWH) have been implemented over the proposed framework. Experimental results reveal that both the developed heuristics are able to restrict packet loss rate below ~ 10% while simultaneously achieving high video qualities and fairness among the transmitted qualities of video flows, on average.
Santhosh Sriram, Arnab Sarkar 0001, Arijit Sur
SMC4
2016 Object based watermarking for H.264/AVC video resistant to rst attacks
Sibaji Gaj, Ashish Singh Patel, Arijit Sur
Multim. Tools Appl.3
2016 MCDCT-TF based video watermarking resilient to temporal and quality scaling
Nilkanta Sahu, Shuvendu Rana, Arijit Sur
Multim. Tools Appl.3
2016 Drift-Compensated Robust Watermarking Algorithm for H.265/HEVC Video Stream
abstract
It has been observed in the recent literature that the drift error due to watermarking degrades the visual quality of the embedded video. The existing drift error handling strategies for recent video standards such as H.264 may not be directly applicable for upcoming high-definition video standards (such as High Efficiency Video Coding (HEVC)) due to different compression architecture. In this article, a compressed domain watermarking scheme is proposed for H.265/HEVC bit stream that can handle drift error propagation both for intra- and interprediction process. Additionally, the proposed scheme shows adequate robustness against recompression attack as well as common image processing attacks while maintaining decent visual quality. A comprehensive set of experiments has been carried out to justify the efficacy of the proposed scheme over the existing literature.
Sibaji Gaj, Aditya Kanetkar, Arijit Sur, Prabin Kumar Bora
ACM Trans. Multim. Comput. Commun. Appl.3
2016 Depth-Based View-Invariant Blind 3D Image Watermarking
abstract
With the huge advance in Internet technology as well as the availability of low-cost 3D display devices, 3D image transmission has become popular in recent times. Since watermarking has become regarded as a potential Digital Rights Management (DRM) tools in the past decade, 3D image watermarking is an emerging research topic. With the introduction of the Depth Image-Based Rendering (DIBR) technique, 3D image watermarking is a more challenging task, especially for synthetic view generation. In this article, synthetic view generation is regarded as a potential attack, and a blind watermarking scheme is proposed that can resist it. In the proposed scheme, the watermark is embedded into the low-pass filtered dependent view region of 3D images. Block Discrete Cosine Transformation (DCT) is used for spatial-filtration of the dependent view region to find the DC coefficient with horizontally shifted coherent regions from the left and right view to make the scheme robust against synthesis view attack. A comprehensive set of experiments have been carried out to justify the robustness of the proposed scheme over related existing schemes with respect to Stereo JPEG compression and different noise addition attacks.
Shuvendu Rana, Arijit Sur
ACM Trans. Multim. Comput. Commun. Appl.2
2015 A three level LTE downlink scheduling framework for RT VBR traffic
Arnab Sarkar 0001, Santhosh Sriram, Arijit Sur
Comput. Networks4
2015 Steganographic algorithm based on randomization of DCT kernel
Swathi Karri, Arijit Sur
Multim. Tools Appl.2
2015 Robust watermarking for resolution and quality scalable video sequence
Shuvendu Rana, Nilkanta Sahu, Arijit Sur
Multim. Tools Appl.3
2015 Detection of motion vector based video steganography
Arijit Sur, Sista Venkat Madhav Krishna, Nilkanta Sahu, Shuvendu Rana
Multim. Tools Appl.1
2014 SIFT based robust image watermarking resistant to resolution scaling
abstract
In this paper, a novel image watermarking scheme is proposed which is robust to resolution scaling. In the proposed method, SIFT (Scale Invariant Feature Transform) features are used which are invariant to rotation, translation and scaling. Most of the existing watermarking algorithms that uses SIFT features, fail to retain the watermark when the image is heavily scaled down. In the proposed scheme, a context coherent object or patch is inserted in the image such that it generates strong SIFT features. These newly generated SIFT features themselves work as a watermark. Since the SIFT features are invariant to scaling, these features can be extracted from any image resolution with high probability. The experimental results support the contention that the proposed method is considerably robust against resolution scaling.
Priyatham Bollimpalli, Nilkanta Sahu, Arijit Sur
ICIP3
2014 Pixel rearrangement based statistical restoration scheme reducing embedding noise
Arijit Sur, Vignesh Ramanathan, Jayanta Mukhopadhyay
Multim. Tools Appl.1
2014 An image steganographic algorithm based on spatial desynchronization
Arijit Sur, Devadeep Shyam, Piyush Goel, Jayanta Mukhopadhyay
Multim. Tools Appl.1
2013 MCRD: Motion coherent region detection in H.264 compressed video
abstract
One of the challenging issues in video watermarking is robustness against collusion attacks. To resist collusion attack, the same video scene should carry the same watermark, whenever and wherever it appears in the video. Inter-frame correlation is more within a short video neighborhood when no scene change is detected. Motion coherency has recently been recognized as a desirable property for watermarks to resist temporal frame averaging attacks. To the best of our knowledge, motion coherent watermarking in compressed domain has not yet been well explored. In this paper, a compressed domain technique has been proposed to detect motion coherent regions within a short video neighborhood. The time complexity of the proposed method is discussed. Simulation results evaluate the effectiveness of the proposed method.
Tanima Dutta, Arijit Sur, Sukumar Nandi
ICME2
2008 A New Approach for Reducing Embedding Noise in Multiple Bit Plane Steganography
Arijit Sur, Piyush Goel, Jayanta Mukhopadhyay
ICISP1
2008 A Novel Steganographic Algorithm Resisting Targeted Steganalytic Attacks on LSB Matching
Arijit Sur, Piyush Goel, Jayanta Mukhopadhyay
IWDW1