Rajiv Ranjan Sahay

dblp:83/1983 · also Rajiv R. Sahay · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0003-0820-0616ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Adversarial learning for unguided single depth map completion of indoor scenes
Moushumi Medhi, Rajiv Ranjan Sahay
Mach. Vis. Appl.2
2024 Hand over Face Gesture Classification with Feature Driven Vision Transformer and Supervised Contrastive Learning
Kankana Roy, Aparna Mohanty, Rajiv Ranjan Sahay
ICPR (4)3
2024 Hand Segmentation With Dense Dilated U-Net and Structurally Incoherent Nonnegative Matrix Factorization-Based Gesture Recognition
abstract
Robust segmentation of hands in a cluttered environment for hand gesture recognition has remained a challenge in computer vision. In this work, a two-stage gesture recognition framework is proposed. In the first stage, we segment hands using the proposed deep learning algorithm, and in the second stage, we use these segmented hands to classify gestures using a novel structurally incoherent nonnegative matrix factorization approach. We propose a new deep learning framework for hand segmentation called densely dilated U-Net. We exploit recently proposed dense blocks and dilated convolution layers in our work. To cope with the scarcity of labeled datasets we extend our densely dilated U-Net for semisupervised hand segmentation using hand bounding boxes as cues. We provide quantitative and qualitative evaluation of proposed hand segmentation model on several public hand segmentation datasets including EgoHands, GTEA, EYTH, EDSH, and HOF. Semisupervised segmentation results are also obtained on two hand detection datasets including VIVA and CVRR. As an extension of our work, we show semisupervised segmentation and gesture recognition results using segmented hands on NUS-II cluttered hand gesture dataset. To validate the efficiency of our semisupervised algorithm we evaluate it on OUHands dataset with real ground truth labels. For gesture classification, we propose a novel structurally incoherent nonnegative matrix factorization algorithm. We propose to use CNN features extracted from segmented images for nonnegative matrix factorization. Experimental results on NUS-II and OUHands datasets demonstrate that our two-stage approach for gesture recognition yields superior results.
Kankana Roy, Rajiv Ranjan Sahay
IEEE Trans. Hum. Mach. Syst.2
2024 Robust static hand gesture recognition: harnessing sparsity of deeply learned features
Aparna Mohanty, Kankana Roy, Rajiv Ranjan Sahay
Vis. Comput.3
2023 Distill-DBDGAN: Knowledge Distillation and Adversarial Learning Framework for Defocus Blur Detection
abstract
Defocus blur detection (DBD) aims to segment the blurred regions from a given image affected by defocus blur. It is a crucial pre-processing step for various computer vision tasks. With the increasing popularity of small mobile devices, there is a need for a computationally efficient method to detect defocus blur accurately. We propose an efficient defocus blur detection method that estimates the probability of each pixel being focused or blurred in resource-constraint devices. Despite remarkable advances made by the recent deep learning-based methods, they still suffer from several challenges such as background clutter, scale sensitivity, indistinguishable low-contrast focused regions from out-of-focus blur, and especially high computational cost and memory requirement. To address the first three challenges, we develop a novel deep network that efficiently detects blur map from the input blurred image. Specifically, we integrate multi-scale features in the deep network to resolve the scale ambiguities and simultaneously modeled the non-local structural correlations in the high-level blur features. To handle the last two issues, we eventually frame our DBD algorithm to perform knowledge distillation by transferring information from the larger teacher network to a compact student network. All the networks are adversarially trained in an end-to-end manner to enforce higher order consistencies between the output and the target distributions. Experimental results demonstrate the state-of-the-art performance of the larger teacher network, while our proposed lightweight DBD model imitates the output of the teacher network without significant loss in accuracy. The codes, pre-trained model weights, and the results will be made publicly available.
Sankaraganesh Jonna, Moushumi Medhi, Rajiv Ranjan Sahay
ACM Trans. Multim. Comput. Commun. Appl.3
2022 A robust multi-scale deep learning approach for unconstrained hand detection aided by skin segmentation
Kankana Roy, Rajiv Ranjan Sahay
Vis. Comput.2
2021 Robust depth map inpainting using superpixels and non-local Gauss-Markov random field prior
Sukla Satapathy, Rajiv Ranjan Sahay
Signal Process. Image Commun.2
2021 A Non-Local Superpatch-Based Algorithm Exploiting Low Rank Prior for Restoration of Hyperspectral Images
abstract
We propose a novel algorithm for the restoration of a degraded hyperspectral image. The proposed algorithm exploits the spatial as well as the spectral redundancy of a degraded hyperspectral image in order to restore it without having any prior knowledge about the type of degradation present. Our work uses superpatches to exploit the spatial and spectral redundancies. We formulate a restoration algorithm incorporating structural similarity index measure as the data fidelity term and nuclear norm as the regularization term. The proposed algorithm is able to cope with additive Gaussian noise, signal dependent Poisson noise, mixed Poisson-Gaussian noise and can restore a hyperspectral image corrupted by dead lines and stripes. As we demonstrate with the aid of extensive experiments, our algorithm is capable of recovering the spectra even in the case of severe degradation. A comparison with the state-of-the-art low rank hyperspectral image restoration methods via experiments with real world and simulated data establishes the competitiveness of the proposed algorithm with the existing methods.
Sourish Sarkar, Rajiv Ranjan Sahay
IEEE Trans. Image Process.2
2018 Rasabodha: Understanding Indian classical dance by recognizing emotions using deep learning
Aparna Mohanty, Rajiv Ranjan Sahay
Pattern Recognit.2
2018 Removal of Eye Blink Artifacts From EEG Signals Using Sparsity
abstract
Neural activities recorded using electroencephalography (EEG) are mostly contaminated with eye blink (EB) artifact. This results in undesired activation of brain-computer interface (BCI) systems. Hence, removal of EB artifact is an important issue in EEG signal analysis. Of late, several artifact removal methods have been reported in the literature and they are based on independent component analysis (ICA), thresholding, wavelet transformation, etc. These methods are computationally expensive and result in information loss which makes them unsuitable for online BCI system development. To address the above problems, we have investigated sparsity-based EB artifact removal methods. Two sparsity-based techniques namely morphological component analysis (MCA) and K-SVD-based artifact removal method have been evaluated in our work. MCA-based algorithm exploits the morphological characteristics of EEG and EB using predefined Dirac and discrete cosine transform (DCT) dictionaries. Next, in K-SVD-based algorithm an overcomplete dictionary is learned from the EEG data itself and is designed to model EB characteristics. To substantiate the efficacy of the two algorithms, we have carried out our experiments with both synthetic and real EEG data. We observe that the K-SVD algorithm, which uses a learned dictionary, delivers superior performance for suppressing EB artifacts when compared to MCA technique. Finally, the results of both the techniques are compared with the recent state-of-the-art FORCe method. We demonstrate that the proposed sparsity-based algorithms perform equal to the state-of-the-art technique. It is shown that without using any computationally expensive algorithms, only with the use of over-complete dictionaries the proposed sparsity-based algorithms eliminate EB artifacts accurately from the EEG signals.
S. R. Sreeja, Rajiv Ranjan Sahay, Debasis Samanta, Pabitra Mitra
IEEE J. Biomed. Health Informatics2
2017 Stereo image de-fencing using smartphones
abstract
Conventional approaches to image de-fencing have limited themselves to using only image data in adjacent frames of the captured video of an approximately static scene. In this work, we present a method to harness disparity using a stereo pair of fenced images in order to detect fence pixels. Tourists and amateur photographers commonly carry smartphones/phablets which can be used to capture a short video sequence of the fenced scene. We model the formation of the occluded frames in the captured video. Furthermore, we propose an optimization framework to estimate the de-fenced image using the total variation prior to regularize the ill-posed problem.
Sankaraganesh Jonna, Sukla Satapathy, Rajiv Ranjan Sahay
ICASSP3
2016 Nrityabodha: Towards understanding Indian classical dance using a deep learning approach
Aparna Mohanty, Pratik Vaishnavi, Prerana Jana, Anubhab Majumdar, Alfaz Ahmed, Trishita Goswami, Rajiv Ranjan Sahay
Signal Process. Image Commun.7
2013 Seeing through the fence: Image de-fencing using a video sequence
abstract
Tourists and amateur photographers are often hindered in capturing their cherished images/videos by a fence/occlusion that limits accessibility to the scene of interest. The situation has been exacerbated by growing concerns of security at public places and a need exists to provide a tool that can be used for post-processing such “fenced videos” to produce a “de-fenced” image. There are several challenges in this problem and in this work, we identify them as 1. Robust detection of the fence/occlusions. 2. Estimating pixel motion of background scene. 3. Filling in the fence/occlusions by utilizing information in multiple frames of the input video. We use a video captured by a camera panning the scene containing a fence and obtain a “de-fenced” image. Our method can effectively remove fences from images as demonstrated for several synthetic and real-world cases.
Vrushali S. Khasare, Rajiv Ranjan Sahay, Mohan Kankanhalli
ICIP2
2011 Dealing With Parallax in Shape-From-Focus
abstract
We propose a new method that extends the capability of shape-from-focus (SFF) to estimate the depth profile of 3-D objects in the presence of structure-dependent pixel motion. Existing SFF techniques work under the constraint that there is no parallax in the captured stack of frames. However, in off-the-shelf cameras, there can be appreciable pixel motion among the observations when there is relative motion between the object and the camera. In such a scenario, the depth estimates will be erroneous if the parallax effect is not factored in. Our degradation model accounts for pixel migration effects in the observations due to parallax resulting in a generalization of the SFF technique. We show that pixel motion and defocus blur therein are tightly coupled to the underlying shape of the 3-D object. Simultaneous reconstruction of the underlying 3-D structure and the all-in-focus image is carried out within an optimization framework using local image operations. The proposed method when tested on many examples, both synthetic and real, is very effective and delivers state-of-the-art performance.
Rajiv Ranjan Sahay, A. N. Rajagopalan 0001
IEEE Trans. Image Process.1
2009 Inpainting in Shape from Focus: Taking a Cue from Motion Parallax
abstract
Shape from focus (SFF) which uses a sequence of space-variantly defocused frames works under the constraint that there is ‘no magnification’ in the stack. In the presence of sensor damage and/or occlusions, there will be missing data in the observations and SFF cannot recover structure in those regions. In many applications, the capability of fillingin missing data is of critical importance. In this paper, we investigate the effect of motion parallax in SFF and demonstrate the interesting possibility of how it can be judiciously used to jointly inpaint image and depth profiles. When there is relative motion between the 3D specimen and the camera, by virtue of the inherent pixel motion in each of the frames, it is possible to obtain a focused image and depth map of the scene despite missing regions in the observations.
Rajiv Ranjan Sahay, A. N. Rajagopalan 0001
BMVC1
2007 High Resolution Image Reconstruction in Shape from Focus
abstract
In the Shape from Focus (SFF) method, a sequence of images of a 3D object is captured for computing its depth profile. However, it is useful in several applications to also derive a high resolution focused image of the 3D object. Given the space-variantly blurred frames and the depth map, we propose a method to optimally estimate a high resolution image of the object within the SFF framework.
Rajiv Ranjan Sahay, A. N. Rajagopalan 0001
ICIP (2)1